Why Dead Code Stays

Most sites carry weight nobody uses: styles for layouts that were never built, fonts for icons that appear twice, scripts for features that were retired years ago. It stays for a simple reason. Adding something is easy to verify - it either shows up or it doesn't. Removing something is hard to verify, because the failure is somewhere else, on a page nobody thought to check.

So the usual approach is spot checks: open a few important pages, scroll, look, ship. That samples the site. It doesn't cover it, and the pages that break are rarely the ones in the sample.

What a Mechanical Check Looks Like

The alternative is to make the comparison mechanical. Capture every page before the change and after it, at more than one screen width, and compare the captures pixel by pixel. No judgment calls, no sampling - a page either matches or it shows exactly where it doesn't.

On this site, that is what made a large cleanup safe. The main stylesheet went from 283KB to 209KB minified, 37.8KB to 33.2KB gzipped, and every page was pixel-diffed before and after. About 21KB of that series was an unused 12-column grid. The icon font - about 220KB across one stylesheet and three font files - was replaced with inline SVG masks, and all 62 icon classes in the markup stayed exactly as they were.

None of that removal depended on remembering which pages used what. The comparison answered it.

What the Check Can't Tell You

A mechanical comparison proves one precise thing: nothing changed since the baseline. It can't say whether the baseline was right. If a regression is already in the "before," it's in the "after" too, and the two match perfectly.

That makes the choice of baseline the whole game. Take it from the last known-good state, before the first change in the series - not before the most recent one. A baseline taken partway through a cleanup quietly certifies everything that came before it.

Two Rules for Removing Things Safely

Inventory every place it can be referenced

Before removing or replacing something, find every way the site can reach it: classes and attributes in the markup, content values and URLs in the stylesheet, and anything a script builds at runtime. The obvious reference is rarely the only one.

Keep one baseline for a multi-step cleanup

A cleanup that happens in several passes needs one baseline, taken before the first pass, with every later pass compared against it. A fresh baseline per step turns each step's check into "nothing changed since the last unverified step."

The Caveat, From Experience

During that same cleanup, a pixel diff passed on every page while two visible regressions were live - blank gallery arrows and misaligned icon rows. The baseline had been captured after the change that caused them, so before and after matched. The full breakdown is in the Field Note The Diff That Compared Two Broken Pages.

Common Questions

Doesn't a pixel comparison flag harmless differences too?

It can. Animations, lazy-loaded images, and timing can make two captures of the same page differ. Freezing animations and re-running any page that differs usually separates real changes from noise, and a page that differs the same way twice is worth a look.

Is a screenshot comparison enough on its own?

It covers what a page looks like, not what it does. Links, forms, and scripts need their own checks - a button can look identical and still stop working.

How long should a baseline be kept?

Until the whole series of changes is verified and shipped. Once a new state is confirmed good, it can become the baseline for the next round of work.

Removal You Can Prove

A leaner site is only an improvement if nothing else moved. The proof is in the baseline.