Privacy Is Still the Web’s Unfinished Business

The Web is old enough to have its own quarter-life crisis. At 25, nytimes.com predates some of the people who build it. Wikipedia has hit 20, and the first browser shipped three decades ago. With more than 4 billion users and nearly 2 billion websites, you would think the fundamental questions about how people and digital technology interact would have settled answers by now. They don’t. If anything, the open questions are accumulating faster than the industry can address them.

Among the most surprising of these unresolved issues is privacy. Offline, we make dozens of privacy judgments daily — whether to read a stranger’s messages over their shoulder, whether to repeat a friend’s intimate disclosure, whether a doctor can share symptoms with an employer — and we rarely hesitate. The context gives us the cues. Online, those cues all but disappear.

Why the Digital Context Breaks Down

Several forces make privacy on the Web harder to reason about than privacy on the street. The first is that privacy is deeply contextual. We instinctively calibrate what we share based on where we are: work, home, a subway car, a doctor’s office. A single glossy slab of plastic now serves as the venue for all of those settings at once. We chat with friends, draft work documents, and look up medical symptoms on the same device, which leaves us without reliable signals for what counts as appropriate.

The second complicating factor is the ubiquity — and ambiguity — of third parties. Modern digital products rarely ship without relying on outside services, and that alone is not a privacy problem. A specialized vendor may well have stronger data protection than an in-house team, and some third parties operate strictly on behalf of the site that hired them, never reusing user data. Those parties are, for privacy purposes, indistinguishable from the first party. Fathom, for instance, is a third-party analytics tool whose anonymization method is publicly documented.

Other third parties insist on acting as independent controllers of the data they collect. They take information from a site’s users and repurpose it for their own entirely separate interests. These are clear privacy violations, yet browsers give users no mechanism to tell the two kinds of third parties apart. Counting “trackers” in an ad blocker tells you little about how privacy-invasive a site actually is. Without a way for browsers to distinguish a loyal service provider from an independent data collector, automated protection at the browser level remains out of reach.

The third factor is the one nobody should pretend away: there is real money in the confusion. Some of the largest technology companies — and many smaller ones — run business models that convert privacy violations into revenue. They have an interest in muddying the distinction between privacy and security, in championing complex consent screens and privacy dashboards, and in keeping the conversation noisy enough that informed improvement crawls forward at a glacial pace. That can be paralyzing, but it shouldn’t be. Not all data collection is a privacy violation, and meaningful progress can come one website at a time.

A practical approach is to begin with a familiar offline context and reason outward from its established norms. There is rarely a perfect match between a physical setting and a website, but starting with a concrete situation structured by everyday expectations is a reliable way to decide where your own line should be.

A Bookshop Thought Experiment

Consider a physical bookshop as a comparison point for an online publication. In a bookshop, the staff sees you walk in. They may recognize you from previous visits and know your name. When you browse, they notice which section you’re in and whether you pull a particular volume off the shelf. They may use that to offer recommendations. If you make a purchase, a payment processor learns a bit about you.

Now, that bookshop installs CCTV cameras. You are a little uncomfortable, but the owner explains that the video runs on a closed local circuit, never leaves the shop, is retained for at most 24 hours, and is watched only to identify theft. Limited access, short retention, a clear and narrow purpose. Assuming you trust the owner, the guarantees are strong and the intrusion is bounded.

But then you notice other, smaller cameras placed throughout the store. These are positioned to capture which books you consider, which blurbs you read, and which titles you ultimately buy. The conversation with the owner goes differently this time. Those feeds go out to several companies. In exchange, they help cover some of the shop’s costs or list it on neighborhood maps. You agreed to all of it, the owner says, the moment you pushed the door open and stepped inside.

The owner insists it is anonymous. The monitoring companies, they explain, use only a hash of your facial biometrics to recognize you from shop to shop. What do you have to hide? It is essential to keeping books affordable; without it, only the wealthy could read. Anything else would be bad for small businesses and the poor. Scanning the list of companies receiving video feeds, you recognize none of them — except for two that receive feeds from every other shop in town as well.

Press the owner further, and you find they are unhappy with those two firms. The companies use the data they collect to compete with the bookshop directly, selling books themselves and recommending rival stores. They also leverage their position to push a model where they provide the infrastructure for all retailers — why should a bookshop owner craft the browsing experience when their value is merely the selection? — and a growing number of shops have adopted that global strip-mall framework. The owner sees it for the threat it is. “When you don’t comply they drop you off the map and send people to your competitors.”

This last piece is slightly contrived, but it is remarkably close to the operational reality of a commercial publisher or online store on today’s Web.

Pick a Default, Set Expectations

Different people will draw their line at different points in this scenario. Some are fully comfortable with the erosion of privacy and trust the promised future of big-tech bureaucracy. Others would rather shop in a staff-less store with no human ever seeing what they browse.

For many, the line falls somewhere near the introduction of the CCTV system. The closed-circuit system, if truly limited in access, retention, and purpose, is bearable when it keeps a beloved source of books alive. Pervasive behavioral surveillance, whose main effect is letting a handful of large corporations absorb or starve out small shops, crosses the threshold. Which position is right is not the point. The point is that a site’s users deserve a default and expectations that match it.

A workable starting default is what might be called the Vegas Rule: what happens on the site stays on the site. That includes the parties working for the site, when they work solely for the site. For the majority of sites performing a tightly scoped set of related services, that is a reasonable baseline to build from. And it is one we can move toward incrementally — taking each site’s privacy properties step by step until they resemble the norms we apply offline without hesitation.

Where Privacy Work Actually Starts

For teams building websites, privacy isn’t a switch you flip. It’s an incremental process. A single “Big Bang” cleanup that removes every third-party script at once is rarely realistic, especially on commercial sites where analytics and advertising are tied to revenue. The practical path is to improve steadily over time.

The first move is diagnosis: understand what data is being collected, by whom, and why. If you’re on the engineering side, it’s easy to resent the marketing team for every new “pixel” that slows down your pages. But those tags usually exist because someone is trying to measure campaign performance or attribute conversions — goals that keep the business running. Instead of complaining, treat this as a chance to build a working relationship. Marketing people often can’t tell a legitimate analytics vendor from a data broker, and they rarely understand the technical mechanics of what they’re buying. You can help with that. If you’re an ally rather than an adversary, you may not eliminate every tracker, but you can start making real cuts.

Once you have credibility (or the decision-making power yourself), hold every remaining third party accountable. The ones that stay should be provably effective. At The Times, working closely with the marketing team cut the amount of audience data shared with third-party controllers by over 90 percent — a win for privacy and for page performance. After that, develop a habit of skepticism. That free social-sharing widget? Likely a data broker. A third-party comments system? Check whether it monetizes user data. Vendors often strike deals to inject additional scripts into your pages, a practice called piggybacking. Run your site through a tool like Blacklight periodically to catch surprises.

Google note asking for key points of Google’s Privacy Policy to be reviewed
Just because your site visitors won’t read the fine print doesn’t mean you shouldn’t either.

There’s also a business case that’s easy to overlook. Audience data is a strategic asset. When you hand it to a social network so you can retarget visitors with ads, that same data — which reveals who is interested in your products — is used to show them your competitors’ ads. If you sell shoes and give a platform a map of shoe shoppers, expect to see rival shoe ads appearing in front of those users. Host your storefront on a marketplace run by a company that also sells its own goods? The data you share about your customers is training data for your competitor. Privacy is not just an ethical stance; it’s sound business strategy when you’re participating in the data economy.

Fixing the Platform

The Web’s core promise is trust: you can click from site to site because your browser protects you from malicious code. That promise holds for security, but it has been broken for privacy. Browsers, however, are pushing back. Most have shipped meaningful tracking prevention, and Chrome — the holdout — has committed to changes. The Global Privacy Control is gaining adoption. As third-party cookies die out, industry groups are proposing standards that let businesses function without wholesale surveillance.

Notable proposals include Apple’s Private Click Measurement, Google’s FLEDGE, Microsoft’s PARAKEET, and The New York Times’s Garuda, with coordination happening in the W3C Privacy CG. Some ideas have stumbled — FLoC ran into trouble — which only highlights the need for a Web community that understands privacy deeply enough to build better solutions.

Those solutions won’t come from engineering alone. They require technologists and policymakers to work together. Technologists can show up as citizens in policy debates, explain complex systems in plain language, and think critically about how their creations behave at scale. Most importantly, when current designs fail people, we should be able to imagine other ways of doing things. Technology is rarely inevitable; it’s a pile of accumulated choices, and we can choose differently.

The goal is to build for users, not to extract from them. Most commercial sites won’t achieve perfect privacy overnight, but that isn’t a reason to stand still. The direction has shifted — a more private Web is now a plausible future, and it’s within our power to build it.

Further Reading