BLACK HAT FIRESIDE CHAT: Websites now leak sensitive data by design – AI is making it worse

By Byron V. Acohido

Public-facing web pages aren’t built in-house anymore. They’re composed from components supplied by outsiders, and the advertising platforms have had their run of them. The exposure is enormous, and AI is accelerating it all.

Related: No easy fixes for AI risk

The bill for all of it is arriving now. Since 2023, hospitals and health systems have paid more than $100 million to settle claims that tracking pixels on their websites leaked patient data to advertising platforms. A court approved $21.5 million against Sutter Health in February. Inova settled for $3.1 million in April. Atrium settled for $1.8 million this month. In each of these cases the pixels came from Meta or Google, and the health system paid the settlement.

Those numbers have teeth. On July 14 a federal judge refused to dismiss claims against Google and Meta over prescription data collected by pixels on a telehealth site, rejecting the argument that users had consented, and pushing both companies into discovery.

Ribeiro

I spoke with Rui Ribeiro, CEO and co-founder of Jscrambler, ahead of Black Hat USA 2026. We talked for the better part of an hour about how companies developed a huge blind spot about exposing sensitive assets via their web pages, and the sharpest stretches are in the accompanying podcast. Please give it a listen.

Fifteen years ago a company built its own public-facing site, in-house or with a few contractors, and could account for every piece of it. Then Chrome made rich, dynamic pages cheap, and companies began bolting on components they did not build. Today a single page is assembled from dozens of outside suppliers, and the browser is where the business happens. Next, AI generates the page and decides on its own what to pull in. At every step the owner had less control over what ran on the page.

Accumulated exposures

What no one accounted for was the access granted to the web page component vendors. A company might bring in one solution for payments, another for retargeting ads, and yet another for a user chatbot. “All of a sudden you have 100 vendors on your website,” Ribeiro said. The practice has been that each one gets handed more access than its job requires, and that surplus is where the exposure lives. Nobody set out to create it. It accumulated, one afternoon at a time.

Adding to this problem, the advertising platforms looked at the same loosening control and spotted an opening.

The ad platforms came for whatever wasn’t locked down. This could be accomplished with a few lines of code, called a pixel, dropped onto any public-facing webpage, usually through a tag manager, usually without a security review. Every time a customer loads that page, the pixel sends a message home describing what is on the screen and who is looking at it.

The hospital lawsuits show how this plays out. A marketing team adds Meta’s pixel to the health system’s website to learn which ads bring patients in. A patient visits that site, signs into the portal, and schedules an oncology appointment. The pixel is sitting on that page. It reports back to Meta which page the patient is on, what they clicked, and an identifier that ties to their Facebook account. Meta now knows a named person booked cancer care. Nobody at the hospital decided to send that. Nobody told the pixel not to. The patients sued the hospital, and the hospital paid.

Jscrambler ran its own study of Meta and TikTok pixels on e-commerce sites and found them reaching for customer email addresses, phone numbers, the products people browsed and the prices they paid. Ribeiro’s summary: the who, the what and the when.

Jscrambler has since extended that runtime research into banking. Across 14 financial-services cases in Europe and the United States, it found tracking operating without a valid consent choice at nine companies, with data flowing to roughly a dozen third parties.

The same pattern runs well beyond banking. He walked me through where that kind of personally identifiable data ends up. A patient books an age-related cancer screening. There is no diagnosis, only an appointment. The platforms watching that page infer one anyway. “They can either use it to advertise miracle cures to you,” he said, “or it can be used down the road to increase your policies, block you from having access to insurance.”

Suppliers and data harvesters

This is the part that separates a supplier from a data harvester. A misconfigured video player leaks by accident. A pixel collects because collecting is the product.

The banking research shows how easily the distinction gets buried inside the machinery. In one Portuguese bank’s credit simulator, a customer selected essential cookies only. But when the journey moved into an embedded frame on another part of the site, the consent choice did not move with it. Pinterest, LinkedIn, TikTok, Meta and Google/DoubleClick received the customer’s loan purpose, requested amount, repayment term, monthly payment and total cost.

The visitor had made a choice. The page architecture failed to honor it.

Consent is the reason all of this has been allowed to unfold. Consent, in this context, is a trick.

It works like this. A consent prompt appears. Accept is one button. Declining means opening a menu, reading through a vendor list and throwing switches one at a time. A decade of that has conditioned American consumers to click accept and get on with their day, which is what the design is for. The click gets recorded. So if and when a customer should sue, that record is the first thing Meta and Google put in front of a judge. See, your honor. They clicked accept.

The framework was written by the advertising industry, for the advertising industry. IAB Europe, the trade group for the sector, built the Transparency and Consent Framework that much of the ad business still runs on.

Three parties, two agree

There is a deeper problem with the consent prompt. Three parties are involved, and only two of them agree to anything.

The visitor. The platform. The company that owns the page.

The consent prompt is an agreement between the visitor and the platform. The visitor clicks accept. The platform starts collecting.

The company that owns the page is not part of that agreement. Yet the company’s proprietary business data leaves anyway. The pricing. The product mix. The margins. The patterns of who buys what, at what price, at what hour. Nobody at the company agreed to send any of it, and nobody was asked.

All of that was true before language-activated AI arrived in November 2022. AI dropped into an environment already spinning, and every party on the web page picked it up at once. Companies added AI chat and personalization. Their component suppliers rebuilt their products around AI. The ad platforms rebuilt their targeting engines around it. Every component on the page got thirstier for data at the same moment, because data is what makes the AI work.

“They will try to get much more data than they did before,” Ribeiro said.

Ribeiro is blunt about where that goes. “AI is going to pull whatever it decides into the supply chain.”

OWASP saw it coming. Its 2025 Top 10 introduced software supply chain failures at number three, ranked first by half the practitioners surveyed, with the highest average incidence rate in the data. OWASP also did something the industry had avoided, extending the definition past the server to the client side.

Compliance is lagging, as usual, but it is moving. PCI DSS 4.0.1 has required payment pages to inventory their scripts and detect unauthorized changes since March 31, 2025. It covers card data and stops there. It’s notable that the healthcare companies writing eight-figure checks were not breached at the payment page.

Ribeiro calls dealing with this non-negotiable, and his reason is not the one a board expects. Leak how a company operates, he argues, the prices it charges, the products that move, who its customers are, and someone replicates it, “All of a sudden you won’t have any business.”

An airline serving 100 million customers is serving 100 million different assemblies of code, each one built fresh in a stranger’s browser. That is the thing to govern, and almost nobody is governing it.

The tools and practices are available. The ball needs to get rolling. I’ll keep watch and keep reporting.

Acohido

Pulitzer Prize-winning business journalist Byron V. Acohido is dedicated to fostering public awareness about how to make the Internet as private and secure as it ought to be.

(Editor’s note: I used Claude and ChatGPT to assist with research compilation, source discovery, and early draft structuring. All interviews, analysis, fact-checking, and final writing are my own. I remain responsible for every claim and conclusion.)

 

Share on FacebookShare on Google+Tweet about this on TwitterShare on LinkedInEmail this to someone