The Cost of Knowing Everything: What Companies Actually Pay to Collect Your Data
Privacy policies tell you what companies collect. Nobody talks about what it costs them. Netflix spends $1-1.3B on AWS, 40% of stored data goes unused, and research shows 33-85% of user history is removable with only 2% quality loss. The economics of data collection are less rational than they appear.
In March, I wrote a teardown of Netflix’s privacy policy after downloading my own data. Seven CSV files. Three years of viewing history. 5,060 individual viewing events tracked down to the millisecond across 13 device types.
The consumer reaction to this kind of disclosure is usually some version of “they know too much about me.” That is a valid concern. But working in Data Architecture, I found myself asking a different question: what does it actually cost Netflix to collect, store, process, and eventually delete all of this?
The answer is more interesting than the outrage.
Update, September 2026. Three points sharpen the argument below. The EDPB’s final report on its right-to-erasure enforcement action (February 18, 2026) covered 764 controllers across 32 authorities and named seven recurring failures, among them treating anonymization as a substitute for deletion and leaving backup retention unresolved. Netflix’s second-quarter 2026 results kept the roughly $3 billion advertising target for 2026 discussed in the revenue section. And a June 2026 paper that analyzed 1,114 open-source Android apps found that leading LLMs reproduce data-minimization-risky code unless given explicit guidelines, which makes the minimization discipline a code-review problem as much as a policy one.
The Scale Nobody Talks About
Netflix’s Keystone pipeline processes over 2 trillion events per day. That is not a typo. Two trillion. The system ingests roughly 3 petabytes and outputs about 7 petabytes daily. Kafka carries approximately 1 million messages per second per topic, all Avro-encoded. This infrastructure scaled 20x over four years.
To run all of this, Netflix spends an estimated $1 billion to $1.3 billion annually on AWS. They operate over 100,000 server instances.
For context, Netflix’s content budget is $18 billion. Their data infrastructure represents about 6-7% of what they spend on shows and movies. It sounds like a rounding error until you remember it is still a billion dollars.
But here is the number that should make every Data Architect pause: Netflix’s own internal research found that at least 40% of stored data never gets used. Storage bills were growing 50% year over year while nearly half the data sat idle. They had to build an internal Garbage Collector (within their Baggins service) just to manage the lifecycle of unused file objects.
Forty percent. At Netflix’s scale, that is hundreds of millions of dollars in infrastructure supporting data that nobody queries, nobody models, and nobody needs.
How the Big Four Compare
Netflix is not an outlier. Every major platform faces this cost-collection tradeoff, though each handles it differently.
| Netflix | Spotify | YouTube | TikTok | |
|---|---|---|---|---|
| Users | 325M subscribers | 675M MAUs | 2.5B MAUs | 1.5B MAUs |
| Daily signals | 2T+ events | Not disclosed | 80B+ signals | Not disclosed |
| Cloud spend | $1-1.3B/year (AWS) | $150M+/year (GCP) | Proprietary (Google infra) | $1B+ (AWS/GCP) |
| Data granularity | Millisecond playtraces, scroll behavior, device fingerprints | Play/pause/skip, playlist creation, listening time | Watch time, completion rates, click patterns | Keystroke patterns, biometric data, clipboard content |
| Primary revenue from data | $1.5B ads + $1B churn savings | Free-tier ad targeting | $31.5B ad revenue | Ad targeting (undisclosed) |
| GDPR fines | EUR 4.75M (Netherlands) | SEK 58M / ~$5.4M (Sweden) | EUR 150M (France, Google cookie fine) | EUR 345M (Ireland) |
The YouTube figure is CNIL’s 2022 fine against Google for cookie-consent practices spanning google.fr and youtube.com. It is a cookie/ePrivacy penalty against Google, not a GDPR fine issued against YouTube specifically.
A few things stand out in this comparison.
TikTok collects the broadest scope of data. Keystroke patterns, biometric identifiers, clipboard content. It also faces the largest regulatory fines and the most government scrutiny. There is a correlation there.
YouTube processes the largest volume at 80 billion signals daily, but benefits from running on Google’s proprietary infrastructure rather than purchasing cloud services. Their $31.5B in ad revenue (2023) dwarfs everyone else, though the infrastructure cost is impossible to separate from Google’s overall spending.
Spotify collects similar behavioral categories to Netflix (play, pause, skip, search) but at lower granularity. No millisecond playtraces. No device fingerprinting. Their cloud spend, built around a 2018 three-year, roughly $450 million Google Cloud contract, was a fraction of Netflix’s AWS bill, though current spend is almost certainly higher. The question is whether that leaner approach produces meaningfully worse recommendations, and from my experience as a user of both, I would argue it does not.
Does More Data Actually Improve the Product?
This is the question that companies rarely answer honestly, because the answer is uncomfortable.
A 2025 research paper by Netflix engineers and a Kellogg School researcher quantified the engagement impact of different recommendation approaches:
| Recommendation Approach | Engagement Impact |
|---|---|
| Random recommendations | -16% engagement |
| Popularity-based algorithms | -12% engagement |
| Matrix factorization (2015-era baseline) | -4% engagement |
| Current personalized system | Baseline |
The jump from random to popularity-based is significant. The jump from popularity-based to modern personalization is meaningful. But the incremental gain from a 2015-era algorithm to today’s more sophisticated one is four percentage points. That gap is about algorithmic sophistication, not the volume of data behind it. How much data those systems actually need is a separate question, answered below.
Netflix’s own history provides the clearest illustration of diminishing returns. The Netflix Prize competition produced a winning ensemble of over 800 models from four teams. Netflix never deployed it. Their VP of Engineering explained: “The additional accuracy gains did not seem to justify the engineering effort needed to bring them into a production environment.” The winning ensemble’s marginal error-rate improvement over the already-deployed system was not worth the engineering complexity.
The most striking evidence comes from a 2025 data minimization study that directly tested how much user interaction data can be removed while maintaining recommendation quality:
- At 2% performance tolerance: 33-51% of user history data could be removed (EASE model), and 51-85% could be removed (simpler ItemKNN model)
- Data necessity is user-dependent: users with longer histories permitted proportionally greater reduction
- At perfect retention (0% tolerance): almost no data removal was possible for complex models
Read that again. In a feasibility study on standard recommender models, a third to four-fifths of user history data could be removed with only a 2% drop in recommendation quality. Netflix collects millisecond-precision playtraces, hover durations, scroll depth, and device network information at a comparable granularity. If a similar tradeoff holds on their production models, a large share of it could be dropped at a similar quality cost.
The Revenue Math: When Collection Pays for Itself
If more data does not dramatically improve recommendations, why do companies keep collecting it?
Because the value is not just in recommendations. It is in advertising, content strategy, and churn prevention.
Netflix generated $1.5 billion in ad revenue in 2025, more than double the prior year. They are targeting $3 billion in 2026. Its ad-supported tier reached 190 million monthly active viewers, and a majority of new sign-ups now choose the ad plan.
The behavioral data that barely moves the needle on recommendations is precisely what makes ad targeting valuable. Viewing patterns, time-of-day habits, content preferences, device information: all of this feeds audience segmentation for advertisers paying premium CPMs on the ad tier, reported well above typical social media rates.
Netflix’s own researchers put the recommendation system’s churn-reduction value at roughly $1 billion a year. Recommendations influence roughly 80% of hours streamed, by Netflix’s own accounting. When subscribers find what they want to watch quickly, they keep paying.
And then there is content commissioning. Netflix’s most famous data-driven bet was House of Cards: $100 million for two seasons without a pilot, based on data showing strong engagement with David Fincher films, Kevin Spacey content, and the British original. That data informs how $18 billion in annual content spending gets allocated.
The simplified equation: $1-1.3B in infrastructure, most of it video encoding and delivery rather than behavioral-data collection, enables $1.5B in ad revenue, $1B in churn savings, and smarter allocation of $18B in content decisions, which makes this an upper bound on what data collection alone could be buying. At Netflix’s scale, the ROI is clearly positive.
For practitioners: But this math has a denominator problem. Most companies collecting extensive behavioral data are not operating at Netflix’s scale, do not have an ad business, and cannot quantify a billion dollars in churn savings. They are paying the infrastructure costs without capturing the revenue upside.
The Bill Nobody Budgets For: Deletion
Collecting data is operationally simple. Deleting it is an engineering nightmare.
GDPR Article 17 gives EU residents the right to erasure. CCPA provides similar rights in California. When a user exercises that right, the company must find and delete their data across every system: production databases, analytics pipelines, recommendation models, backup systems, ad-tech integrations, clean rooms, and data lakes.
Gartner estimates the average cost of manually processing a single Data Subject Access Request at $1,524, and complex requests spanning many systems run far higher. Privacy request volumes have grown 246% in two years.
Netflix has responded by building a centralized data deletion platform: 1,300 datasets now fall under its management, with 76.8 billion rows deleted to date and no data loss incidents reported so far. Engineers presented the architecture at QCon SF in November 2025, treating “deletion as a first-class architectural concern.” That is a significant engineering investment on top of the collection infrastructure.
For practitioners: The deletion cost scales with the collection scope. Every new data point you collect creates a future liability: it must be stored, secured, governed, and eventually deletable on demand. The more systems that touch the data, the more expensive deletion becomes.
Then there is the problem nobody has solved: machine unlearning. ML models trained on user data retain learned patterns even after the source data is deleted. Naive retraining from scratch is computationally prohibitive for frequent deletion requests. Techniques like SISA (Sharded, Isolated, Sliced, Aggregated) enable faster selective retraining, but the field acknowledges that aggressive unlearning risks degrading model quality for remaining users.
The EDPB selected right to erasure as its 2025 enforcement priority, investigating 764 controllers across 32 European data protection authorities. Their finding: companies rely on anonymization as a substitute for actual deletion, and inconsistent practices around retention periods remain widespread.
The Data Minimization Alternative
GDPR Article 5(1)(c) requires personal data to be “adequate, relevant, and limited to what is necessary.” In practice, most companies interpret “necessary” as broadly as possible.
What this looks like in practice. Apple offers the most visible counter-example. They process data on-device whenever possible, use differential privacy to learn aggregate trends without individual-level data, and have built features like Apple Intelligence (Genmoji, Image Playground, Writing Tools) using synthetic data generation rather than user behavioral data. Their privacy posture is not charity. It is a product differentiator: in Common Sense Media’s 2021 review of streaming apps, Apple TV+ scored 79% on privacy while the Netflix app scored 46%.
The academic research supports Apple’s approach. The data minimization study I cited earlier found that 33-85% of user history can be removed with minimal quality impact. Netflix’s own 40% unused data figure confirms that even within their current architecture, a massive amount of collection produces no measurable value.
The tension is organizational, not technical. Product teams default to “collect everything because storage is cheap.” And marginal storage costs are cheap. But when you factor in the full lifecycle (ingestion infrastructure, processing pipelines, governance overhead, compliance staff, DSAR processing, deletion engineering, regulatory risk), the total cost of ownership for a data point is far higher than its S3 storage fee.
What This Means for Data Teams
If you are building data infrastructure or setting Data Governance policy, the research points to three actionable conclusions.
First, audit what you actually use. Netflix discovered 40% waste despite having one of the most sophisticated data engineering organizations on the planet. Most companies have never done this audit. Before your next infrastructure budget review, run a usage analysis on your data lake. You will almost certainly find entire pipelines feeding tables that nobody queries.
Second, design for deletion from day one. Netflix invested in a centralized deletion platform after years of ad-hoc approaches across 1,300 datasets. The earlier you build deletion into your architecture, the cheaper it is. Treat data retention policies as architectural requirements, not compliance paperwork.
Third, challenge the “more data is better” assumption. The recommendation research shows clear diminishing returns. If your team is collecting device fingerprints, scroll depth, hover duration, or millisecond-precision interaction logs, ask whether those signals meaningfully improve the product. If nobody can demonstrate the improvement, you are paying infrastructure and compliance costs for data that exists because nobody questioned the default.
The privacy conversation usually focuses on what companies collect and whether users consent. That framing misses the economic reality. Companies are spending billions to collect data, a measurable portion of which they never use, while simultaneously building expensive infrastructure to delete it when regulators or users demand it. The “collect everything” default is not a strategy. It is an organizational habit that persists because the costs are distributed across enough line items that nobody sees the full picture.
Netflix’s numbers tell the story clearly. Forty percent waste. Thirty-three to eighty-five percent of user history removable at 2% quality cost, depending on the recommendation model. $1,524 per deletion request. These are not abstract principles. They are the actual economics of knowing everything about your users.
None of this is an argument for collecting nothing: granular playback signals also feed adaptive bitrate and encoding decisions, quality-of-experience monitoring, fraud and account-sharing detection, and A/B testing infrastructure, so the minimization case applies to what gets retained beyond those operational uses, not to the raw telemetry itself.
The companies that figure out how to collect less and extract more value per data point will have a structural cost advantage. Everyone else will keep paying the bill for data they never needed.
Sources & References
- Netflix Tech Blog - Keystone Real-time Stream Processing Platform(2024)
- Netflix Tech Blog - Navigating the Netflix Data Deluge(2024)
- CloudZero - Inside Netflix's AWS Strategy(2025)
- arXiv - The Value of Personalized Recommendations: Evidence from Netflix(2025)
- arXiv - What Data is Really Necessary? Inference Data Minimization for Recommender Systems(2025)
- InfoQ - Netflix Tackles Data Deletion at Scale(2025)
- DataGrail - The Cost of Data Privacy Continues to Rise(2024)
- AdExchanger - Netflix Doubled Its Ad Revenue(2026)
- IAPP - Apple's New Privacy Push Focuses on Data Minimization(2025)
- Common Sense Media - Privacy Analysis of Top Streaming Apps(2021)
- CNBC - Spotify Will Spend Nearly $450 Million on Google Cloud Over 3 Years(2018)
- CNBC - Google Hit with EUR 150 Million French Fine for Cookie Breaches(2022)
- EDPB - CEF 2025 Report on Right to Erasure(2025)
- FlowingData - Why Netflix's $1M Algorithm Never Went to Production(2012)
- Variety - Netflix Content Spending 2025(2025)
Stay in the loop
Get new articles on data governance, AI, and engineering delivered to your inbox.
No spam. Unsubscribe anytime.