Strip an employee’s name from a decade of HR threads and payroll records, and apparently that data is worth real money. Google just paid $10 million for exactly that kind of corporate ghost — the entire digital footprint of bankrupt Spirit Airlines, including roughly 100 million internal emails and 500 million Microsoft Teams items, acquired through a competitive bankruptcy auction and destined to train Google’s AI models. This isn’t a data breach story. It’s something newer and stranger.
The Data Google Actually Bought
The scope of what Google purchased goes well beyond a simple email archive.
- Approximately 100 million company emails across around 80,000 accounts
- Around 500 million Microsoft Teams chats and related items
- Roughly 17 million OneDrive files and SharePoint items
- Payroll and crew scheduling records, customer service call recordings, internal software code, financial audits, and fraud reports
- Explicitly excluded: passenger PII, frequent flyer and loyalty program data
Google’s spokesperson framed the purchase as buying “part of an enterprise dataset” that “can be helpful in improving our products and AI models.” Translation: feeding a large language model real corporate workflows — airline-industry jargon, org-wide communication patterns, operational complexity — the kind of domain-specific texture that makes AI tools genuinely useful in enterprise settings like Gmail, Docs, and Workspace. A third party scrubs the data before Google receives it, per PJT Partners VP Dylan Friesner’s court filing. Clean hands, officially.
“We will not receive any personal information from this dataset. Any data we receive will be rigorously scrubbed of any personally identifiable information by a third party before receipt.” — Google spokesperson, via Yahoo Finance
The auction played out like a going-out-of-business sale where even the filing cabinets had bidders. Google opened at $5 million. AI hiring startup Mercor countered at $7.5 million, reportedly willing to handle deidentification themselves. Google raised to $10 million and won — and that difference in who controls the scrubbing process likely mattered to the bankruptcy court. The result: $10 million is now the public benchmark for a large enterprise data corpus.
When Deidentified Doesn’t Mean Risk-Free
Legal cleanliness and practical privacy risk are not always the same thing.
Spirit’s flight attendants’ union has filed objections, per Bloomberg Law reporting, and the concern is legitimate. Decades of HR records, payroll histories, and internal communications — even stripped of names — can carry reidentification risk when processed by sophisticated AI. Think of it like how Netflix’s supposedly anonymous viewing data famously wasn’t: enough behavioral detail makes deidentification fragile. Whether workers should have any say in their personal histories becoming machine learning material is a question the courts haven’t squarely answered. Judge Sean H. Lane in the Southern District of New York still needs to approve the deal.
The precedent being set here matters more than the $10 million.
More bankrupt companies’ archives will likely be packaged and marketed as deidentified AI training assets. Regulators haven’t established clear guardrails yet. Workers have no obvious recourse once the company folds. Your Slack messages, your Teams threads, your work emails — they’re not just records. In the right bankruptcy filing, they’re inventory.






























