header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

After the company went bankrupt, how it made money by selling employee data

Read this article in 35 Minutes
A Rebirth of 600 Million Messages.
Original Title: "How to Make Money by Selling Employee Data After a Company Goes Bankrupt"
Original Author: Insight from Beating


Emails, chat records, project documents, work orders—previously just digital remnants awaiting cleanup after a company shutdown. Now, they are being reevaluated, packaged, sold, and fed into an AI company's training pipeline.


A mortician preserves a dead person's last shred of dignity. The posthumous dignity of a corporation lies in proving that what they left behind still holds value.


On August 17, at a bankruptcy asset auction, Google bid $10 million to acquire all of Spirit Airlines' corporate data. Competing against them was a company named Mercor, bidding $7.5 million, falling short by $2.5 million.


The auction items were divided into three parts. The first part consisted of approximately 100 million employee emails. The second part included 500 million Microsoft Teams messages. Together, these two parts totaled 600 million messages. The third part comprised calendars, spreadsheets, financial databases, project files, operational records, and a batch of internal software. Passenger profiles and frequent flyer information were excluded.


600 million messages—if a person speaks 100 sentences a day, they would have to speak continuously for over 16,000 years.


Calculating it out, one message was sold for 1.67 cents. Americans refer to a 1-cent coin as a penny, often not bothering to pick it up when dropped on the ground. In other words, one sentence spoken by a Spirit employee in Teams was worth one and a half pennies.


Let's rewind to 1980 when Spirit Airlines was born in Detroit. It originated from a trucking company called Charter One, which later shifted its focus to aviation. In 1992, it rebranded as Spirit, meaning "soul." Over the next 34 years, it set the industry standard for the imitation of the ultra-low-cost carrier model in the U.S.


Its foundation was once strong—205 all-Airbus fleet, approximately 300 flights a day, with a 2024 revenue of around $5 billion. However, in the same year, it suffered a net loss of about $1.2 billion, carrying about $9 billion in debt when filing for bankruptcy.


In a late-night May 2026 announcement, the company declared a halt to its flights. The next day, about 17,000 employees learned from the news that they were now jobless.


The process is still ongoing, with transactions awaiting approval from a bankruptcy judge. Spirit is undergoing a liquidation closure under Chapter 11, without a bankruptcy trustee taking over. The company is still under court supervision and is selling off its remaining assets one by one. Involving sensitive data like employee emails and Teams messages, the information must first be processed by an independent entity to remove personally identifiable information such as names and email addresses. Google selected this entity, and Google is also covering the costs.


Furthermore, scouring through public reports, the bankrupt company selling internal data to an AI company had never been seen before. This deal is likely the first of its kind in history.


The long-standing U.S. tech media Gizmodo titled this deal as follows: "Spirit may be dead, but its data will haunt Google's servers for generations to come."


A New Business Opportunity


There have always been people who deal with the remains of dead companies. Lawyers, liquidators, auction houses—doing this for decades. Aircraft, furniture, trademarks, patents—all that could be sold has long been sold.


What's truly new is that starting this year, even a company's internal employee data has been put on the table.


This emerging business has two main reasons behind it.


First, the number of defunct companies has increased.


In the first quarter of 2024, the failure rate of U.S. startups increased by 58% compared to the previous year, and the number of active venture capital firms decreased by 62% from its peak.


The money hasn't actually decreased; it's just flowing more and more heavily toward AI. In 2024, U.S. AI startups raised a record $97 billion in financing. The capital markets still have money, but they are increasingly unwilling to spend it outside of AI.


As a result, a group of companies that could have continued to survive on financing have started hitting a wall earlier. In August 2024, the fintech company Tally, backed by a16z, announced its closure. Having raised a total of $172 million with a peak valuation of $855 million, reaching Series D, it still couldn't secure the next round of funding.


Second, data has become more valuable.


The consumption of data for large-scale model training has reached an extreme level. The training data for GPT-4 is approximately 130 trillion tokens. For comparison, over more than forty years, Google Books scanned around 40 million books, which is roughly 4 trillion tokens. This means that the amount of data used for one GPT-4 training session is equivalent to over three Google Books datasets.


Epoch AI did the math and estimated that high-quality language data from books, news, and Wikipedia will be depleted around 2026. The amount of high-quality text available on the public internet for training is rapidly diminishing, with synthetic data becoming increasingly prevalent. Consequently, AI companies are forced to look beyond the public internet for new data. The internal company archives of years of emails, chat records, and work documents have now come into view.


As for the AI training dataset market, research agencies predict it could reach $9.7 billion by 2030. Considering all kinds of licensing, the total market size is estimated to reach $67.5 billion.


On one side, more and more companies are shutting down and liquidating, leaving behind a large amount of internally held data that was previously not priced; on the other side, AI companies have an increasing demand for data beyond the public internet. The simultaneous occurrence of these two factors is what has enabled for the first time internal corporate data to be ready for scalable transactions.


What truly adds value to this type of data is the development of Enterprise Agents. Gartner predicts that by 2026, 40% of enterprise applications will embed task-specific AI Agents, whereas the previous year this ratio was less than 5%.


The training material required for Enterprise Agents is not the same as that for regular large models. While public webpages can provide knowledge, language, and finished content, it is difficult to replicate a company's actual workflow. Communication and collaboration involve many real details, such as how a requirement is proposed, how multiple people discuss it, how tasks are assigned, how issues are resolved, and ultimately how delivery is made.


These processes are largely preserved in a company's internal emails, chat records, work orders, and project documents.


When this type of data starts to have clear buyers and use cases, things that were previously deleted directly when the company shut down now have standalone trading value.


Corpse Collectors


Previously carried out by liquidation lawyers, there are now three additional types of individuals.


The first type is called a Dissolution Service Provider, with an example being SimpleClosure.



This company focuses on one thing: helping startup companies die gracefully. In 2023, the company started its journey, raising $1.5 million in a pre-seed round, and in May 2025, it secured a $15 million Series A round led by TTV Capital. Even Carta, which heavily managed equity and corporate affairs for many U.S. startups, halted its own shutdown services and instead invested in SimpleClosure, redirecting this segment of customer demand to them.


By October 2025, SimpleClosure had already handled the funerals of over a thousand companies. Crunchbase gave it a nickname, "A Better Way To Fail," a kinder way of failing. While American entrepreneurs love to say "fail fast," SimpleClosure believes that fast failure is not enough; it must also be done gracefully. Their website even features a pricing calculator where you can input your company's situation to determine how much it would cost for one dignified demise.


A funeral is never in vain. In April 2026, SimpleClosure launched the Asset Hub, specifically to handle intangible assets left behind after a company's closure. In addition to the brand, software, and customer list, for the first time, items such as Slack messages, emails, and Jira tickets, which are internal work data, were explicitly put on the shelf.


This indicates that even before Spirit, the market had already begun to try to price the internal data of dead companies, although at that time it was still startups, small transactions, and private matching.


There is a ready-made case. When the transcription and captioning company cielo24 shut down, it sold Slack messages, internal emails, and Jira tickets accumulated over the past 13 years through SimpleClosure. CEO Shanna Johnson later told Forbes that this data set was eventually sold for hundreds of thousands of dollars.


For a company that had already decided to close its doors, this was originally a batch of data that needed to be cleaned up but ultimately became an asset that could be recovered during liquidation.


For the data sale process, SimpleClosure handed it over to Protege. This is a data exchange market specializing in AI training data licensing, which raised $30 million in January 2026, with lead investment from a16z. The founder is Bobby Samuels. Protege's initial entry point was medical imaging, and within 30 days, they assembled millions of images for pre-training for the buyer.


Now, Protege is starting to apply this data licensing and trading capability to the internal communication data of closing companies. SimpleClosure is responsible for company closure and asset organization, while Protege is responsible for finding buyers, completing data licensing, and transactions.


The second type of participant is bankruptcy courts and liquidation lawyers. For decades, they have inventoried airplanes, tables, chairs, trademarks, and patents. Now, the list is beginning to include emails, Slack messages, and other internal data.


According to U.S. bankruptcy law, this data can be included as intangible assets in the bankruptcy estate and sold under court supervision. Related law firms have also begun to establish dedicated teams to handle data preservation, discovery, and organization in bankruptcy cases. Redgrave LLP is one such restructuring and discovery firm.


The third type is e-discovery service providers responsible for technical execution. Companies like KLDiscovery, Epiq, and Consilio are usually involved in collecting, organizing, hosting, and reviewing enterprise data. The content in emails, Teams, and SharePoint needs to be exported, archived, and organized into data packages that can enter the transaction process.


This industry already has a mature set of pricing models. The EDRM regularly releases pricing surveys, with common billing units including data collection per GB, hosting per GB per month, and document review fees.


Spirit’s 600 million messages eventually turned into an auction item, relying on this kind of foundational work. However, unlike court documents and auction bids, this part rarely appears in public reports. The specific details of who handles it and how it is handled are usually invisible to the outside world.


Autopsy Checklist


For a corporate Agent, what it needs to learn is not only knowledge and standard answers, but also judgment, collaboration, error correction, and execution processes in a real working environment.


The finished product tells the model what it ultimately became, while internal records tell it how it was achieved.


This shift is already reflected in Agent training data. In the past, a single-line of code training sample might only be a few hundred tokens, involving modifying a few lines of code; now, a single Agent training sample often needs to include the entire process of requirements understanding, document locating, code modification, and test validation.


Training data is moving from a single correct answer to a complete task execution record. For the Agent, the end result is certainly important, but the judgment, actions, and feedback left during task completion are more valuable for training.


Compared to a company operating normally, the data of a closed-down company is more likely to enter the transaction process. While the company is still operating, selling internal communications would involve trade secrets, employee privacy, competitive risks, and customer relationships, and legal and management teams are usually very cautious. After entering liquidation, the company's main goal becomes to recover as many remaining assets as possible to secure more value for creditors.


This is also why the same set of internal data, which was difficult to sell when the company was alive, may be revalued as part of the assets in the closure stage.


The value of internal corporate data has long been recognized by the industry. Salesforce has always regarded the enterprise communications accumulated in Slack as important data assets, and Microsoft CEO Satya Nadella has repeatedly emphasized that when companies use AI, the truly valuable part is their proprietary data and work context.


In the past, this data mainly served the companies themselves. Now, as AI companies actively seek enterprise internal data, they are finding clearer external buyers for the first time.


The Art of Pricing


While this business is still in its early stages, some reference prices have emerged in the market.


SimpleClosure and Protege handle data from shut-down startups, with individual transactions typically ranging from $10,000 to $100,000. Mercor has provided quotes to acquired startup employees based on chat logs and emails, with the highest offer reaching $300,000.



Spirit has pushed the price into the tens of millions of dollars. Mercor bid $7.5 million, and Google eventually bid $10 million, competing for around 600 million internal communications and other corporate data. Based solely on these 600 million messages, the average price per message is approximately 1.67 cents.


This unit price is not high. In 2024, Reuters reported that Photobucket negotiated licensing of around 13 billion photos and videos with an AI company, with prices ranging from about 5 cents to $1 per photo and over $1 per video. In the B2B data market, a single contact's information can be sold for anywhere from a few cents to a few dollars depending on completeness and accuracy.


However, these prices are still not enough to form a unified standard. The value of a shut-down company's internal data is currently mainly a matter of negotiation. Factors such as data volume, industry, time span, completeness, uniqueness, and what the buyer intends to do with it will all affect the final price. Spirit's $10 million deal appears more like one of the few publicly visible large-scale examples.


When compared to the established data licensing market, the gap becomes even more apparent. Reddit licensed user posts and comments to Google for about $60 million annually; News Corp's content licensing agreement with OpenAI is worth around $250 million over five years; xAI's collaboration with Telegram amounts to $300 million; Apple's acquisition of Shutterstock image licenses ranges from $25 million to $50 million.


These markets already have stable buyers, licensing methods, and pricing experiences. Transactions of shut-down startups' internal data are just beginning, and there are currently no clear rules on which data is most valuable and whether valuation should be by item, volume, or as a whole.


Looking at Spirit's $10 million deal in the context of its own scale is another story. In 2024, Spirit's annual revenue was close to $5 billion, averaging around $13.7 million per day. The amount Google paid for this data is less than what it typically earns in a day of normal operations.


For a bankruptcy liquidation, this is just a small recovery from the remaining assets; for an AI company, what was acquired is a batch of long-unseen data containing the real-world operational processes of the business.


Cleansing and Handover


After the data transaction, it cannot be directly handed over to the buyer. Before the formal handover, it usually needs to go through several steps such as extraction, de-identification, organization, packaging, court approval, and final delivery.


The first step is extraction. Slack's enterprise data is usually exported as a ZIP file, including JSON files organized by channel, member information, and attachments.


Microsoft 365, on the other hand, can export emails and Teams messages through an e-discovery tool. Spirit involves approximately 600 million internal communications, which is a large amount of data and requires processing in batches in practice. Whether this will be carried out by an e-discovery service provider or the buyer's engineering team is undisclosed in public documents at the moment.


Next is de-identification, which involves removing as much information as possible that can be traced back to specific employees. Common methods include identifying and masking personal information such as names, phone numbers, addresses, or replacing sensitive fields with identifiers that do not directly correspond to individuals. In some statistical and training scenarios, additional random perturbation is used to further reduce the possibility of reidentifying individuals.


However, de-identification does not guarantee complete anonymity. Even if names and emails have been removed, as long as the text retains enough occupational, temporal, locational, or behavioral features, there is still a possibility of reidentifying individuals through other information.


Therefore, in this deal with Spirit, who is responsible for this step is crucial. According to the current transaction arrangement, a third-party independent entity will handle the de-identification process, but the chosen institution is selected by Google, with Google bearing the cost. Google has also committed not to reidentify individuals using this set of data.


Selling data during a bankruptcy by a company is not a new concept. Over the past two decades, from Toysmart, Borders, RadioShack to 23andMe, cases have arisen where customer data was dealt with or sold during bankruptcy proceedings. Such transactions involving consumer privacy usually undergo stricter scrutiny by courts, regulatory bodies, and state governments.


What sets Spirit apart is that this sale is not centered around the passenger list but rather around employee emails, Teams messages, and other internal work data. Existing bankruptcy procedures do not have established rules for handling this type of data as they do for consumer information.


Controversies have already arisen. Spirit's flight attendants' union has raised objections to the transaction, leading the court to postpone approval. A batch of corporate data that was initially intended for sale as bankruptcy assets is now starting to involve employee rights and the boundaries of AI usage.


If the transaction is ultimately approved, the data will only enter the final delivery stage. However, there is little public disclosure of the data between the finalized data package and Google's systems. It cannot currently be confirmed how the data is transmitted, whether further cleansing will occur, and in what form it will ultimately enter product development or model training.


At this point, the process completes the true transformation of data from bankruptcy assets to AI assets.


Buyer


The most explicit buyers of this type of data at present are model companies like Google and training data service providers like Mercor.


Google's official statement regarding this transaction with Spirit is for product enhancement and AI development. For Google, the value of this data lies in documenting the real operational processes of a large enterprise, such as scheduling, collaboration, internal communication, project progress, issue resolution, and management decisions.


These contents are difficult to obtain from public web pages. Especially for enterprise agents, beyond the final result, what is more important is how tasks are actually progressed and completed within a real organization. The internal records left by Spirit happen to contain a large amount of such process data.


Another bidder, Mercor, better illustrates the direction in which this business is heading.


Founded in 2023, Mercor's three founders, Brendan Foody, Adarsh Hiremath, and Surya Midha, have been debate teammates since high school. The company initially focused on AI recruitment, using models to help companies screen and interview candidates. By September 2024, Mercor had evaluated around 300,000 job applicants, achieving a valuation of $2.5 billion.


Subsequently, the company gradually shifted its focus to AI training data. In November 2025, with the three founders each only 22 years old, as the company's valuation rose, they became some of the youngest self-made billionaires globally at the time. In the first half of 2026, Mercor's revenue exceeded $614 million, with about 90% coming from top AI labs like OpenAI. By July, the company was seeking a valuation of around $20 billion and had acquired Deeptune, which specializes in building training environments for AI agents.


Mercor acquires training data mainly through two avenues.


One is by directly purchasing internal corporate records. It has made offers to acquired or closed-down startups to buy employee chat logs and emails, with the highest bid to a single company being around $300,000.



Another is to hire people with actual work experience. TechCrunch reported in October 2025 that the AI lab recruited former employees through Mercor, allowing them to convert their work experience into training tasks and feedback at an hourly rate of about $200. Mercor's CEO has stated that the company pays out over $1.5 million daily to individuals participating in AI training.


On one side, acquiring the work records left by companies, and on the other, buying the work experience held by employees. Mercor's core assets are essentially real-world work processes.


This also explains why it showed up on the Spirit auction block. For Mercor, 6 billion pieces of internal communication serve as another, larger-scale source of training data.


Such business activities also come with risks. In 2025, Scale AI sued Mercor, alleging that former employees took trade secrets with them; subsequently, the company has faced training-related information leaks and partnership suspensions. Data is not only Mercor's business but also its most essential asset to defend.


Finally, there is a numerical contrast. Spirit has been operating under this name for 34 years, while Mercor, participating in the auction of its internal data, had three founders who were only 22 years old at the time.


Broken Bench


The Spirit transaction is not yet complete, the court has not yet signed off, and Google has not obtained that set of data. There are union disputes to follow, hearings to be held, and many procedures to go through.


But this auction has already made something that was rarely discussed before more concrete.



In the past, after a company closed its doors, many things would indeed disappear. Financial statements remain, trademarks remain, patents remain. However, a company's day-to-day experiences usually do not remain. How a department conducts meetings, how a manager makes decisions, how dozens of people coordinate after a delay, why a set of processes ended up being the way they are today – these things are rarely formally documented.


As a company dissolves, people leave, email accounts are deactivated, chat groups are closed, and they disappear along with them.


So, a company has always been a peculiar organization.


It can survive for decades, accumulating the experiences of thousands of people, but what can truly be inherited is often just a very thin slice of it. The next company will have to rehire, make the same mistakes, and relearn many things all over again.


AI might change just that. If emails, meetings, tickets, code changes, and internal discussions can truly be organized into training data, then a company's experiential knowledge, previously only transferable through people, has for the first time found another way to persist.


Where this all ends up is difficult to say. Perhaps in the future, businesses will voluntarily retain this data, maybe employment contracts will be rewritten, perhaps bankruptcy law will add new constraints, maybe companies will even separately value internal working records during financing and mergers.


Spirit raised this issue early.


Bankruptcy has an etymology with a widely circulated explanation. Italian medieval merchants conducted business sitting behind benches. If they couldn't pay their debts, their bench was publicly destroyed. banca rotta, broken bench.


For centuries, when the bench broke, it was all over.


Spirit's bench has already broken. Its 600 million messages are being revalued.


At a penny and a half each.


Original Article Link


Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia

举报 Correction/Report
Choose Library
Add Library
Cancel
Finish
Add Library
Visible to myself only
Public
Save
Correction/Report
Submit