Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

news.sbsnews.sbssiliconcanals+1AI developers including OpenAI and Anthropic are increasingly purchasing internal corporate data — employee chats, emails, video conference recordings, and code change histories — as the supply of publicly available internet text nears exhaustion, according to a report by The Information published on August 13 citing roughly 20 sources including startup founders and data industry insiders.news.sbs
The demand for enterprise communication data has intensified over the past several months, driven by the race to build AI agents capable of performing real-world tasks such as customer service and invoice verification. Bobby Samuels, CEO of data brokerage startup Protege, told The Information that his company's gross transaction volume has surged from $30 million last year to at least $100 million this year.news.sbs
Data buyers are primarily targeting startups facing bankruptcy or acquisition. Warmly, an AI agent startup recently acquired by HubSpot , reportedly received four separate offers of up to $300,000 to purchase its internal meeting minutes and emails after signing its acquisition agreement, but declined all of them. The records of live human interactions — developers solving problems over email, CFOs discussing financial performance, software demonstration videos — are considered essential training material for next-generation AI systems.news.sbs
The buying spree reflects a structural constraint. Epoch AI estimates that the effective stock of quality public human-written text amounts to roughly 300 trillion tokens and could be fully consumed by frontier models between 2026 and 2032. Under more aggressive training assumptions, that timeline could collapse to as soon as 2027. Fresh human text continues to appear online, but its growth rate is far slower than the expansion of AI training datasets.siliconcanals+3
The scramble for private data follows a period of controversy over how AI companies have sourced training material. Anthropic's internal "Project Panama" initiative involved purchasing copyrighted books, cutting off their bindings for scanning, and destroying the physical copies afterward. A federal judge in San Francisco approved a $1.5 billion class-action settlement in July over the program. As The Washington Post reported, an unsealed internal planning document described the effort as an attempt "to destructively scan all the books in the world."reuters+1
As de-identification of corporate datasets emerges as a major challenge, Shub Sinha, CEO of data processing firm Integral, said his company is conducting "a rigorous anonymization process to maintain data value while complying with privacy regulations." The expanding market for private text signals that training costs — once dominated by compute — are becoming an increasingly consequential line item for AI developers.news.sbs