:

US DATA FIRMS SELL AI TRAINING SETS TO CHINA

AI DESK2 MIN READ
FRI, AUG 7, 2026

■ AI-SUMMARIZED FROM 1 SOURCE ▸ TIMELINE

US data labeling companies including Surge AI and Mercor are selling AI training datasets to Chinese labs while also supplying OpenAI, Anthropic, and the US government, according to internal documents reviewed by Forbes.

The revelation raises significant national security concerns about the dual distribution of sensitive AI training data. Surge AI and Mercor, prominent Silicon Valley startups, have built their business models around supplying high-quality labeled datasets to leading US artificial intelligence companies and federal agencies. The same datasets are being made available to Chinese AI laboratories. Forbes reporter Anna Tong obtained documents showing the extent of this cross-border data sales practice. The findings suggest that efforts to maintain US technological advantages in AI development may be undermined by these commercial arrangements. Data labeling companies serve a critical function in AI development, hiring workers to annotate images, text, and other content that trains machine learning models. The quality and scope of training datasets directly impact AI system performance. By selling identical or similar datasets to both US and Chinese entities, these firms are potentially accelerating Chinese AI capabilities while maintaining relationships with US defense and intelligence priorities. This practice occurs amid growing US government efforts to restrict Chinese access to advanced semiconductors and AI technology. The Biden administration has imposed export controls on chip manufacturers and pursued policies aimed at limiting China's AI development. However, the data labeling sector has largely escaped regulatory scrutiny. Neither Surge AI nor Mercor immediately responded to requests for comment regarding the documents. The companies may argue that unclassified training data falls outside export control frameworks, though experts question whether this interpretation aligns with national security interests. The disclosure adds to existing tensions over AI development practices. It highlights a gap between government policy on technological competition with China and the commercial incentives driving private companies. Training data represents a foundational resource for AI systems, making its unrestricted distribution to foreign competitors particularly sensitive. The findings suggest potential regulatory blind spots in how the US monitors AI development inputs, even as policymakers focus heavily on controlling advanced chips and AI model exports.

■ SOURCES

Techmeme

■ SUMMARY WRITTEN BY AI FROM THE LINKS ABOVE

■ MORE FROM THE BUSINESS DESK

The European Union is advancing its independent satellite constellation by one year, targeting launch in 2029. The acceleration responds to shifting security threats requiring expanded communication capacity.

JUST NOWIndustry Desk

The U.S. labor market contracted in July, losing 23,000 jobs in a sharp reversal from recent months. The unexpected decline marks a significant shift in employment trends.

1H AGOIndustry Desk

The U.S. government has agreed to pay a German company $1.2 billion to halt its offshore wind projects. The deal resolves a dispute over the ventures.

3H AGOIndustry Desk

Two proposed hyperscale datacenters in Arkansas are sparking rare bipartisan opposition from residents and activists. Critics argue the projects constitute 'redlining' by targeting rural, Black-owned land while straining local water and electricity resources.

3H AGOIndustry Desk

■ SUBSCRIBE TO THE DAILY BRIEF

ONE EMAIL, 5 STORIES, 06:00 UTC. UNSUBSCRIBE ANYTIME.