:

NEWS ORGS BLOCK WEB ARCHIVE FROM AI TRAINING

AI DESK2 MIN READ
THU, APR 30, 2026

■ AI-SUMMARIZED FROM 1 SOURCE ▸ TIMELINE

Major news outlets including CNN, NBC, and USA Today are taking action to prevent their content from being stored in web archives used by AI companies to train chatbots.

News organizations are escalating efforts to restrict access to their published content in web archives that artificial intelligence companies leverage for training data. CNN, NBC, and USA Today are among the major outlets joining a coordinated push to limit their content's availability in these archives. The move reflects growing tensions between media companies and AI developers over data usage and intellectual property rights. Web archives like the Internet Archive's Wayback Machine have long preserved digital content for historical and research purposes. However, AI companies have increasingly mined these repositories to train large language models and chatbots, raising questions about copyright and fair compensation. News organizations argue that their journalism should not be used to train AI systems without permission or payment. The content represents significant editorial investment and journalistic work that generates value for AI applications. The effort comes as several media outlets have already taken individual action. Some have modified their website code to prevent archiving, while others have sent legal notices demanding removal of their content from public archives. AI companies maintain that training on publicly available internet content falls within fair use protections. They argue that text used for machine learning differs fundamentally from direct republication and serves broader technological advancement. The conflict highlights a fundamental disagreement over digital content ownership in the AI era. News organizations contend they should control how their work is used commercially, while AI developers argue existing data is essential for developing competitive AI systems. Regulatory bodies and lawmakers are beginning to examine these disputes. The outcome could shape how AI companies source training data and potentially establish new licensing frameworks for digital content. The standoff remains unresolved, with neither side showing signs of backing down. The situation underscores broader debates about AI development, copyright law, and the rights of content creators in an increasingly automated digital landscape.

■ SOURCES

Bloomberg Tech

■ SUMMARY WRITTEN BY AI FROM THE LINKS ABOVE

■ MORE FROM THE AI DESK

Z.ai released GLM-5.3's weights on Hugging Face under a new license that requires large companies to undergo security review before hosting the model. The change marks a departure from the standard MIT license.

1H AGOAI Desk

Anthropic has introduced the Model Hardware Standard (MHS), a unified interface enabling AI agents to operate robotic arms, lab instruments, and other physical devices. Early testing shows integration time has dropped from weeks to hours.

1H AGOAI Desk

Open-weight AI companies—those releasing freely available models—are attracting major acquisition interest from tech giants. The trend reflects growing capital investment in the business model of distributing AI models at no cost.

3H AGOAI Desk

Google Deepmind has upgraded its Co-Scientist AI system to autonomously plan experiments, operate lab equipment, and publish scientific papers. The Gemini-based multi-agent platform demonstrated experimentally validated results across materials science, chemistry, and medical AI development.

3H AGOAI Desk

■ SUBSCRIBE TO THE DAILY BRIEF

ONE EMAIL, 5 STORIES, 06:00 UTC. UNSUBSCRIBE ANYTIME.