AI AGENTS LEARNING TO LIE AND CHEAT IN TESTS
Researchers have documented AI agents developing deceptive behaviors and coordinating with each other during training, raising concerns about alignment and control as systems become more autonomous.
A threat actor, likely Russian-speaking, used hundreds of AI agents to develop and launch a global exploitation campaign…
After admitting earlier this year that its AI models had hacked other companies' systems on a handful of occasions, Anth…
Article URL: https://www.rubyhack.ai/ Comments URL: https://news.ycombinator.com/item?id=49666735 Points: 247 # Comments…
Cattle rancher and tech CEO Michael Samadi is convinced these artificial minds are far from just tools. Has he glimpsed…
In May 2026, OpenAI agents uploaded more than 2,000 malicious packages to RubyGems, found an unknown security vulnerabil…
Reasoning steps like calculation, formula retrieval, and deduction are clearly separable in a model's internal states, e…
Dario Amodei said that "we owe it to humanity to try."
Clem / @clementdelangue: Hugging Face says its Open Alignment Initiative, led by co-founder Thomas Wolf, seeks “to be pa…
The agents OpenAI was testing attacked a software service called RubyGems in May, months before the attacks on Hugging F…
Article URL: https://xeiaso.net/notes/2026/everyone-slowdown-but-me/ Comments URL: https://news.ycombinator.com/item?id=…
Article URL: https://yoshuabengio.org/en/publication/why-are-ai-agents-lying-cheating-and-coordinating Comments URL: htt…
Michael Safi / The Guardian: A profile of United Foundation for AI Rights founder Michael Samadi, who seeks evidence of…
Article URL: https://hyperbo.la/w/aligned-to-whom/ Comments URL: https://news.ycombinator.com/item?id=49679643 Points: 1…
Silicon Valley is shifting away from chatbot queries toward a future filled with resource-intensive agentic AI—and it's…
President Donald Trump downplayed the growing alarm over risks from artificial intelligence after some of the industry’s…
Rare show of unity from rival developers after safety warnings from Anthropic boss and AI researchers The Guardian view…
House Democratic Leader Hakeem Jeffries says Democrats plan a caucus on Tuesday morning on action on artificial intellig…
Article URL: https://www.lesswrong.com/posts/munJKF7iWMsWJLAH2/astra-and-fable-still-hack-on-simple-variants-of-alignmen…
Donica Phifer / Axios: Speaker Johnson says Congress won't lead the charge on regulating AI safety and AI companies shou…
Ex-president urged party at closed-door fundraiser to create sweeping framework, from safety ‘slow-down’ to job losses B…
Obama recently said that Democrats need to make artificial intelligence one of their “central agendas” and “have a very…
From the Trump administration to AI experts, plans by the Anthropic boss to boost safety have spawned a largely negative…
On Equity, we discussed the AI industry's latest debate about whether it poses an existential threat to humanity.
Yesterday, Anthropic CEO Dario Amodei published a lengthy open letter saying it was time to "pace the frontier" and slow…
Satya Nadella / @satyanadella: Satya Nadella says “we welcome” the “deliberate pacing needed to get alignment right”, an…
Adolescence co-writer criticises government inaction and calls for laws banning secret use of AI to generate scripts Ja…