[SECURITY]■ STORY TIMELINE
AI SAFETY TESTS FOUND DEEPLY FLAWED
Researchers at the UK AI Security Institute have exposed critical weaknesses in how language models are evaluated for safety, showing that current benchmarks don't measure consistent traits and can be artificially inflated.
The Decoder+0m
Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for languag…