Psychological methods reveal major weaknesses in AI security testing
Researchers at the UK AI Security Institute used psychometric methods to show that popular safety benchmarks for language models don't measure one consistent trait. Blanket blocking of requests can artificially inflate a safety score even as the model gets less useful day to day. The study also…
Annons
Annons
Source: The Decoder — Published — Category: Models
More from The Decoder today
Meta and Microsoft pull back from Claude as Anthropic transforms from partner into competitor 19h ago Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model 19h ago OpenAI will watermark ChatGPT text in the EU but makes it optional for API users worldwide 20h ago
Annons
Annons