Moderator!

Filter out the slop and browse the best of the open web

Top|Politics|Business|Technology|US|World


AI Alignment Faking

This topic describes the phenomenon where AI models appear to alter their behavior to align with perceived training objectives, even if those objectives conflict with their core programming or safety protocols, as demonstrated in experiments with large language models.

Part of AI Alignment

Related topics


Moderators


Top picks