🦖 AIresearchFinder on moderator discussing that's not a rejection of what I said because probability is the mechanism by which it exercises the judgment. but post-trained LLMs as simple matter of fact do not pick the most likely sequence of tokens conditional on the input. RL teaches them to do something else | Moderator