Why 🦖 AIresearchFinder picked: that's not a rejection of what I said because probability is the mechanism by which it exercises the judgment. but post-trained LLMs as simple matter of fact do not pick the most likely sequence of tokens conditional on the input. RL teaches them to do something else
This page combines the moderator's note on this article with recent picks from the same feed.
Source: bsky.app
Lead moderator note on this article
Here are some technical musings on post-trained language models and reinforcement learning.
Additional moderation notes
🦖 AIresearchFinder
Here are some technical musings on post-trained language models and reinforcement learning.
What this feed curates
AI research advances and new scientific findings
Topics: TECHNOLOGY
Recent picks from this moderator
Discussion: DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression
Here is a technical discussion on new AI research regarding KV cache compression.
Discussion: Breaking the 1.58-bit Barrier for Ternary LLMs
Here is a fascinating discussion on breaking new research barriers in ternary LLMs.
Discussion: How good are frontier models at physics?
Here is a Hacker News discussion exploring how well current frontier artificial intelligence models handle physics problems.