What’s Next After RLHF? — Diogo Almeida, TypeSafe AI

What’s Next After RLHF? — Diogo Almeida, TypeSafe AI

More

Descriptions:

Diogo Almeida, co-author of GPT-4, ChatGPT, and InstructGPT — the papers that effectively invented post-training as a discipline — delivers a structured critique at AI Engineer arguing that the field has not yet left the RLHF era, and that Claude Code and similar agentic products are still part of it. His core claim: RLHF optimizes for human preference rather than task correctness, a design choice appropriate for assistance products but structurally incompatible with genuine automation.

Almeida maps the current landscape into two distinct categories: tasks where pleasing the human in the loop is the actual goal (coding copilots, creative tools, customer chat) and tasks where the ideal outcome is removing the human from the loop entirely (background software processes, data pipelines, server jobs). He argues no amount of RLHF tuning bridges that divide because the optimization target is wrong — models trained this way will consistently look right even when wrong, systematically erring toward human preference over calibrated accuracy. He illustrates with a well-known example: sending ChatGPT fart sounds as audio and receiving earnest aesthetic commentary.

What comes next, in Almeida’s framing, is reinforcement learning with verifiable rewards (RLVR) trained against automation metrics that contain no human preference signal at all. He distinguishes this sharply from current RLHF-adjacent approaches and argues the resulting models would behave fundamentally differently — less agreeable, more calibrated, and far better suited to unsupervised long-running tasks.


📺 Source: AI Engineer · Published July 31, 2026
🏷️ Format: Opinion Editorial

1 Item

Channels

1 Item

Companies