Intelligence + Continual Learning = Expertise — Yu Su, NeoCognition

Intelligence + Continual Learning = Expertise — Yu Su, NeoCognition

More

Descriptions:

Scheduling a meeting is not finding a shared slot on everyone’s calendar. It is a constraint optimization over authority, priority, and urgency, and an expert sees that immediately where a very capable model does not. Yu Su uses examples like that one to separate two things the field keeps collapsing together. Intelligence is reasoning through an unfamiliar problem from the context you were handed, which frontier models keep getting better at and where each episode stands alone. Expertise is accumulated, situated competence, and almost nobody is scaling it.

His account of why coding agents work while everything else stays brittle is a modern Moravec’s paradox. Code is already a language native world, symbolic and structured, with tests standing in for rewards. The rest of digital work is millions of micro worlds, each with its own local physics, far too heterogeneous for one static model to compress. The slide he calls the most important plots raw intelligence against expertise and finds them roughly orthogonal: scale intelligence alone and you get what he calls the world’s smartest novice, brilliant at whatever is put in front of it and accumulating nothing between problems. Intelligence expands the search, spinning up a hundred parallel attempts. Expertise compresses it, because the shortcuts are already learned. The provocation he leaves is unbounded expertise from bounded intelligence: if continual learning gets good enough past some threshold of raw capability, the thing worth scaling stops being the model.

Speaker info:
– https://x.com/ysu_nlp
– https://www.linkedin.com/in/ysu1989/
– https://ysu1989.github.io/

Timestamps:
0:00 – Why coding agents work and little else does
1:30 – Agents before language models
2:06 – What multimodal language agents changed
3:26 – Why code was the ideal first market
4:06 – Leaving the privileged world of code
4:46 – A modern Moravec’s paradox
5:25 – Millions of micro worlds, each with local physics
6:44 – Defining intelligence
7:20 – Defining expertise
8:00 – Experts see differently, not just more
8:39 – Conditionality, judgment, and knowing when to stop
9:53 – Expanding the search against compressing it
10:32 – Continual learning as the bridge
11:08 – Four parts of a working definition
13:10 – The world’s smartest novice
14:32 – Unbounded expertise from bounded intelligence
15:51 – Reliability against plasticity
16:32 – Parametric and nonparametric learning together
17:08 – Specialization as the next data opportunity
17:53 – Making expertise abundant

1 Item

Channels

1 Item

Companies