Evals for taste: Hill-climbing a slide-generation agent

Evals for taste: Hill-climbing a slide-generation agent

More

Descriptions:

Built rubric-driven replayable eval system from real user projects giving quality, cost, latency, error, token signals in under 6 hours per model change. Evolved into dev flywheel powered by real user dissatisfaction signals.

1 Item

Channels

1 Item

Companies