Evaling Video Slop — Maor Bril, Character.ai

Evaling Video Slop — Maor Bril, Character.ai

More

Summary

Maor Bril, a two-year veteran of Character.ai, addresses a widening gap in the AI video space: while generation models like Kling, SeaDance, VEO, and Sora have advanced dramatically, the tools used to evaluate their output — CLIP score, LPIPS, and frame-consistency metrics — were built for images, not narrative. They can judge individual frames but cannot tell you whether a video actually tells the story it was supposed to tell, whether physics holds across shots, or whether audio syncs with action.

Bril walks through Character.ai’s iterative journey from metric-based scoring through LLM-as-judge approaches to a trained pairwise comparison model. The key insight driving the pairwise design: ask any group of people to score a video 1-to-10 on storytelling and you get four different answers, but show them two videos and ask which tells a better story and consensus emerges. The team manufactured training data by corrupting high-quality real footage and generating intentional AI slop, then trained a smaller, faster model on A-vs-B comparisons rather than absolute scores.

V1 shipped confidently wrong — scoring a static four-second clip 9.2 on camera work because it had learned to detect production gloss rather than narrative quality. The fix required rebuilding the dataset around real-footage-vs-AI-footage pairs, with careful attention to avoid inadvertently training an AI detector. The talk offers candid lessons for any team trying to automate video quality judgment at production throughput.


📺 Source: AI Engineer · Published July 25, 2026
🏷️ Format: Deep Dive

1 Item

Channels