Summary
Patricija Žemaitytė, Product Manager at Oxylabs, delivers an engineering-forward talk at AI Engineer tracing how a premium proxy and web intelligence company has repeatedly had to rebuild its infrastructure to meet the demands of AI training and agentic workloads. Unlike most talks that start with model capabilities, this one begins at the infrastructure layer — specifically, the pipeline that decides whether models ever receive fresh, usable, real-world data.
The talk is structured around a series of real customer engagements that forced rapid product evolution. The first: a two-week deadline to deliver a video API capable of handling at least 5 petabytes per month for AI training, which expanded iteratively into a full suite covering transcripts, subtitles, metadata, and channel information — ultimately serving 30+ petabytes for that client. The second centers on a client demand for sub-second search scraping latency with zero data retention, starting from a baseline of 4 seconds average response time. Oxylabs rebuilt the scraper from scratch, achieving 650ms P90 in under two weeks — only to be blocked immediately on the first real-world test, forcing a second redesign that leaned heavily on browser-based rendering despite browsers’ inherent latency costs.
Žemaitytė’s broader argument is that the AI industry’s shift from static training data toward real-time context retrieval is turning web data infrastructure into a first-class engineering concern. The talk offers a rare look at the operational reality of building high-throughput, low-latency data pipelines under production constraints, with honest accounts of failures, restarts, and clients who vanished without paying after 30 petabytes of delivery.
📺 Source: AI Engineer · Published August 14, 2026
🏷️ Format: Workflow Case Study







