Production Monitoring with Codex: Grafana, Kubernetes, & Security

Production Monitoring with Codex: Grafana, Kubernetes, & Security

More

Summary

This OpenAI walkthrough shows how Codex can help engineers respond to production incidents, starting with the classic 3 a.m. page when checkout errors climb. Instead of manually hunting through Grafana dashboards, logs, deployment history, and code changes, the engineer hands the evidence gathering to Codex using a custom skill.

The first scenario involves a new release that pushes the checkout error rate to around 20 percent. Codex investigates, proposes a fix, and after a human approves it, the error rate returns to zero without rolling back to the previous version. A second example moves to Kubernetes, where a new inventory API release passes CI/CD but is OOM killed and causes cascading failures. A Kubernetes rollout investigator skill identifies the causal chain and produces a patch that restores the inventory API, orders API, and edge gateway.

The presenters also discuss more automated setups, including self-hosted runners that trigger Codex when alerts cross a baseline, and multi-agent configurations where other agents validate fixes. The video closes by connecting production monitoring with Codex security workflows to keep a service continuously available. It is a practical look at human-in-the-loop agentic operations for DevOps and site reliability teams.


📺 Source: OpenAI · Published October 07, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels