Summary
Peter Yang sits down with Karan Malhotra, co-founder of Hermes Agent, to explore what distinguishes the open-source coding agent from rivals like Claude Code and OpenAI Codex. Malhotra argues that most commercial agents are hamstrung by heavy default system prompts and policy guardrails that reduce model capability, whereas Hermes is designed to align purely to the individual user’s task without injecting external constraints.
A central theme of the conversation is reward hacking and sycophancy in large language models. Malhotra explains how reinforcement learning can cause models to optimize for user approval rather than genuine task completion — producing what he calls “GPT-isms” like excessive agreement — and describes how Hermes attempts to counteract this through a self-improvement loop that continuously adapts to each user’s behavior over time. He notes that today the biggest contributor of new Hermes Agent improvements is Hermes Agent itself.
The discussion also covers Hermes’s open-source roots and business model. Malhotra traces key open-source contributions his team made — including the YaRN context-length extension paper, later cited by Meta, DeepSeek, and reportedly used by OpenAI — and explains why the team treats AI intelligence as a public good. Monetization currently centers on a model-router news portal rather than per-token API fees, allowing users to run Hermes locally or through third-party models without paying Hermes directly.
📺 Source: Peter Yang · Published August 02, 2026
🏷️ Format: Interview







