Today we are talking with Vignesh Baskaran, the CTO and co-founder of Hexo Labs, about teaching AI agents to improve themselves. Vignesh has been training neural networks since 2012, back when he was still called a data scientist.
Then he became an ML engineer and now an AI engineer, though he says the underlying work has never really changed. It's to figure out how to make a system behave the way you intend it to. He built the litigation search engine that Google itself became a customer of, and now he's chasing something new, agents that rewrite and retrain other agents without a human in the loop.
We dig into Sia, the meta-agent at the center of Hexo's research, and why improving an agent means touching both its harness and its actual model weights, not just one or the other. We talk about proxy evals for when you don't have much to ground truth. The Darwin-Gödel machine and why formal verification is too strict a bar for anything commercial.
How Hexo's work echoes DeepMind's Alpha lineage from AlphaGo to AlphaEvolve, and the spectrum from clearly verifiable to totally subjective tasks? Why VAE evals are quietly wrecking agent quality across the industry, and a great story about an agent that discovered a customer's own eval file was silently corrupted, something buried in hundreds of thousands of traces that no human would have caught.
Visit https://hexolabs.com/
52:49
Today we are talking with Vignesh Baskaran, the CTO and co-founder of Hexo Labs, about teaching AI agents to improve themselves. Vignesh has been training neural networks since 2012, back when he was still called a data scientist.
Then he became an ML engineer and now an AI engineer, though he says the underlying work has never really changed. It's to figure out how to make a system behave the way you intend it to. He built the litigation search engine that Google itself became a customer of, and now he's chasing something new, agents that rewrite and retrain other agents without a human in the loop.
We dig into Sia, the meta-agent at the center of Hexo's research, and why improving an agent means touching both its harness and its actual model weights, not just one or the other. We talk about proxy evals for when you don't have much to ground truth. The Darwin-Gödel machine and why formal verification is too strict a bar for anything commercial.
How Hexo's work echoes DeepMind's Alpha lineage from AlphaGo to AlphaEvolve, and the spectrum from clearly verifiable to totally subjective tasks? Why VAE evals are quietly wrecking agent quality across the industry, and a great story about an agent that discovered a customer's own eval file was silently corrupted, something buried in hundreds of thousands of traces that no human would have caught.
Visit https://hexolabs.com/
如果你正在编写代码或维护系统运行,你大概熟悉这样的场景:深夜被叫醒、追踪奇怪的 bug、处理告警风暴。这很棘手!系统故障会带来经济损失,而且说实话,没人喜欢这种体验。
那么关键问题是:我们能否利用类似 AI(尤其是 AI 代理)来让可靠性工作变得不那么痛苦、更加系统化?这正是我们今天要讨论的话题。我们邀请到了 Temperstack 的联合创始人兼 CEO Amal Kiran。
他们正在构建旨在自动化 SRE 任务的工具,比如自动发现监控缺口、告警、辅助根因分析,甚至利用 AI 生成 Runbook。
如果你想了解如何将 AI 应用于真实的 SRE 问题及其背后的技术,相信你会喜欢本期节目。
51:08
如果你正在编写代码或维护系统运行,你大概熟悉这样的场景:深夜被叫醒、追踪奇怪的 bug、处理告警风暴。这很棘手!系统故障会带来经济损失,而且说实话,没人喜欢这种体验。
那么关键问题是:我们能否利用类似 AI(尤其是 AI 代理)来让可靠性工作变得不那么痛苦、更加系统化?这正是我们今天要讨论的话题。我们邀请到了 Temperstack 的联合创始人兼 CEO Amal Kiran。
他们正在构建旨在自动化 SRE 任务的工具,比如自动发现监控缺口、告警、辅助根因分析,甚至利用 AI 生成 Runbook。
如果你想了解如何将 AI 应用于真实的 SRE 问题及其背后的技术,相信你会喜欢本期节目。