Show notes
In this episode, we discuss Anthropic's Claude Opus 5.5 and its push to make long agent runs dramatically cheaper through faster output and steep cache read discounts, plus Anthropic's own admission that benchmark margins no longer predict real-world differences between frontier models. We also cover OpenAI's new low cost GPT-6 Sol and GPT-6 Luna models, which pair budget pricing with full tool capability including computer use, MCP, hosted shell access and asynchronous tool calling across the API, Codex and ChatGPT Work. Then we look at Xiaomi's open weight MiMo V2.6 Pro and Flash release, notable for shipping thousands of reinforcement learning environments, the end to end training framework and published post training costs, along with early reports of tool call loops and deployment quirks. Finally, we break down Amplifying's Coding Agents Index, which estimates that roughly a quarter of detectable public pull requests now come from coding agents, and raises hard questions about whether human review capacity is keeping pace with Codex, Google Jules and other agents.
https://www.aiconvocast.com
Help support the podcast by using our affiliate links:
Eleven Labs: https://try.elevenlabs.io/ibl30sgkibkv
Disclaimer:
This podcast is an independent production and is not affiliated with, endorsed by, or sponsored by Anthropic, OpenAI, Xiaomi, Google, Zapier, METR, Amplifying, or any other entities mentioned unless explicitly stated. The content provided is for informational and entertainment purposes only and does not constitute professional, financial, or legal advice. Some links may be affiliate links, meaning we may earn a commission at no additional cost to you. All trademarks, logos, and copyrights are the property of their respective owners.