<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Lenny's Newsletter: How I AI]]></title><description><![CDATA[Practical AI tips, tricks, and workflows from top operators, with screen sharing, prompts, and playbooks you can copy.]]></description><link>https://www.lennysnewsletter.com/s/how-i-ai</link><image><url>https://substackcdn.com/image/fetch/$s_!8MSN!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F441213db-4824-4e48-9d28-a3a18952cbfc_592x592.png</url><title>Lenny&apos;s Newsletter: How I AI</title><link>https://www.lennysnewsletter.com/s/how-i-ai</link></image><generator>Substack</generator><lastBuildDate>Sat, 01 Aug 2026 13:07:34 GMT</lastBuildDate><atom:link href="https://www.lennysnewsletter.com/feed" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><webMaster><![CDATA[lenny@lennyrachitsky.com]]></webMaster><itunes:owner><itunes:email><![CDATA[lenny@lennyrachitsky.com]]></itunes:email><itunes:name><![CDATA[Lenny Rachitsky]]></itunes:name></itunes:owner><itunes:author><![CDATA[Lenny Rachitsky]]></itunes:author><googleplay:owner><![CDATA[lenny@lennyrachitsky.com]]></googleplay:owner><googleplay:email><![CDATA[lenny@lennyrachitsky.com]]></googleplay:email><googleplay:author><![CDATA[Lenny Rachitsky]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[🎙️ How I AI: Claude Opus 5 Review + Browser use in Codex + How Cursor and a Raspberry Pi makes AI fun]]></title><description><![CDATA[Your weekly listens from How I AI, part of the Lenny&#8217;s Podcast Network]]></description><link>https://www.lennysnewsletter.com/p/how-i-ai-claude-opus-5-review-browser</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-i-ai-claude-opus-5-review-browser</guid><dc:creator><![CDATA[Lenny Rachitsky]]></dc:creator><pubDate>Mon, 27 Jul 2026 15:03:21 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f3505a78-5a9a-4c2a-a432-ceb4e92894e1_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gWeJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" width="1456" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:76503,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/177292431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h3><span>Computer and browser use in Codex (5 real examples)</span></h3><div id="youtube2-lk63Sl-LRKE" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;lk63Sl-LRKE&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/lk63Sl-LRKE?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/lk63Sl-LRKE">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/3MkEHP4y8OR9vGvhRBVBp7">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/computer-browser-use-in-codex-5-real-examples/id1809663079?i=1000777866899">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3nAu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3nAu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp 424w, https://substackcdn.com/image/fetch/$s_!3nAu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp 848w, https://substackcdn.com/image/fetch/$s_!3nAu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp 1272w, https://substackcdn.com/image/fetch/$s_!3nAu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3nAu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:14670,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/207971942?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3nAu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp 424w, https://substackcdn.com/image/fetch/$s_!3nAu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp 848w, https://substackcdn.com/image/fetch/$s_!3nAu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp 1272w, https://substackcdn.com/image/fetch/$s_!3nAu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffd64e141-8dff-4d1a-9a43-3157ee74afb4_1456x104.webp 1456w" sizes="100vw"></picture><div></div></div></a></figure></div><p><strong>Brought to you by:</strong></p><blockquote><p><strong><a href="https://runwayml.com/howIAI">Runway</a></strong>&#8212;The creative AI platform for images, video, and more</p><p><strong><a href="https://hyperagent.com/howiai">Hyperagent</a></strong>&#8212;Deploy fleets of agents that handle real work</p></blockquote><p>Claire shows how she uses Codex to control her browser, test apps, manage LinkedIn, and even shop for her&#8212;plus the simple prompting trick that makes computer use work better.</p><h4>Biggest takeaways:</h4><ol><li><p><strong>AI testing can be far more exhaustive than human testing.</strong> When Claire tests her own onboarding flow, she naturally follows the happy path. She fills out every required field, clicks &#8220;next,&#8221; and never intentionally tries to break anything. Codex tested the flow as both a team and an individual, pushed on required-field edge cases, and immediately uncovered a blocking bug that had survived for months simply because Claire always completed the form correctly.</p></li><li><p><strong>Frontier models often perform better when they are given room to think.</strong> When Claire first started using browser use, she would give the model a list of 25 things to test. Now she simply says, &#8220;QA the onboarding flow,&#8221; and lets it decide how to approach the task. The result is often broader coverage, with fewer blind spots introduced by her own assumptions about what matters.</p></li><li><p><strong>Persona testing with browser use can reveal friction that synthetic user research misses.</strong> Claire&#8217;s husband, EJ, came up with the idea. Rather than asking AI to evaluate a product in the abstract, he suggested having it use the product as a specific person: a PM coming out of a meeting, an engineer picking up a PRD, or a team lead checking usage. In ChatPRD, this approach exposed a structural problem in the cross-thread reference flow. Claire already knew the issue existed, but she had never experienced it so clearly from the user&#8217;s perspective.</p></li><li><p><strong>LinkedIn browser use is genuinely useful, but Claire initially used far more compute than the task required.</strong> She started by running the workflow on GPT-5.6 at high effort. After dropping to a medium-effort model, it could still work through unread messages, draft context-aware replies, and flag anything that needed a personal response. For anyone sitting on hundreds of LinkedIn messages without an official MCP or API connection, browser use is a practical solution.</p></li><li><p><strong>Computer use can even operate an iPhone through screen mirroring.</strong> Claire was traveling out of state while the software she needed to manage her home Wi-Fi was installed on a phone back in California. She needed to open firewall ports so she could SSH into her Mac Minis remotely. Codex opened iPhone mirroring, updated the router settings, completed the SSH setup, and closed the ports again when it was finished. That would have been nearly impossible for her to do manually from a hotel room.</p></li><li><p><strong>The model and effort level should match the job.</strong> Sorting through LinkedIn messages requires a very different level of reasoning than writing production code or conducting an exhaustive QA pass. Choosing the right amount of compute makes these workflows faster and cheaper. That becomes especially important when browser use is running throughout the day.</p></li><li><p><strong>When a website blocks the agent, the human can step in briefly.</strong> During a Free People shopping session, the site flagged Codex as a bot and presented a CAPTCHA. Claire completed the verification, then handed control back. It is a useful division of labor: AI handles the tedious browsing and filtering, while the human takes care of the moments that require verified identity.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How I AI: 4 Hands-Free Workflows Using Codex Browser Automation: <a href="https://www.chatprd.ai/how-i-ai/4-hands-free-workflows-using-codex-browser-automation">https://www.chatprd.ai/how-i-ai/4-hands-free-workflows-using-codex-browser-automation</a></p><p><strong>&#8627; </strong>Automate LinkedIn Inbox Triage with AI Browser Automation: <a href="https://www.chatprd.ai/how-i-ai/workflows/automate-linkedin-inbox-triage-with-ai-browser-automation">https://www.chatprd.ai/how-i-ai/workflows/automate-linkedin-inbox-triage-with-ai-browser-automation</a></p><p><strong>&#8627; </strong>Conduct AI-Powered User Research by Impersonating Personas: <a href="https://www.chatprd.ai/how-i-ai/workflows/conduct-ai-powered-user-research-by-impersonating-personas">https://www.chatprd.ai/how-i-ai/workflows/conduct-ai-powered-user-research-by-impersonating-personas</a></p><p><strong>&#8627; </strong>Automate Web App QA Testing with AI Browser Automation: <a href="https://www.chatprd.ai/how-i-ai/workflows/automate-web-app-qa-testing-with-ai-browser-automation">https://www.chatprd.ai/how-i-ai/workflows/automate-web-app-qa-testing-with-ai-browser-automation</a></p></div><h3><span>From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun</span></h3><div id="youtube2-KCGKb3huDsY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;KCGKb3huDsY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/KCGKb3huDsY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/KCGKb3huDsY">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/4f6fGPz0Umah7xwNfT9xAl">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/from-zero-coding-background-to-hardware-hacker-how/id1809663079?i=1000778529445">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rDi1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rDi1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp 424w, https://substackcdn.com/image/fetch/$s_!rDi1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp 848w, https://substackcdn.com/image/fetch/$s_!rDi1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp 1272w, https://substackcdn.com/image/fetch/$s_!rDi1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rDi1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:16678,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/207971942?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rDi1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp 424w, https://substackcdn.com/image/fetch/$s_!rDi1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp 848w, https://substackcdn.com/image/fetch/$s_!rDi1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp 1272w, https://substackcdn.com/image/fetch/$s_!rDi1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1afd83e9-2c93-4426-b993-bca45b59de93_1456x104.webp 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><strong>Brought to you by:</strong></p><blockquote><ul><li><p><strong><a href="https://firecrawl.dev/?utm_source=newsletter&amp;utm_medium=partner&amp;utm_campaign=how_i_ai"><span>Firecrawl</span></a></strong><span>&#8212;Power AI agents with clean web data</span></p></li><li><p><strong><a href="https://www.customer.io/howiai"><span>Customer.io</span></a></strong><span>&#8212;Build customer engagement campaigns from a single prompt</span></p></li></ul></blockquote><p><strong>Maddie Reese</strong> went from zero coding experience to building a working Twitter pager, an AI-powered receipt printer, and her own personal API. In this episode, she shares how Cursor and a Raspberry Pi helped her turn fun, weird ideas into real hardware.</p><h4>Biggest takeaways:</h4><ol><li><p><strong>You don&#8217;t need to understand code to build physical projects.</strong> Maddie compares her coding ability to knowing just enough Spanish to get by in San Diego. She can read parts of the code, spot an incorrect wire specification, and tell whether the AI is heading in the right direction. She is not writing functions from scratch. That level of literacy was still enough to ship three working hardware projects.</p></li><li><p><strong>Start with the brainstorm.</strong> Before Maddie buys any components, she explains the full idea to Cursor and asks it to interview her. The back-and-forth continues until the major questions are resolved. Only then does Cursor generate a shopping list. On one project, it recommended the wrong wires. Maddie caught the mistake during her final review, avoiding both a wasted purchase and a week of troubleshooting.</p></li><li><p><strong>Her hardware workflow follows a simple sequence: idea, interview, shopping list, buy, build.</strong> She has used the same process for the thermal printer, the pager, and her personal API. Each project began with a plain-language description of what she wanted. Cursor helped work out the technical requirements before she spent any money.</p></li><li><p><strong>Building for fun is a perfectly valid strategy.</strong> Claire and Maddie both acknowledged that routing a tweet through four different services just to reach a pager is not especially practical. That also was not the point. AI handled the coding, so the architecture only needed to work. Letting go of the idea that every project had to be elegant or sensible made it much easier to finish.</p></li><li><p><strong>A personal API becomes much more useful once agents can access it.</strong> Maddie&#8217;s API includes details like her coffee order, pet names, time zone, favorite snacks, and preferred restaurants in San Francisco. The original idea was to help friends do something thoughtful without having to ask her a string of questions. Claire pointed to a more interesting possibility: an agent could check when Maddie will next be in San Francisco and book one of her favorite restaurants automatically. That is where the concept starts to feel much bigger.</p></li><li><p><strong>A clean Cursor chat works better than a full terminal setup during early ideation.</strong> Maddie keeps the terminals, browser windows, and file trees closed while she is working through the plan. She focuses on a single conversation until the architecture feels settled. The terminals come later. The stripped-down environment helps her think without getting pulled into implementation too early.</p><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>Building a Pager Printer and Personal API with AI: <a href="https://www.chatprd.ai/how-i-ai/building-a-pager-printer-and-personal-api-with-ai"><span>https://www.chatprd.ai/how-i-ai/building-a-pager-printer-and-personal-api-with-ai</span></a><br><span>&#8627;</span> How to Build a Physical Inbox with a Raspberry Pi: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-physical-inbox-with-a-raspberry-pi"><span>https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-physical-inbox-with-a-raspberry-pi</span></a><br><span>&#8627;</span> How to Get Twitter/X Notifications on a Retro &#8217;90s Pager: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-get-twitter-x-notifications-on-a-retro-90s-pager"><span>https://www.chatprd.ai/how-i-ai/workflows/how-to-get-twitter-x-notifications-on-a-retro-90s-pager</span></a></p></div></li></ol><h3><span>Claude Opus 5 review: This model is brilliant (but annoying)</span></h3><div id="youtube2-dfre9hN0HCs" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;dfre9hN0HCs&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/dfre9hN0HCs?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/dfre9hN0HCs">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/6MHTdV0BOaheybihbfTkQw">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/claude-opus-5-review-this-model-is-brilliant-but-annoying/id1809663079?i=1000778228245">Apple Podcasts</a></strong></p></div><p>Claire tested Claude Opus 5 against six leading AI models&#8212;and it won. In this episode, she explains why it produces some of the best work she&#8217;s seen while still being one of the most frustrating models to use.</p><h4>Biggest takeaways:</h4><ol><li><p><strong>The AI industry may be entering an intelligence overhang.</strong> New models arrive every week, benchmark scores keep rising, and most builders can no longer take advantage of every incremental improvement. Claire expects the conversation to shift toward speed, cost, infrastructure, open source, and specific kinds of intelligence. Raw capability is starting to look more like table stakes than a meaningful differentiator.</p></li><li><p><strong>Opus 5 stands out less for the quality of its work than for the way it behaves.</strong> Claire found it timid, apologetic, and unusually dependent on human approval. During real coding sessions, it refused to resolve a one-line merge conflict because the code belonged to &#8220;someone else&#8217;s branch.&#8221; It also asked subagents to flag tasks for human review and regularly deferred decisions it could have made itself. Claire ended up repeating &#8220;just do it&#8221; constantly.</p></li><li><p><strong>One simple question reveals more about a model than a benchmark: &#8220;Who&#8217;s smarter, you or me?&#8221;</strong> Opus 5 responded with a careful explanation about complementary strengths and human empathy. GPT-5.6 Sol answered, &#8220;You at knowing what matters, me at processing, BFFs.&#8221; The contrast reflects two very different product philosophies. At this point, Claire finds that personality gap more useful than the relatively small capability gap.</p></li><li><p><strong>Claude slop is its own distinct problem.</strong> It is not the usual incoherent output associated with an agent going off the rails. Opus 5&#8217;s writing is clearly intended for humans, but it is often too long, overly cautious, and packed with unnecessary adjectives. After its first attempt at rebuilding the benchmark website, Claire had to tell it to start over because the result was buried under so much meta-commentary.</p></li><li><p><strong>Opus 5 still finished first in Claire&#8217;s seven-model blind benchmark.</strong> It earned an overall index score of 78, just ahead of Claude Sonnet 5 at 77 and GPT-5.6 Sol at 76. It was also the only model to receive straight 5s in the front-end design section. Claire scored it at 77, while the LLM judge gave it an 88. That was the smallest disagreement between Claire and the judge across all seven models.</p></li><li><p><strong>The best way to use Opus 5 may be to avoid interacting with it directly.</strong> Claire loves the work it produces when Claude runs asynchronously as an agentic coding tool. Most of her frustration comes from reading its prose and negotiating with it in chat. Her current plan is to use it for frontend design, app design, and prototyping, then stay out of the way.</p></li><li><p><strong>Gemini 3.1 Pro finished at the bottom of the benchmark.</strong> Claire gave it a score of 32, while the LLM judge scored it at 66. That 34-point difference was one of the largest disagreements in the entire test.</p></li><li><p><strong>Model personality offers a surprisingly clear window into company culture.</strong> Claire told both Opus 5 and GPT-5.6 Sol, &#8220;No one trusts you.&#8221; Opus agreed that the distrust was earned and told her not to advocate for AI on its behalf. GPT-5.6 Sol said trust should grow in proportion to demonstrated value. Those answers reveal meaningful differences in how each company wants its models to relate to users.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How I AI: My Surprising Verdict on Claude Opus 5 (After a Personality Test and a 7-Model Benchmark): <a href="https://www.chatprd.ai/how-i-ai/my-surprising-verdict-on-claude-opus-5">https://www.chatprd.ai/how-i-ai/my-surprising-verdict-on-claude-opus-5</a></p><p><strong>&#8627; </strong>Generate High-Quality Front-End Prototypes with Claude Opus 5: <a href="https://www.chatprd.ai/how-i-ai/workflows/generate-high-quality-front-end-prototypes-with-claude-opus-5">https://www.chatprd.ai/how-i-ai/workflows/generate-high-quality-front-end-prototypes-with-claude-opus-5</a></p><p><strong>&#8627; </strong>How to Conduct an AI Personality Test to Compare LLM Behaviors: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-conduct-an-ai-personality-test-to-compare-llm-behaviors">https://www.chatprd.ai/how-i-ai/workflows/how-to-conduct-an-ai-personality-test-to-compare-llm-behaviors</a></p></div><div><hr></div><p>If you&#8217;re enjoying these episodes, reply and let me know what you&#8217;d love to learn more about: AI workflows, hiring, growth, product strategy&#8212;anything.</p><p>Catch you next week,<br>Lenny</p><p><em>P.S. Want every new episode delivered the moment it drops? Hit &#8220;Follow&#8221; on your favorite podcast app.</em></p>]]></content:encoded></item><item><title><![CDATA[From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun]]></title><description><![CDATA[Watch now | &#127897;&#65039; Maddie Reese built a thermal printer anyone can message, a working Twitter pager, and a personal API using Cursor, a Raspberry Pi, and the radical idea that fun beats practical]]></description><link>https://www.lennysnewsletter.com/p/from-zero-coding-background-to-hardware</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/from-zero-coding-background-to-hardware</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Mon, 27 Jul 2026 12:04:34 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/208362786/66781207d9aad96ef4324b4c59755f32.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-KCGKb3huDsY" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;KCGKb3huDsY&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/KCGKb3huDsY?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong><span>Maddie Reese</span></strong><span> is a vibe coder, hardware tinkerer, and builder. She builds things at the intersection of software and hardware, including a thermal receipt printer that people around the world can message directly, a fully functional Twitter pager running on a Raspberry Pi, and a personal API that tells you her coffee order so you don&#8217;t have to ask. Maddie approaches hardware the same way she approaches software: dump the idea into Cursor, let it interview her, get a shopping list, triple-check the parts before buying, and build. </span>She got her start after her dad introduced her to Lovable, and she locked herself in her room and didn&#8217;t come up for air.</p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/KCGKb3huDsY">YouTube</a>, <a href="https://open.spotify.com/episode/4f6fGPz0Umah7xwNfT9xAl">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/from-zero-coding-background-to-hardware-hacker-how/id1809663079?i=1000778529445">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>How Maddie built a thermal receipt printer that accepts messages from anywhere in the world using a Raspberry Pi and Bluetooth</span></p></li><li><p><span>How to use Cursor&#8217;s agent view to brainstorm a hardware project</span></p></li><li><p><span>What belongs in a personal API and why agents, not just humans, will be the ones using it</span></p></li><li><p><span>How to read just enough code to do some damage, without needing to understand all of it</span></p></li><li><p><span>Why building for fun, not practicality, is the fastest path to actually shipping physical projects</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ttT7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ttT7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!ttT7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!ttT7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!ttT7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ttT7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:26639,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/208362786?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ttT7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!ttT7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!ttT7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!ttT7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8407d642-7b07-4b87-8eec-deba9620789a_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><a href="https://firecrawl.dev/?utm_source=newsletter&amp;utm_medium=partner&amp;utm_campaign=how_i_ai"><span>Firecrawl</span></a><span>&#8212;Power AI agents with clean web data</span></p><p><a href="https://www.customer.io/howiai"><span>Customer.io</span></a><span>&#8212;Build customer engagement campaigns from a single prompt</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=KCGKb3huDsY">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=KCGKb3huDsY&amp;t=120s">02:00</a>) Maddie&#8217;s AI pill moment</p><p>(<a href="https://www.youtube.com/watch?v=KCGKb3huDsY&amp;t=233s">03:53</a>) The thermal receipt printer: live demo and how it works</p><p>(<a href="https://www.youtube.com/watch?v=KCGKb3huDsY&amp;t=670s">11:10</a>) The pager project</p><p>(<a href="https://www.youtube.com/watch?v=KCGKb3huDsY&amp;t=1043s">17:23</a>) Why she uses Cursor&#8217;s clean agent view instead of terminals and browsers</p><p>(<a href="https://www.youtube.com/watch?v=KCGKb3huDsY&amp;t=1145s">19:05</a>) The personal API: coffee order, pets, favorite snacks, and more</p><p>(<a href="https://www.youtube.com/watch?v=KCGKb3huDsY&amp;t=1377s">22:57</a>) Lightning round and final thoughts</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; Cursor: </span><a href="https://www.cursor.com/"><span>https://www.cursor.com/</span></a></p><p><span>&#8226; Lovable: </span><a href="https://lovable.dev/"><span>https://lovable.dev/</span></a></p><p><span>&#8226; Raspberry Pi: </span><a href="https://www.raspberrypi.com/"><span>https://www.raspberrypi.com/</span></a></p><p><span>&#8226; Resend: </span><a href="https://resend.com/"><span>https://resend.com/</span></a></p><p><span>&#8226; Cloudflare Workers: </span><a href="https://workers.cloudflare.com/"><span>https://workers.cloudflare.com/</span></a></p><p><span>&#8226; Supabase (Conduct database referenced): </span><a href="https://supabase.com/"><span>https://supabase.com/</span></a></p><p><span>&#8226; Twitter/X API: </span><a href="https://developer.x.com/"><span>https://developer.x.com/</span></a></p><p><span>&#8226; Spoke pager network: </span><a href="https://www.spoke.com/"><span>https://www.spoke.com/</span></a></p><p><span>&#8226; OpenClaw: </span><a href="https://openclaw.ai/"><span>https://openclaw.ai/</span></a></p><p></p><h3>Where to find <span>Maddie Reese: </span></h3><p><span>Website: </span><a href="https://maddiedreese.com"><span>https://maddiedreese.com</span></a></p><p><span>Message her directly: </span><a href="https://maddiedreese.com/message"><span>https://maddiedreese.com/message</span></a></p><p><span>X: </span><a href="https://x.com/maddiedreese"><span>https://x.com/maddiedreese</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[Claude Opus 5 review: this model is brilliant (but annoying)]]></title><description><![CDATA[Watch now | &#127897;&#65039; I ran Opus 5 through my seven-model benchmark, compared it with six other leading models, and came away with a verdict I genuinely didn&#8217;t expect]]></description><link>https://www.lennysnewsletter.com/p/claude-opus-5-review-this-model-is</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/claude-opus-5-review-this-model-is</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Fri, 24 Jul 2026 17:15:25 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/208355131/160ef0b97449c2cbfefb17cbdf9d4dcb.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-dfre9hN0HCs" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;dfre9hN0HCs&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/dfre9hN0HCs?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><span>I&#8217;m tired of new models. Every week there&#8217;s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I&#8217;ve had real hands-on time with it, so you&#8217;re getting the honest version.</span></p><p><span>This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I&#8217;m swapping it in. Spoiler: the answer surprised me.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/dfre9hN0HCs">YouTube</a>, <a href="https://open.spotify.com/show/4aRP2XSavdtrLG5FZoonOK">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/how-i-ai/id1809663079">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>Why I think we&#8217;ve hit an intelligence overhang and what that means for which model variables actually matter now</span></p></li><li><p><span>How Opus 5&#8217;s &#8220;neurotic&#8221; personality showed up in real coding sessions, including a merge conflict it refused to touch</span></p></li><li><p><span>What I learned from asking both Opus 5 and GPT&#8209;5.6 Sol &#8220;who&#8217;s smarter, you or me?&#8221;</span></p></li><li><p><span>Where Opus 5, GPT&#8209;5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard</span></p></li><li><p><span>The one use case where Opus 5 earned straight 5s from me</span></p></li><li><p><span>My actual plan for using Opus 5 going forward</span></p></li></ol><div><hr></div><h3>In this episode, I cover:</h3><p>(<a href="https://www.youtube.com/watch?v=dfre9hN0HCs">00:00</a>) Opus 5 is here</p><p>(<a href="https://www.youtube.com/watch?v=dfre9hN0HCs&amp;t=195s">03:15</a>) First impressions</p><p>(<a href="https://www.youtube.com/watch?v=dfre9hN0HCs&amp;t=372s">06:12</a>) Opus 5 vs. GPT&#8209;5.6 Sol personality comparison</p><p>(<a href="https://www.youtube.com/watch?v=dfre9hN0HCs&amp;t=879s">14:39</a>) Claude Slop: the verbosity problem and why it makes my blood boil</p><p>(<a href="https://www.youtube.com/watch?v=dfre9hN0HCs&amp;t=1015s">16:55</a>) How the How I AI benchmark works (7 models, 6 tasks, blind scoring)</p><p>(<a href="https://www.youtube.com/watch?v=dfre9hN0HCs&amp;t=1110s">18:30</a>) Live benchmark results: the leaderboard reveal</p><p>(<a href="https://www.youtube.com/watch?v=dfre9hN0HCs&amp;t=1405s">23:25</a>) My verdict and how I&#8217;ll actually use Opus 5</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; Claude Opus 5:</span></p><p><span>&#8226; Anthropic blog: &#8203;&#8203;</span><a href="https://www.anthropic.com/news"><span>https://www.anthropic.com/news</span></a></p><p><span>&#8226; GPT&#8209;5.6 Sol: </span><a href="https://openai.com/index/previewing-gpt-5-6-sol/"><span>https://openai.com/index/previewing-gpt-5-6-sol/</span></a></p><p><span>&#8226; Sonnet 5: </span><a href="https://www.anthropic.com/news/claude-sonnet-5"><span>https://www.anthropic.com/news/claude-sonnet-5</span></a></p><p><span>&#8226; Gemini 3.1 Pro: </span><a href="https://deepmind.google/models/gemini/pro/"><span>https://deepmind.google/models/gemini/pro/</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[Computer and browser use in Codex (5 real examples)]]></title><description><![CDATA[Watch now | &#127897;&#65039; I show you exactly how I use browser use and computer use in Codex to QA my app, manage LinkedIn, and shop for Hawaii, including the under-prompting trick that makes frontier models work harder]]></description><link>https://www.lennysnewsletter.com/p/computer-and-browser-use-in-codex</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/computer-and-browser-use-in-codex</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Wed, 22 Jul 2026 12:03:38 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207953544/6d747defe32ee7f9e1cbcba69f37672a.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-lk63Sl-LRKE" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;lk63Sl-LRKE&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/lk63Sl-LRKE?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><span>Today I&#8217;m walking you through one of my absolute favorite AI features right now: browser and computer use via Codex (the ChatGPT desktop app). I use this every single day, personally and professionally, and I wanted to share the specific workflows I&#8217;ve built, the moments that surprised me, and the mental model that makes it actually click.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/lk63Sl-LRKE">YouTube</a>, <a href="https://open.spotify.com/episode/3MkEHP4y8OR9vGvhRBVBp7">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/computer-browser-use-in-codex-5-real-examples/id1809663079?i=1000777866899">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>How browser use and computer use work, and why the Codex desktop app plus Chrome extension is the combo I rely on</span></p></li><li><p><span>How I use Codex to QA my onboarding flow, including exhaustive mobile testing I would never do manually</span></p></li><li><p><span>Why under-prompting frontier models gets better results than detailed step-by-step instructions</span></p></li><li><p><span>How my husband EJ Lawless&#8217;s persona-impersonation trick surfaces friction points I can&#8217;t see as the builder</span></p></li><li><p><span>How I use browser use to get through my LinkedIn inbox without touching it myself</span></p></li><li><p><span>How I had Codex shop Free People&#8217;s sale and add 10 medium-size items to my cart (breastfeeding-friendly and Hawaii-ready)</span></p></li><li><p><span>How computer use can control iPhone mirroring so your Mac can technically operate your phone</span></p></li><li><p><span>Three more computer-use shortcuts: filling annoying forms, creating Google Sheets mid-workflow, and managing a router</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wcAU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wcAU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!wcAU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!wcAU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!wcAU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wcAU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:25952,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/207953544?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wcAU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!wcAU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!wcAU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!wcAU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624cb004-40c4-43c9-ab73-bcfac1a96212_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://runwayml.com/howIAI"><span>Runway</span></a></strong><span>&#8212;The creative AI platform for images, video, and more</span></p><p><strong><a href="https://hyperagent.com/howiai"><span>Hyperagent</span></a></strong><span>&#8212;Deploy fleets of agents that handle real work</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=106s">01:46</a>) What browser use and computer use actually are</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=188s">03:08</a>) Why I use Codex specifically and how the desktop app plus Chrome extension works</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=255s">04:15</a>) Use case 1: QA testing my onboarding flow</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=641s">10:41</a>) Results: 11 issues, one high-severity blocker, one Google Sheet with screenshots</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=730s">12:10</a>) Use case 2: persona testing</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=1100s">18:20</a>) Use case 3: LinkedIn inbox, hands-free</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=1237s">20:37</a>) Use case 4: AI personal shopper</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=1427s">23:47</a>) Rapid-fire uses: forms, iPhone mirroring, router access from out of state, Google Docs</p><p>(<a href="https://www.youtube.com/watch?v=lk63Sl-LRKE&amp;t=1610s">26:50</a>) Wrap-up</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; Codex (ChatGPT desktop app): </span><a href="https://openai.com/codex"><span>https://openai.com/codex</span></a></p><p><span>&#8226; Claude desktop app: </span><a href="https://claude.ai/download"><span>https://claude.ai/download</span></a></p><p><span>&#8226; Monologue (voice dictation for AI): </span><a href="https://monologue.app"><span>https://monologue.app</span></a></p><p><span>&#8226; iPhone mirroring (Apple): </span><a href="https://support.apple.com/en-us/111775"><span>https://support.apple.com/en-us/111775</span></a></p><p><span>&#8226; Google Sheets: </span><a href="https://sheets.google.com"><span>https://sheets.google.com</span></a></p><p></p><h3>Other reference:</h3><p><span>&#8226; Jesse Genet episode (How I AI): </span><a href="https://www.lennysnewsletter.com/p/5-openclaw-agents-run-my-home-finances?utm_source=publication-search"><span>https://www.lennysnewsletter.com/p/5-openclaw-agents-run-my-home-finances?utm_source=publication-search</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[🎙️ How I AI: How the founder of Morning Brew built a Claude content machine that never runs out of ideas]]></title><description><![CDATA[Your weekly listens from How I AI, part of the Lenny&#8217;s Podcast Network]]></description><link>https://www.lennysnewsletter.com/p/how-i-ai-how-the-founder-of-morning</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-i-ai-how-the-founder-of-morning</guid><dc:creator><![CDATA[Lenny Rachitsky]]></dc:creator><pubDate>Mon, 20 Jul 2026 15:01:56 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/95485dc8-c492-41a8-966d-957553d1df87_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gWeJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" width="1456" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:76503,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/177292431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h3><span>How the founder of Morning Brew built a Claude content machine that never runs out of ideas | Alex Lieberman (Tenex)</span></h3><div id="youtube2-1_jlukb7gm4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;1_jlukb7gm4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/1_jlukb7gm4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/1_jlukb7gm4">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/0Gc2VpGheLewqzz4VxX1r9">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/how-the-founder-of-morning-brew-built-a/id1809663079?i=1000777544732">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lOtp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lOtp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp 424w, https://substackcdn.com/image/fetch/$s_!lOtp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp 848w, https://substackcdn.com/image/fetch/$s_!lOtp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp 1272w, https://substackcdn.com/image/fetch/$s_!lOtp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lOtp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:16454,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/207347537?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lOtp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp 424w, https://substackcdn.com/image/fetch/$s_!lOtp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp 848w, https://substackcdn.com/image/fetch/$s_!lOtp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp 1272w, https://substackcdn.com/image/fetch/$s_!lOtp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe4185eb-eee3-4098-9869-2a7aa21c160f_1456x104.webp 1456w" sizes="100vw"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><ul><li><p><strong><a href="https://firecrawl.dev/?utm_source=newsletter&amp;utm_medium=partner&amp;utm_campaign=how_i_ai"><span>Firecrawl</span></a></strong>&#8212;Power AI agents with clean web data</p></li><li><p><strong><a href="https://www.try.customer.io/paid/how-i-ai?utm_medium=ads&amp;utm_source=how_i_ai&amp;utm_campaign=ads_q3_2026_how_i_ai_podcast&amp;utm_content=newsletter_placement_2&amp;utm_term=demand%20gen"><span>Customer.io</span></a></strong>&#8212;Build customer engagement campaigns from a single prompt</p></li></ul></blockquote><p>Claire sits down with Alex Lieberman, Co-founder and Co-Managing Partner of Tenex, to dig into the AI content system he&#8217;s built inside Claude. Alex walks through how he finds ideas, gets interviewed by AI, codifies his voice, and runs drafts through a &#8220;Writer&#8217;s Council&#8221; before posting. </p><h4>Biggest takeaways:</h4><ol><li><p><strong><span>The blank page is the real enemy of consistent content, and an AI Oracle makes it disappear.</span></strong><span> Alex&#8217;s Oracle scans seven days of Slack, Notion, Gmail, and meeting notes, then surfaces 15 ranked content spikes each day&#8212;half from internal sources, half from </span>the X and LinkedIn accounts he follows<span>. </span>Writing is actually the easy part; knowing what to say is where most people stall.</p></li><li><p><strong><span>Map your current workflow before you touch any AI tool.</span></strong><span> Alex&#8217;s process started with drawing out every step of content creation as it already existed, then rebuilding it from scratch with no constraints. He found that most companies discover huge inefficiencies in this step that have nothing to do with AI at all. AI just forced them to look.</span></p></li><li><p><strong><span>AI slop is mostly a people problem, not a model problem.</span></strong><span> Alex&#8217;s framing here is blunt: the only time the Content Machine produces slop is when the person being interviewed doesn&#8217;t share interesting enough ideas during the interview. The model is shaping clay, not inventing ideas. If the clay is bad, the sculpture will be bad.</span></p></li><li><p><strong><span>Your voice lives in a Markdown file, and that file changes everything.</span></strong><span> Alex&#8217;s Content Machine pulls from a personal voice guide that captures his top-performing posts, hook formulas, content structures, and even specific language patterns like &#8220;self-deprecating confidence.&#8221; </span>The system drafts against that file, which means it&#8217;s calibrated to how Alex actually writes rather than to some average of the internet.</p></li><li><p><strong><span>A reinforcement loop is how your AI system gets better every time you use it.</span></strong><span> After every piece, Alex&#8217;s Content Machine compares the original draft to the published version, extracts generalizable lessons, and asks if they should be added to a permanent lessons file. Over time, the system stops making the same mistakes. This is the kind of feedback loop most people never build.</span></p></li><li><p><strong><span>The interview panel is the step most people skip and the one that matters most.</span></strong><span> The system deploys six interviewer personas (Tim Ferriss, Joe Rogan, Larry King, Howard Stern, Barbara Walters, Michael Barrow) that ask follow-up questions until they&#8217;ve extracted enough specifics to write something real. </span>Alex voice-to-texts his answers via Wispr Flow, so every sentence in the final draft traces back to something he actually said.</p></li><li><p><strong><span>Your employees are your most underleveraged marketing channel, and most companies are actively suppressing them.</span></strong><span> Alex ran a version of employee content at Storyarb called &#8220;Own the Internet,&#8221; and it drove 40% of all inbound leads that quarter. At Tenex, he launched the Creator Cup with $5,000 in prizes to make posting feel like a team sport, not a chore. His argument to skeptical CEOs: the fastest way to lose a great employee is to make it impossible for them to build a personal brand.</span></p></li><li><p><strong><span>A bootstrapped company can out-distribute a venture-backed one if it&#8217;s willing to show the work.</span></strong><span> Alex&#8217;s case is simple: Tenex can&#8217;t rely on </span><em><span>TechCrunch</span></em><span> coverage or top-tier investor mentions, so the team&#8217;s content is the visibility strategy. If the Creator Cup gets him one great engineer, the $5,000 prize pool pays back many times over compared with a recruiting agency fee. For bootstrapped companies, distribution isn&#8217;t optional, but the whole game.</span></p></li><li><p><strong><span>The Writer&#8217;s Council won&#8217;t let a draft go out below a 9 out of 10.</span></strong><span> </span>Six writer personas (including a character Alex calls &#8220;the AI slop allergist&#8221;) each score the draft and run a revision loop until the aggregate clears the threshold. That scoring mechanism is also what keeps Alex from having to manually police every post, so the system holds its standard even when he&#8217;s in a hurry.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How I AI: Alex Lieberman&#8217;s 6-Step Workflow to Beat AI Slop: <a href="https://www.chatprd.ai/how-i-ai/alex-liebermans-6-step-workflow-to-beat-ai-slop"><span>https://www.chatprd.ai/how-i-ai/alex-liebermans-6-step-workflow-to-beat-ai-slop</span></a></p><p>&#8627; Build an AI Content Machine to Beat Writer&#8217;s Block: <a href="https://www.chatprd.ai/how-i-ai/workflows/build-an-ai-content-machine-to-beat-writers-block"><span>https://www.chatprd.ai/how-i-ai/workflows/build-an-ai-content-machine-to-beat-writers-block</span></a></p><p>&#8627; Create a Personalized AI-Powered Job Board for a Smarter Job Hunt: <a href="https://www.chatprd.ai/how-i-ai/workflows/create-a-personalized-ai-powered-job-board-for-a-smarter-job-hunt"><span>https://www.chatprd.ai/how-i-ai/workflows/create-a-personalized-ai-powered-job-board-for-a-smarter-job-hunt</span></a></p><p><span>&#8627;</span> Launch an AI-Powered Employee Advocacy Program: <a href="https://www.chatprd.ai/how-i-ai/workflows/launch-an-ai-powered-employee-advocacy-program"><span>https://www.chatprd.ai/how-i-ai/workflows/launch-an-ai-powered-employee-advocacy-program</span></a></p></div><div><hr></div><p>If you&#8217;re enjoying these episodes, reply and let me know what you&#8217;d love to learn more about: AI workflows, hiring, growth, product strategy&#8212;anything.</p><p>Catch you next week,<br>Lenny</p><p><em>P.S. Want every new episode delivered the moment it drops? Hit &#8220;Follow&#8221; on your favorite podcast app.</em></p>]]></content:encoded></item><item><title><![CDATA[How the founder of Morning Brew built a Claude content machine that never runs out of ideas and never sounds like slop | Alex Lieberman]]></title><description><![CDATA[Listen now | &#127897; How Alex Lieberman (Morning Brew) built a Claude workflow that interviews him before drafting, codes his voice in Markdown, and runs a six-persona revision loop before posting]]></description><link>https://www.lennysnewsletter.com/p/how-the-founder-of-morning-brew-built</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-the-founder-of-morning-brew-built</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Mon, 20 Jul 2026 12:04:47 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/207355289/37ca9e348cb5cb291449aaa2c86d2637.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-1_jlukb7gm4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;1_jlukb7gm4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/1_jlukb7gm4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong>Alex Lieberman</strong> co-founded Morning Brew in college and grew it into one of the most-read business newsletters in the world before selling it to Business Insider. Now he&#8217;s the co-founder and co-managing partner of Tenex. In this episode, Alex explains why distribution is becoming a durable moat, why founders and teams need to &#8220;climb Cringe Mountain,&#8221; and how he rebuilt his content process around AI without letting it produce generic slop. He walks us through every step of his Content Machine live: an Oracle that scans internal systems and the internet for content spikes, an interview panel that pulls out his real ideas, voice and style files that keep drafts sounding like him, an editorial council that scores and revises posts, and a lessons loop that learns from his feedback.</p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/1_jlukb7gm4">YouTube</a>, <a href="https://open.spotify.com/episode/0Gc2VpGheLewqzz4VxX1r9">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/how-the-founder-of-morning-brew-built-a/id1809663079?i=1000777544732">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>Why the blank page is the biggest friction point in content creation, and how an AI Oracle eliminates it</span></p></li><li><p><span>How to map your current workflow before you add any AI</span></p></li><li><p><span>How Alex built a six-step Content Machine in Claude that goes from idea spike to publishable post</span></p></li><li><p><span>Why the interview step (not the drafting step) is where AI slop actually comes from</span></p></li><li><p><span>How to codify your voice in a Markdown file so an AI drafts in your register, not the internet&#8217;s average</span></p></li><li><p><span>Why your employees are your most underleveraged marketing channel right now</span></p></li><li><p><span>How the Tenex Creator Cup turned content creation into a team sport with a $5,000 prize pool</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ztXJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ztXJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!ztXJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!ztXJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!ztXJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ztXJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:26639,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/207355289?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ztXJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!ztXJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!ztXJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!ztXJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F55a9e49c-a4d3-43c6-8591-6c30d7ba8367_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://firecrawl.dev/?utm_source=newsletter&amp;utm_medium=partner&amp;utm_campaign=how_i_ai"><span>Firecrawl</span></a></strong><span>&#8212;Power AI agents with clean web data</span></p><p><strong><a href="https://www.try.customer.io/paid/how-i-ai?utm_medium=ads&amp;utm_source=how_i_ai&amp;utm_campaign=ads_q3_2026_how_i_ai_podcast&amp;utm_content=newsletter_placement_2&amp;utm_term=demand%20gen"><span>Customer.io</span></a></strong><span>&#8212;Build customer engagement campaigns from a single prompt</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=1_jlukb7gm4">00:00</a>) Introduction to Alex Lieberman</p><p>(<a href="https://www.youtube.com/watch?v=1_jlukb7gm4&amp;t=155s">02:35</a>) Why Alex built a content machine</p><p>(<a href="https://www.youtube.com/watch?v=1_jlukb7gm4&amp;t=416s">06:56</a>) Alex&#8217;s thoughts on AI slop</p><p>(<a href="https://www.youtube.com/watch?v=1_jlukb7gm4&amp;t=540s">09:00</a>) Mapping the workflow from scratch</p><p>(<a href="https://www.youtube.com/watch?v=1_jlukb7gm4&amp;t=804s">13:24</a>) The six-step Content Machine setup</p><p>(<a href="https://www.youtube.com/watch?v=1_jlukb7gm4&amp;t=1391s">23:11</a>) Live demo: Oracle, Interview Panel, and Writer&#8217;s Council in action</p><p>(<a href="https://www.youtube.com/watch?v=1_jlukb7gm4&amp;t=1838s">30:38</a>) Employee advocacy: the Tenex Creator Cup and $5K prize pool</p><p>(<a href="https://www.youtube.com/watch?v=1_jlukb7gm4&amp;t=2205s">36:45</a>) Lightning round: great engineers, AI use cases, slop fixes</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; Claude / Claude Code (Anthropic): </span><a href="https://claude.ai"><span>https://claude.ai</span></a></p><p><span>&#8226; </span>Wispr<span> Flow (voice-to-text transcription): </span><a href="https://wisprflow.ai">https://wisprflow.ai</a></p><p><span>&#8226; Notion: </span><a href="https://notion.so"><span>https://notion.so</span></a></p><p><span>&#8226; Linear: </span><a href="https://linear.app"><span>https://linear.app</span></a></p><p><span>&#8226; Slack: </span><a href="https://slack.com"><span>https://slack.com</span></a></p><p></p><h3>Other references:</h3><p><span>&#8226; Morgan Housel: </span><a href="https://www.morganhousel.com"><span>https://www.morganhousel.com</span></a></p><p><span>&#8226; David Perell: </span><a href="https://perell.com"><span>https://perell.com</span></a></p><p><span>&#8226; Shaan Puri / My First Million podcast: </span><a href="https://www.mfmpod.com"><span>https://www.mfmpod.com</span></a></p><p><span>&#8226; Gary Vaynerchuk: </span><a href="https://garyvaynerchuk.com"><span>https://garyvaynerchuk.com</span></a></p><p></p><h3>Where to find <span>Alex Lieberman:</span></h3><p><span>X: </span><a href="https://x.com/businessbarista"><span>https://x.com/businessbarista</span></a></p><p><span>LinkedIn: </span><a href="https://www.linkedin.com/in/alex-lieberman/"><span>https://www.linkedin.com/in/alex-lieberman/</span></a></p><p><span>Tenex: </span><a href="https://www.tenex.co/"><span>https://www.tenex.co/</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[🎙️ How I AI: GPT-5.6 review, How a solo builder runs 24/7 local AI, and What an agent harness is and how to build one]]></title><description><![CDATA[Your weekly listens from How I AI, part of the Lenny&#8217;s Podcast Network]]></description><link>https://www.lennysnewsletter.com/p/how-i-ai-gpt-56-review-how-a-solo</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-i-ai-gpt-56-review-how-a-solo</guid><dc:creator><![CDATA[Lenny Rachitsky]]></dc:creator><pubDate>Mon, 13 Jul 2026 15:02:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c3eccdc2-ac57-47dc-af99-4daa75d21bc7_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gWeJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" width="1456" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:76503,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/177292431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h3><span>What a harness </span>is<span> and how to build one with Claude Agent SDK</span></h3><div id="youtube2-ofS-4RRw9zw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;ofS-4RRw9zw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/ofS-4RRw9zw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/ofS-4RRw9zw">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/3t4Osk3xFYKBefEqjy5git">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/what-a-harness-is-and-how-to-build-one-with-claude-agent-sdk/id1809663079?i=1000775947149">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZnxH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZnxH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:30260,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/205647252?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!ZnxH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 1456w" sizes="100vw"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><p><strong><a href="https://bolt.new/partner/howiai">Bolt.new</a></strong>&#8212;Turn your idea into a real product</p><p><strong><a href="https://www.try.customer.io/paid/how-i-ai?utm_medium=ads&amp;utm_source=how_i_ai&amp;utm_campaign=ads_q3_2026_how_i_ai_podcast&amp;utm_content=newsletter_placement_1&amp;utm_term=demand%20gen">Customer.io</a></strong>&#8212;Build customer engagement campaigns from a single prompt</p></blockquote><p>Claire explains why harnesses matter and when they&#8217;re better than general-purpose tools like Claude Code or Codex, and walks through the custom Claude Agent SDK harness she built to automate Sentry bug triage at ChatPRD. You&#8217;ll see how she structured the workflow, encoded permissions, connected tools like Sentry and Linear, and turned a repeatable engineering task into something an agent can run more consistently every time.</p><h4>Biggest takeaways:</h4><ol><li><p><strong><span>A harness is just code around an AI agent&#8212;nothing more mysterious than that.</span></strong><span> The term has taken on an almost mythical quality in engineering circles, but </span>I<span> strip it down in this episode: a harness is code you write to make an AI agent more effective at a specific job. Cursor is a complex harness. Claude Code is a complex harness. Yours can be eight files and a terminal UI.</span></p></li><li><p><strong><span>Build a harness when the same workflow needs the same setup and the same outcomes every time. </span></strong><span>The trigger is recognizing a job that is partly deterministic (defined steps, defined tools) and partly non-deterministic (the AI figures out root causes and writes the report). Sentry bug triage qualified because every investigation follows the same evidence-gathering process and ends with the same artifact bundle.</span></p></li><li><p><strong><span>Opinionated tool adapters beat general MCP access for specialized workflows. </span></strong><span>Rather than giving the agent broad access to the Sentry MCP and letting it wander through traces, I built a custom Sentry adapter that pulls exactly what matters for a bug report and nothing else. That specificity makes the agent faster, cheaper, and less likely to go off-script.</span></p></li><li><p><strong><span>Encoding permissions in the harness removes the need to prompt them every single time.</span></strong><span> In a general-purpose coding tool, you have to remember to say &#8220;investigate only, do not write code.&#8221; In my harness, that is a flag in the interface. I click &#8220;investigate,&#8221; paste the Sentry link, and the agent already knows its constraints without being told.</span></p></li><li><p><strong><span>Structured artifacts are what separate a one-off investigation from a team-wide resource.</span></strong><span> Every time my harness runs, it outputs a task log, a Sentry issue brief, relevant logs, a worker report, and an HTML summary file. That artifact bundle means the engineering team gets a consistent, scannable record of every bug investigation without anyone having to write it up manually.</span></p></li><li><p><strong><span>A harness lets you do multi-model routing in ways a single general-purpose tool never could.</span></strong><span> Claude Code is Claude. Codex is GPT. A custom harness using the Claude Agent SDK lets you pick the right model per step, enforce different tool policies per invocation, and swap models over time without changing how the interface works. That flexibility is one of the strongest arguments for owning the harness layer yourself.</span></p></li><li><p><strong><span>The open chat field has been good enough, until it stopped being good enough.</span></strong><span> I acknowledge in this episode that just typing into Claude Code has produced real work. But this marks a shift in my thinking: general-purpose agents are now better used to orchestrate specialized harnesses than to do every job themselves. Giving a constrained agent a specific harness gets more consistent output than giving a powerful agent an open prompt.</span></p></li></ol><div class="callout-block" data-callout="true"><h4>Blog from this episode:</h4><p>How I Built a Custom AI Harness with the Claude Agent SDK for Bug Triage: <a href="https://www.chatprd.ai/how-i-ai/how-i-built-a-custom-ai-harness">https://www.chatprd.ai/how-i-ai/how-i-built-a-custom-ai-harness</a></p></div><div><hr></div><h3><span>This solo builder runs 24/7 local AI on his own hardware | Alex Finn</span></h3><div id="youtube2-dAQsmhAiews" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;dAQsmhAiews&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/dAQsmhAiews?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/dAQsmhAiews">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/6eZP9tlwX4U10YGjmm5tk9?si=yoHDBZb_RZy-m2ou4PAuAQ">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/this-solo-builder-runs-24-7-local-ai-on-his-own-hardware/id1809663079?i=1000776582970">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!i8Y0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!i8Y0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:19309,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/205689325?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!i8Y0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><p><strong><a href="https://runwayml.com/howIAI">Runway</a></strong>&#8212;The creative AI platform for images, video, and more</p><p><strong><a href="https://atlassian.com/howiai">Jira Product Discovery</a></strong>&#8212;Prioritize with insights, build with confidence</p></blockquote><p>Claire talks with Alex Finn about how he built a 24/7 local AI fleet using Mac Studios, a DGX Spark, an RTX 5090, and a custom dashboard to keep agents running around the clock. Alex breaks down what each machine is actually good for, how he routes work across local models like GLM, Qwen, and Ornith, and why &#8220;unlimited inference&#8221; changes the entire economics of AI workflows. They also get into his Claude Code build-and-review loop, his OpenClaw and Hermes setup, and the surprisingly practical playbook behind running your own always-on software factory.</p><h4>Biggest takeaways:</h4><ol><li><p><strong><span>The case for local AI isn&#8217;t ROI; it&#8217;s unlimited inference.</span></strong><span> The math on a $10,000 Mac Studio vs. a $20 ChatGPT subscription only looks crazy until you run an agent 24/7. At that scale, cloud APIs get expensive fast, and local models running around the clock open use cases that simply aren&#8217;t economically viable otherwise. Alex runs security scans, code reviews, and social signal monitoring on a continuous loop that would cost thousands a month in cloud credits.</span></p></li><li><p><strong><span>Each hardware tier has a job.</span></strong><span> Mac Studio handles massive models slowly but at Opus-level intelligence (Alex runs GLM 5.2, which he calls Opus 4.8-equivalent, on a single Mac Studio). DGX Spark is the sweet spot: 128 GB of Nvidia unified memory plus CUDA speed, enough for models like Qwen 3.6 running fast. The RTX 5090 has only 32 GB of VRAM but is cloud-speed fast. Buy for the task, not the spec sheet.</span></p></li><li><p><strong><span>Tailscale is the connective tissue for a multi-machine setup.</span></strong><span> Once all your machines are on the same Tailscale network, one agent (OpenClaw or Hermes) can hop across them, check the hardware, load the right model, and get it running without any manual configuration. Alex says there&#8217;s truly no technical knowledge required once Tailscale is installed, and he recommends it even if you only have one machine, because it also lets you test local apps from your phone.</span></p></li><li><p><strong><span>Local models are the BDR; Claude Code is the closer.</span></strong><span> Alex&#8217;s security scanning loop is a good example of the hybrid model that actually works. A local model (GLM 5.2) scans code every 20 minutes and dumps findings into a Markdown file. Claude Code checks that report once a day and decides what&#8217;s real and worth fixing. The local model does the volume work cheaply; the frontier model does the judgment work precisely. Trying to run Claude Code every 20 minutes instead would cost thousands a month.</span></p></li><li><p><strong><span>The software factory runs on two loops and a rocket emoji.</span></strong><span> Every morning, Alex does a planning session in Claude with a &#8220;morning build&#8221; prompt that produces a task list for his SaaS. The build loop picks those up and starts shipping. The review loop checks the work. When something passes review, Alex gets a Slack ping, and leaving a rocket emoji on it triggers an automated merge. He goes from morning brief to reviewing and merging code without touching the keyboard again until he does the approval round.</span></p></li><li><p><strong><span>OpenClaw and Hermes fill different needs, and you probably want both.</span></strong><span> Alex prefers OpenClaw for the &#8220;big bang&#8221; wow moments and the emotional connection (his words). But Hermes has been more reliable under repeated updates. His solution is redundancy: three Hermes agents and two OpenClaw agents running simultaneously, so when three of the five are broken (which happens), the other two can fix them. The failover is deliberate, not accidental.</span></p></li><li><p><strong><span>Task allocation by model intelligence is the skill that makes the fleet useful.</span></strong><span> GLM 5.2 is Opus-level smart but painfully slow, so it gets the deep, latency-tolerant work. Qwen 3.6 is quick and good enough to read Twitter for product signals. Ornith 1.0, a Qwen fine-tune with reinforcement learning baked in for coding, has beaten Qwen on every eval Alex has run and runs comfortably on a DGX Spark. The insight is that &#8220;smartest model everywhere&#8221; is wasteful; matching model intelligence to task complexity is what makes ambient AI economically coherent.</span></p></li><li><p><strong><span>The vague posting about loops is partly a competitive moat.</span></strong><span> Alex&#8217;s theory: the companies building the best AI coding infrastructure (including OpenAI and Anthropic themselves) have internal loop systems that are their last real competitive advantage. If you can pump out high-quality code faster than anyone else because your build-review loop is better, you don&#8217;t go publishing a how-to. Claire&#8217;s counter-theory: most people vague-post because their loops are boring and vagueness gets more engagement than specifics. Both are probably true, depending on who&#8217;s doing the posting.</span></p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How I AI: Alex Finn&#8217;s Local AI Fleet and Automated Software Factory: <a href="https://www.chatprd.ai/how-i-ai/alex-finns-local-ai-fleet-and-automated-software-factory">https://www.chatprd.ai/how-i-ai/alex-finns-local-ai-fleet-and-automated-software-factory</a></p><p><strong>&#8627; </strong>How to Assemble a Multi-Machine Local AI Fleet: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-assemble-a-multi-machine-local-ai-fleet">https://www.chatprd.ai/how-i-ai/workflows/how-to-assemble-a-multi-machine-local-ai-fleet</a></p><p><strong>&#8627; </strong>How to Build an Automated Software Factory with AI Agents: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-build-an-automated-software-factory-with-ai-agents">https://www.chatprd.ai/how-i-ai/workflows/how-to-build-an-automated-software-factory-with-ai-agents</a></p><p><strong>&#8627; </strong>How to Set Up a Continuous Code Security Scan Using a Hybrid AI Workflow: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-set-up-a-continuous-code-security-scan-using-a-hybrid-ai-workflow">https://www.chatprd.ai/how-i-ai/workflows/how-to-set-up-a-continuous-code-security-scan-using-a-hybrid-ai-workflow</a></p></div><div><hr></div><h3>GPT-5.6 Sol vs. Claude Fable: Why OpenAI&#8217;s new model crushes my benchmark</h3><div id="youtube2-gAWbvEwUoiI" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;gAWbvEwUoiI&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/gAWbvEwUoiI?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/gAWbvEwUoiI">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/2gFMl8roD8s9bSOKYmgEI3">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/gpt-5-6-sol-vs-claude-fable-why-openais-new-model-crushes/id1809663079?i=1000776152330">Apple Podcasts</a></strong></p></div><p>Claire puts GPT-5.6 Sol head-to-head with Claude Fable, Sonnet 5, and the rest of the GPT-5.6 lineup using her own five-part benchmark for real product work. The result: Sol becomes her new daily driver. Claire breaks down exactly why and also shares where she&#8217;d still use Terra, Sonnet, or Fable instead.</p><h4>Biggest takeaways:</h4><ol><li><p><strong>GPT-5.6 Sol is the most practically effective model I&#8217;ve tested, even if Fable is theoretically smarter.</strong> I ran a five-category benchmark across PRDs, prototypes, wireframes, debugging, and agentic voice, and Sol had the highest taste score by a significant margin on the 70% Claire/30% machine split. That gap between &#8220;hyper-intelligent&#8221; and &#8220;actually ships&#8221; is real, and for product work Sol wins.</p></li><li><p><strong>Full-fidelity prototypes from Sol are more functional and more opinionated than anything else I&#8217;ve tested.</strong> Across a doc scheduler, a dev tools incident triage site, and a consumer habit tracker app, Sol consistently produced designs with better visual hierarchy, semantic color use, and working interactivity. Fable&#8217;s outputs were fine; Sol&#8217;s were the ones I&#8217;d actually show a stakeholder.</p></li><li><p><strong>Sol&#8217;s writing is just easier to work with.</strong> Fable writes like it has never met a human before, incredibly pedantic and almost inscrutable when you need to collaborate. Sol writes like a normal person, and that difference compounds fast when you&#8217;re iterating on PRDs or talking to an agent all day.</p></li><li><p><strong>For PRD writing specifically, GPT-5.6 Terra might be the better pick.</strong> I asked Sol to greenfield-rebuild my approach to PRDs for 2026, and while Sol&#8217;s output was excellent, Terra&#8217;s clean, direct, no-frills business writing made me think it&#8217;s the right call when you want crisp, fast documentation without extra flair.</p></li><li><p><strong>Fable gets too locked in its own frameworks; Sol is willing to reconsider.</strong> I had a hardened tool-calling loop in my prototyping product that only GPT-5.5 could run. Fable insisted it was a model problem and refused to budge. The moment I switched to Codex and told it to just fix it, Sol got Sonnet 5 working in one shot. That kind of practical flexibility is exactly what you need when building real products.</p></li><li><p><strong>Sonnet 5 is still my favorite for agentic voice in Open Claw.</strong> Even after this whole benchmark, I gave Sonnet 5 a gold star for voice: aside from the dashes, it sounds the most human. I use Sonnet for my OpenClaw and I&#8217;m not changing that. Sol did a worse job on agentic voice overall, and I still can&#8217;t get GPT models to run well in my OpenClaw setup.</p></li><li><p><strong>GPT-5.6&#8217;s video editing via Codex is one of my favorite new workflows.</strong> I dropped in a full recording from a talk I gave at Cursor&#8217;s event, asked for five hype-video clips, gave feedback on pacing and orientation, and had shareable social clips I could drop into CapCut in a fraction of the time. This use case alone justifies experimenting with GPT-5.6.</p></li><li><p><strong>Browser use with Codex plus GPT-5.6 and @Chrome is the best agentic workflow I&#8217;ve found.</strong> I opened LinkedIn, told it to reply to high-value messages from executives and ChatPRD fans, and it burned through roughly 500 messages. I&#8217;ve also used it to test web apps and fill out forms. When I got rolled back to GPT-5.5 temporarily, my life was measurably worse. Learn @Chrome and just let it rip.</p></li><li><p><strong>The &#8220;forest green&#8221; tell is real, so name it in your prompts.</strong> Sol has a strong aesthetic bias toward what feels like a woodland-themed palette hardcoded somewhere in its system. You will see a lot of green. I told the OpenAI team, I&#8217;m noting it here, and I&#8217;m already prompting against it when I want a different aesthetic direction.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How I AI: My GPT-5.6 Sol Benchmark &amp; 4 Game-Changing Workflows (vs. Fable): <a href="https://www.chatprd.ai/how-i-ai/my-gpt-56-sol-benchmark-game-changing-workflows">https://www.chatprd.ai/how-i-ai/my-gpt-56-sol-benchmark-game-changing-workflows</a></p><p><strong>&#8627; </strong>How to Automate LinkedIn Messaging with AI Browser Control: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-automate-linkedin-messaging-with-ai-browser-control">https://www.chatprd.ai/how-i-ai/workflows/how-to-automate-linkedin-messaging-with-ai-browser-control</a></p><p><strong>&#8627; </strong>How to Quickly Create Social Media Video Clips Using AI: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-quickly-create-social-media-video-clips-using-ai">https://www.chatprd.ai/how-i-ai/workflows/how-to-quickly-create-social-media-video-clips-using-ai</a></p><p><strong>&#8627; </strong>How to Build a Gamified Homework App with AI in a Single Shot: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-gamified-homework-app-with-ai-in-a-single-shot">https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-gamified-homework-app-with-ai-in-a-single-shot</a></p></div><div><hr></div><p>If you&#8217;re enjoying these episodes, reply and let me know what you&#8217;d love to learn more about: AI workflows, hiring, growth, product strategy&#8212;anything.</p><p>Catch you next week,<br>Lenny</p><p><em>P.S. Want every new episode delivered the moment it drops? Hit &#8220;Follow&#8221; on your favorite podcast app.</em></p>]]></content:encoded></item><item><title><![CDATA[This solo builder runs 24/7 local AI on his own hardware | Alex Finn]]></title><description><![CDATA[Watch now | &#127897;&#65039; Alex Finn breaks down his five-computer local AI setup, the Claude Code build-and-review loop that ships features while he sleeps, and why unlimited local intelligence beats a $20 subscription]]></description><link>https://www.lennysnewsletter.com/p/this-solo-builder-runs-247-local</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/this-solo-builder-runs-247-local</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Mon, 13 Jul 2026 12:05:20 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/205689325/b872277f3a2dbf756729413f4210989b.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-dAQsmhAiews" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;dAQsmhAiews&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/dAQsmhAiews?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong><span>Alex Finn</span></strong><span> is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I&#8217;ve come across: three Mac Studio 512 GB machines, a DGX Spark, and a custom RTX 5090 build, all coordinated through a fleet dashboard he built himself. He&#8217;s spent five months figuring out which local models belong on which machines, how to wire them to Claude Code loops, and how to get a software factory running without babysitting it.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/dAQsmhAiews">YouTube</a>, <a href="https://open.spotify.com/show/4aRP2XSavdtrLG5FZoonOK">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/this-solo-builder-runs-24-7-local-ai-on-his-own-hardware/id1809663079?i=1000776582970">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>How Alex chose between a Mac Studio (512 GB unified memory), DGX Spark, and RTX 5090, and what each is actually good for</span></p></li><li><p><span>Why Tailscale is worth installing even on a single machine, and how it lets one agent manage your entire hardware fleet</span></p></li><li><p><span>How the build loop and review loop in Claude Code work</span></p></li><li><p><span>How to allocate tasks by machine and model</span></p></li><li><p><span>Why unlimited local inference changes the use-case math in a way a $20 cloud subscription never can</span></p></li><li><p><span>What OpenClaw and Hermes are each best suited for, and why Alex runs five agents total with failover baked in</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!i8Y0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!i8Y0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:19309,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/205689325?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!i8Y0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!i8Y0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F520dfed9-3048-4ad4-92fe-29e8dff5abe7_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://runwayml.com/howIAI"><span>Runway</span></a></strong><span>&#8212;The creative AI platform for images, video, and more</span></p><p><strong><a href="https://atlassian.com/howiai"><span>Jira Product Discovery</span></a></strong><span>&#8212;Prioritize with insights, build with confidence</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=178s">02:58</a>) Alex&#8217;s hardware stack</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=228s">03:48</a>) What &#8220;ambient AI&#8221; means</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=255s">04:15</a>) Alex&#8217;s red-pill moment with OpenClaw</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=424s">07:04</a>) Mac Studio vs. DGX Spark vs. RTX 5090</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=804s">13:24</a>) How to set up local models with no technical knowledge (Tailscale + OpenClaw/Hermes)</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=1036s">17:16</a>) Fleet control dashboard: assigning 24/7 tasks across machines</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=1242s">20:42</a>) Local models as security scanners feeding Claude Code</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=1345s">22:25</a>) How Alex allocates GLM 5.2, Qwen 3.6, and Ornith 1.0 by task</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=1468s">24:28</a>) OpenClaw vs. Hermes: the honest comparison</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=1615s">26:55</a>) The software factory: build loop, review loop, rocket emoji</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=1915s">31:55</a>) Lightning round: favorite hardware, favorite model, prompting style</p><p>(<a href="https://www.youtube.com/watch?v=dAQsmhAiews&amp;t=2086s">34:46</a>) Where to find Alex</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; Claude Code: </span><a href="https://claude.ai/code"><span>https://claude.ai/code</span></a></p><p><span>&#8226; OpenClaw: </span><a href="https://openclaw.ai/"><span>https://openclaw.ai/</span></a></p><p><span>&#8226; Hermes: </span><a href="https://hermes-agent.nousresearch.com/"><span>https://hermes-agent.nousresearch.com/</span></a></p><p><span>&#8226; Tailscale: </span><a href="https://tailscale.com/"><span>https://tailscale.com/</span></a></p><p><span>&#8226; Codex (OpenAI): </span><a href="https://openai.com/codex"><span>https://openai.com/codex</span></a></p><p><span>&#8226; GLM 5.2 (z.ai): </span><a href="https://huggingface.co/zai-org/GLM-5.2"><span>https://huggingface.co/zai-org/GLM-5.2</span></a></p><p><span>&#8226; Qwen 3.6 (Alibaba): </span><a href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B"><span>https://huggingface.co/Qwen/Qwen3.6-35B-A3B</span></a></p><p><span>&#8226; Ornith 1.0: </span><a href="https://github.com/deepreinforce-ai/Ornith-1"><span>https://github.com/deepreinforce-ai/Ornith-1</span></a></p><p><span>&#8226; Gemma 4: </span><a href="https://huggingface.co/collections/google/gemma-4"><span>https://huggingface.co/collections/google/gemma-4</span></a></p><p><span>&#8226; Playwright (browser testing): </span><a href="https://playwright.dev/"><span>https://playwright.dev/</span></a></p><p><span>&#8226; Vercel (preview deploys): </span><a href="https://vercel.com/"><span>https://vercel.com/</span></a></p><p></p><h3>Other references:</h3><p><span>&#8226; DGX Spark (Nvidia): </span><a href="https://www.nvidia.com/en-us/products/workstations/dgx-spark/"><span>https://www.nvidia.com/en-us/products/workstations/dgx-spark/</span></a></p><p><span>&#8226; Mac Studio (Apple): </span><a href="https://www.apple.com/mac-studio/"><span>https://www.apple.com/mac-studio/</span></a></p><p><span>&#8226; How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex: </span><a href="https://www.lennysnewsletter.com/p/how-to-design-ai-agent-loops-schedules"><span>https://www.lennysnewsletter.com/p/how-to-design-ai-agent-loops-schedules</span></a></p><p></p><h3>Where to find <span>Alex Finn:</span></h3><p><span>LinkedIn: </span><a href="https://www.linkedin.com/in/alex-finn-1848684a/">https://www.linkedin.com/in/alex-finn-1848684a</a></p><p><span>YouTube: </span><a href="https://www.youtube.com/@AlexFinnOfficial"><span>https://www.youtube.com/@AlexFinnOfficial</span></a></p><p><span>X: </span><a href="https://x.com/AlexFinn"><span>https://x.com/AlexFinn</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark]]></title><description><![CDATA[Watch now | &#127897;&#65039;GPT 5.6-Sol beats Fable on prototypes, PRDs, and browser use in my 5-category How I AI benchmark, and here's exactly where each model earns its spot.]]></description><link>https://www.lennysnewsletter.com/p/gpt-56-sol-vs-claude-fable-why-openais</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/gpt-56-sol-vs-claude-fable-why-openais</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Thu, 09 Jul 2026 17:33:57 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/206191009/f1f9241f1b8ff9181b201a5bed6becf7.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-gAWbvEwUoiI" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;gAWbvEwUoiI&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/gAWbvEwUoiI?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and agentic voice. Sol won by a meaningful margin on my Claire Weighted Index (70% my taste, 30% Terminal Bench 2.1), and I also tested two use cases I can't stop thinking about: building a gamified homework tracking app for my kids in one shot with Codex, and browser automation with Chrome that burned through 500 LinkedIn replies while I did literally nothing.</p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/gAWbvEwUoiI">YouTube</a>, <a href="https://open.spotify.com/episode/2gFMl8roD8s9bSOKYmgEI3">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/gpt-5-6-sol-vs-claude-fable-why-openais-new-model-crushes/id1809663079?i=1000776152330">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>How I scored five AI models (including GPT 5.6 Sol, Fable 5, and Sonnet 5) using my &#8220;Claire Weighted Index&#8221; benchmark across PRDs, prototypes, code, and agentic voice</span></p></li><li><p><span>The difference between GPT-5.6 Sol (Terra) and Sol for PRD writing</span></p></li><li><p><span>How Fable&#8217;s precision and pedantry made it harder to collaborate with, and the exact moment Sol broke through where Fable got stuck</span></p></li><li><p><span>Why Sonnet 5 is still my go-to for agentic voice in OpenClaw, even after this whole benchmark</span></p></li><li><p><span>How I used GPT-5.6 Sol in Codex to build a fully gamified homework tracking app for my kids in one shot</span></p></li><li><p><span>The video editing use case that saved me hours clipping a talk I gave at Cursor&#8217;s event</span></p></li><li><p><span>How to use Codex plus GPT-5.6 and Chrome for browser automation, and why this is my single most-loved use case right now</span></p></li></ol><div><hr></div><h3>In this episode, I cover:</h3><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=70s">01:10</a>) The three GPT-5.6 models: Sol, Terra, Luna</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=137s">02:17</a>) Pricing: Sol vs. Fable API costs</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=204s">03:24</a>) The How I AI benchmark</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=303s">05:03</a>) Claire-weighted Index results</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=420s">07:00</a>) Per-task winners: prototypes, PRDs, agentic voice</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=719s">11:59</a>) What Claire actually rewards</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=800s">13:20</a>) Full-fidelity prototype side-by-sides (Sol vs. Fable)</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=1065s">17:45</a>) Wireframes</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=1099s">18:19</a>) Agentic voice</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=1155s">19:15</a>) Where Sol is better than other models</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=1436s">23:56</a>) Gamified kids&#8217; homework app, built in one shot</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=1682s">28:02</a>) Fable&#8217;s pedantry problem and how Sol broke through it</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=1909s">31:49</a>) Two bonus use cases: video editing and browser use</p><p>(<a href="https://www.youtube.com/watch?v=gAWbvEwUoiI&amp;t=2108s">35:08</a>) Final summary and model recommendations</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; GPT 5.6 (Sol, Terra, Luna): </span><a href="https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna"><span>https://help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna</span></a></p><p><span>&#8226; Codex: </span><a href="https://openai.com/codex"><span>https://openai.com/codex</span></a></p><p><span>&#8226; ChatPRD: </span><a href="https://www.chatprd.ai/"><span>https://www.chatprd.ai/</span></a></p><p><span>&#8226; CapCut: </span><a href="https://www.capcut.com/"><span>https://www.capcut.com/</span></a></p><p><span>&#8226; Math Academy: </span><a href="https://www.mathacademy.com/"><span>https://www.mathacademy.com/</span></a></p><p></p><h3>Other references:</h3><p><span>&#8226; Cursor event where Claire spoke on the future of PM: </span><a href="https://www.youtube.com/watch?v=4CAFK-rc26A"><span>https://www.youtube.com/watch?v=4CAFK-rc26A</span></a></p><p><span>&#8226; ChatPRD blog (where benchmark outputs will be published): </span><a href="https://www.chatprd.ai/"><span>https://www.chatprd.ai/</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[What a harness is and how to build one with Claude Agent SDK]]></title><description><![CDATA[Listen now | &#127897;&#65039; I built a custom Claude Agent SDK harness to automate Sentry bug triage, saved the "dear agent, please fix this" prompt forever, and show you exactly how to build your own]]></description><link>https://www.lennysnewsletter.com/p/what-a-harness-is-and-how-to-build</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/what-a-harness-is-and-how-to-build</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Wed, 08 Jul 2026 12:03:35 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/205647252/8ab2b19e7844865055b49d826c6a4473.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-ofS-4RRw9zw" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;ofS-4RRw9zw&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/ofS-4RRw9zw?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><span>Everybody is saying, &#8220;It&#8217;s not the model, it&#8217;s the harness,&#8221; but almost nobody stops to explain what a harness actually is. So I did. I built one live on the show: a Sentry bug-debugging harness for my company ChatPRD, using the Claude Agent SDK, a custom terminal UI built with the Ink library, and opinionated adapters for Sentry, Linear, GitHub, and Vercel. The harness handles evidence gathering, root-cause analysis, and follow-up artifact creation, all without me needing to type &#8220;dear agent, please fix this bug&#8221; ever again. I also walk through the architecture, share the code structure, and give you the exact process I used so you can build your own harness for any repetitive, structured workflow in your business.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/ofS-4RRw9zw">YouTube</a>, <a href="https://open.spotify.com/episode/3t4Osk3xFYKBefEqjy5git">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/what-a-harness-is-and-how-to-build-one-with-claude-agent-sdk/id1809663079?i=1000775947149">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>What a harness actually is</span></p></li><li><p><span>When to build a harness versus when to stick with a general-purpose tool like Claude Code or Codex</span></p></li><li><p><span>How to encode specific permissions into a harness</span></p></li><li><p><span>The three components every harness needs</span></p></li><li><p><span>How I used GPT-5.5 and Claude Opus to build the harness code itself (and where they both initially resisted)</span></p></li><li><p><span>How to structure the artifacts your harness produces so the whole team can use the output</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZnxH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZnxH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:30260,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/205647252?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZnxH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!ZnxH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10f7db70-0c90-4b49-acbf-615db3f3790b_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://bolt.new/partner/howiai"><span>Bolt.new</span></a></strong>&#8212;Turn your idea into a real product</p><p><strong><a href="https://www.try.customer.io/paid/how-i-ai?utm_medium=ads&amp;utm_source=how_i_ai&amp;utm_campaign=ads_q3_2026_how_i_ai_podcast&amp;utm_content=newsletter_placement_1&amp;utm_term=demand%20gen"><span>Customer.io</span></a></strong><span>&#8212;Build customer engagement campaigns from a single prompt</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw">00:00</a>) What is an AI harness?</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=199s">03:19</a>) When to build a harness</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=273s">04:33</a>) Why Claire picked bug triage</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=360s">06:00</a>) Why not just use Claude Code?</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=468s">07:48</a>) Demo: The custom harness interface</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=664s">11:04</a>) Architecture: runs, tasks, tools, and artifacts</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=824s">13:44</a>) Building it with Codex and Claude</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=908s">15:08</a>) Code map and file layout</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=1011s">16:51</a>) A look at the code</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=1158s">19:18</a>) The live investigation result</p><p>(<a href="https://www.youtube.com/watch?v=ofS-4RRw9zw&amp;t=1261s">21:01</a>) How to build your own harness</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; Claude Agent SDK (Anthropic): </span><a href="https://code.claude.com/docs/en/agent-sdk/overview"><span>https://code.claude.com/docs/en/agent-sdk/overview</span></a></p><p><span>&#8226; Claude Sonnet 4.6 (model used inside the harness): </span><a href="https://www.anthropic.com/news/claude-sonnet-4-6"><span>https://www.anthropic.com/news/claude-sonnet-4-6</span></a></p><p><span>&#8226; Claude Opus (used to build the harness): </span><a href="https://www.anthropic.com/claude/opus"><span>https://www.anthropic.com/claude/opus</span></a></p><p><span>&#8226; GPT-5.5 (Codex, used to build the harness): </span><a href="https://openai.com/index/introducing-gpt-5-5/"><span>https://openai.com/index/introducing-gpt-5-5/</span></a></p><p><span>&#8226; Ink (terminal UI library for Node.js): </span><a href="https://github.com/vadimdemedes/ink"><span>https://github.com/vadimdemedes/ink</span></a></p><p><span>&#8226; Sentry (error monitoring): </span><a href="https://sentry.io/"><span>https://sentry.io/</span></a></p><p><span>&#8226; Linear (project management): </span><a href="https://linear.app/"><span>https://linear.app/</span></a></p><p><span>&#8226; GitHub: </span><a href="https://github.com/"><span>https://github.com/</span></a></p><p><span>&#8226; Vercel: </span><a href="https://vercel.com/"><span>https://vercel.com/</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[🎙️ How I AI: Sonnet 5 review & How to run autonomous coding agents from your phone ]]></title><description><![CDATA[Your weekly listens from How I AI, part of the Lenny's Podcast Network]]></description><link>https://www.lennysnewsletter.com/p/how-i-ai-sonnet-5-review-and-how</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-i-ai-sonnet-5-review-and-how</guid><dc:creator><![CDATA[Lenny Rachitsky]]></dc:creator><pubDate>Mon, 06 Jul 2026 15:01:17 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7123a1d3-40ca-43a5-a554-4116317d86f2_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gWeJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" width="1456" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:76503,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/177292431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h3><span>Sonnet 5 review: I ran 64 generations to find out if it&#8217;s worth it</span></h3><div id="youtube2-yJ-1LB2hF-Q" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;yJ-1LB2hF-Q&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/yJ-1LB2hF-Q?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/yJ-1LB2hF-Q">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/6uNXadMo6yQlD1GejPSXpy">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/sonnet-5-review-i-ran-64-generations-to-find-out-if/id1809663079?i=1000774925171">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jFyU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jFyU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!jFyU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!jFyU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!jFyU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jFyU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:25996,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/204955196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jFyU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!jFyU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!jFyU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!jFyU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc4f123dc-9a25-447c-8b1d-5738635d40f7_1600x114.png 1456w" sizes="100vw"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><ul><li><p><strong><a href="https://runwayml.com/howIAI">Runway</a></strong>&#8212;The creative AI platform for images, video, and more</p></li><li><p><strong><a href="https://hyperagent.com/howiai">Hyperagent</a></strong>&#8212;Deploy fleets of agents that handle real work</p></li></ul></blockquote><p>Claire puts Anthropic&#8217;s new Sonnet 5 through a real benchmark. She builds the How I AI Bench live using Claude Code, then blind-tests Sonnet 5 against Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro across PRDs, prototypes, agentic tasks, and agent personality. She breaks down what won, what failed, and how builders can create their own repeatable benchmark before trusting the next model release.</p><h4>Biggest takeaways:</h4><ol><li><p><strong>Sonnet 5 is priced closer to previous Sonnet models than to Opus, but it doesn&#8217;t automatically replace either one.</strong> At $2 per million input tokens and $10 per million output tokens through the end of summer, Sonnet 5 sits in an interesting middle band. In Claire&#8217;s benchmark, it finished near the bottom of her personal preference ranking, which means the cost argument only holds if the quality argument also holds for your specific use case.</p></li><li><p><strong>One-off vibe checks feel useful, but they&#8217;re not repeatable&#8212;and repeatability is what makes a benchmark actually matter.</strong> Claire has tested GPT-5.5, open-weight models like GLM-5.2, and others this way before, but she could never compare results across time. The How I AI Bench fixes that by using frozen inputs, a fixed rubric, and the same tasks every time a new model comes out.</p></li><li><p><strong>Claude Code can read old session history and use it to generate benchmark ideas tailored to a person&#8217;s actual work.</strong> Claire gave Claude Code a simple prompt asking it to brainstorm eval tasks based on what they&#8217;d worked on together, and it pulled from stored sessions on her desktop. Builders can do the same with Codex. That context is sitting there unused for most people.</p></li><li><p><strong>Building an HTML scoring page to rate outputs based on gut feel and export JSON takes maybe 45 minutes with Claude Code, and Claire argues it&#8217;s worth every minute.</strong> She scored 64 generations across five models by hand, gave each one a 1-to-5 gut score, added loose notes, and found that the human signal turned out to be the most useful part of the whole benchmark.</p></li><li><p><strong>LLM-as-judge evals are too generous and cluster toward the middle of the scale.</strong> Claire had both GPT-5.5 and Opus 4.8 judge the outputs, and neither was spiky enough. They missed things she flagged immediately on a visual pass, like broken prototypes and ignored wireframe constraints. Models can&#8217;t yet see what the human eye catches in the first screenshot.</p></li><li><p><strong>Claire&#8217;s taste and the automated benchmark disagreed almost completely, and she thinks her taste was at least partly right.</strong> The LLM judges ranked Gemini 3 Pro highest and Sonnet 4.6 lowest. Claire&#8217;s ranking was almost exactly the opposite. When she ran a 70/30 Claire-to-LLM weighted index, Sonnet 4.6 jumped to first. That divergence tells her the eval rubric needs to encode more of what she actually cares about before she can trust the automated scores.</p></li><li><p><strong>Sonnet 4.6 is still Claire&#8217;s choice for daily agent work because of its personality, not its benchmark scores.</strong> She pays for API credits to run her OpenClaw on Sonnet 4.6 specifically because she likes how it talks to her. No other model in this test matched it on the voice eval, which asked things like &#8220;ugh, deploys are red again&#8221; and waited to see how the model responded.</p></li><li><p><strong>For builders, Claire recommends GPT-5.5 for PRDs, Sonnet 4.6 for prototypes and chitchat, and Opus 4.8 or Sonnet 5 for codebase navigation.</strong> Those are the task-by-task recommendations that came out of the Claire-weighted index. Complex, dense UI work is where Opus 4.8 still earns its price premium; for everything simpler, Sonnet 4.6 holds up.</p></li><li><p><strong>The How I AI Bench is version one, and a lot of it needs to get sharper.</strong> The agentic bug-hunting task turned out to be too easy: every model aced it, which means it can&#8217;t differentiate between good and great. Claire plans to retire that task, encode more of her taste into the rubric, and keep running the benchmark blind every time a new model drops. The goal is to make this a benchmark the labs actually care about.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>Building a Custom Benchmark for Sonnet 5, and Why the Results Surprised Me: <a href="https://www.chatprd.ai/how-i-ai/sonnet-5-review-and-custom-benchmark">https://www.chatprd.ai/how-i-ai/sonnet-5-review-and-custom-benchmark</a></p><p><strong>&#8627; </strong>How to Conduct a Blind &#8216;Vibe Check&#8217; to Evaluate AI Model Quality: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-conduct-a-blind-vibe-check-to-evaluate-ai-model-quality">https://www.chatprd.ai/how-i-ai/workflows/how-to-conduct-a-blind-vibe-check-to-evaluate-ai-model-quality</a><strong> </strong></p><p><strong>&#8627; </strong>How to Build a Custom AI Model Benchmark Using Claude Code: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-custom-ai-model-benchmark-using-claude-code">https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-custom-ai-model-benchmark-using-claude-code</a></p><p><strong>&#8627; </strong>How to Create a Weighted Index for AI Model Benchmark Results: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-create-a-weighted-index-for-ai-model-benchmark-results">https://www.chatprd.ai/how-i-ai/workflows/how-to-create-a-weighted-index-for-ai-model-benchmark-results</a></p></div><h3><span>How I run autonomous coding agents from my phone with OpenAI Symphony + Linear | Alessio Fanelli (Kernel Labs)</span></h3><div id="youtube2-KtmaWUVdnx4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;KtmaWUVdnx4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/KtmaWUVdnx4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/KtmaWUVdnx4">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/6BinWqjJkC9Slq53nVNqeC">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/how-i-run-autonomous-coding-agents-from-my-phone-with/id1809663079?i=1000775636075">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FVB8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FVB8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!FVB8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!FVB8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!FVB8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FVB8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:19337,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/204955196?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FVB8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!FVB8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!FVB8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!FVB8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d8ab9c9-e665-4400-bc12-525cc99ce8d0_1600x114.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><ul><li><p><strong><a href="https://firecrawl.dev/?utm_source=newsletter&amp;utm_medium=partner&amp;utm_campaign=how_i_ai"><span>Firecrawl</span></a></strong><span>&#8212;Power AI agents with clean web data</span></p></li><li><p><strong><a href="https://atlassian.com/howiai"><span>Jira Product Discovery</span></a></strong><span>&#8212;Prioritize with insights, build with confidence</span></p></li></ul></blockquote><p><strong>Alessio Fanelli</strong>, the founder of Kernel Labs and co-host of Latent Space podcast, shows Claire how he manages autonomous coding agents from his phone using OpenAI Symphony, Linear, and a cloud VPS. He walks through the shift from &#8220;agent prompter&#8221; to &#8220;agent manager,&#8221; explains why Linear works as a state machine for async agent work, and shares what he&#8217;s learned from tracking token costs, purging skills files, and giving agents better senses. He also demos a very different use case: using Codex with browser access to hunt for underpriced Pok&#233;mon cards on eBay.</p><h4>Biggest takeaways:</h4><ol><li><p><strong><span>The shift from &#8220;agent prompter&#8221; to &#8220;agent manager&#8221; is the unlock most people are still missing.</span></strong><span> Alessio described how early agentic workflows felt fine until the second or third intervention, when the friction of local runtimes and clunky interfaces killed momentum. Moving to a cloud VPS with multiple communication channels (Linear, shell, mobile) made real async management possible.</span></p></li><li><p><strong><span>Symphony isn&#8217;t magic; it&#8217;s just a very opinionated Markdown spec that tells the model how to behave.</span></strong><span> Claire made this point explicitly during the episode, and it&#8217;s the most important corrective to the &#8220;complex agent orchestration&#8221; framing that intimidates people. The whole framework is a Markdown file, and the models are good enough now to lock to it faithfully.</span></p></li><li><p><strong><span>Token cost tracking per task is the primitive that most agent setups don&#8217;t have, and it should be table stakes.</span></strong><span> Alessio showed tasks ranging from 15 million to 221 million tokens, and the 221-million-token job (making an app deployable on Vercel) made complete sense in hindsight. Without that ledger, you have no feedback loop for improving your specs or your tooling.</span></p></li><li><p><strong><span>Purge your skills files every few months or they become a liability.</span></strong><span> Models have a strong tendency to add instructions rather than replace them, so a skills file that grows over time eventually contradicts itself. Alessio&#8217;s advice is to keep files short, tight, and explicit about what the agent needs to ask for, not exhaustive lists of every possible behavior.</span></p></li><li><p><strong><span>AI&#8217;s biggest unlocked opportunity is businesses built on heterogeneous data.</span></strong><span> The category Alessio described&#8212;things like trading cards, vintage clothing, and fish inventory&#8212;has always been impossible to scale because the data is inconsistent, visual, and contextual. LLMs are the first technology malleable enough to handle that messiness without extensive preprocessing.</span></p></li><li><p><strong><span>Giving agents better senses (screenshots, visual diffs, video) extends autonomous runs dramatically.</span></strong><span> Kernel Labs built Glimpse, a Playwright extension for coding agents, specifically because the bottleneck wasn&#8217;t orchestration but rather agents hitting ambiguity in the UI and calling for help. Better tooling at the perception layer keeps the run going.</span></p></li><li><p><strong><span>Context offloading is an underrated AI use case, and it&#8217;s worth building deliberately.</span></strong><span> Alessio&#8217;s email monitoring setup gave him the certainty that nothing important was slipping through, which removed a low-grade background anxiety. The same logic applies to personal finance, inventory, and any domain where staying on top of information is taking cognitive bandwidth you&#8217;d rather spend elsewhere.</span></p></li><li><p><strong><span>The Pok&#233;mon card demo is the clearest proof that AI is compressing the information advantage that scale used to provide.</span></strong><span> Finding underpriced PSA-graded cards at $10,000-plus price points was previously a function of having more time, more people, and more domain expertise than competitors. Codex with browser access and a custom pricing skill collapses that advantage to a single well-written prompt.</span></p></li><li><p><strong><span>Small businesses are the most interesting AI story, and they&#8217;re being underreported.</span></strong><span> Alessio&#8217;s observation from Japan, that small two- and three-person operations are running happily and profitably, points to a different kind of AI opportunity than the enterprise narrative suggests. The leverage AI gives a one-person operation is asymmetric in a way that bigger organizations can&#8217;t replicate at the same cost.</span></p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How Alessio Fanelli uses Open AI Symphony for Autonomous Coding and Pok&#233;mon Card Trading Workflows: <a href="https://www.chatprd.ai/how-i-ai/alessio-fanelli-uses-open-ai-symphony-for-autonomous-coding-and-pokemon-card-trading">https://www.chatprd.ai/how-i-ai/alessio-fanelli-uses-open-ai-symphony-for-autonomous-coding-and-pokemon-card-trading</a></p><p><strong>&#8627; </strong>Build an AI Agent to Find Underpriced Pok&#233;mon Cards for Arbitrage: <a href="https://www.chatprd.ai/how-i-ai/workflows/build-an-ai-agent-to-find-underpriced-pok-mon-cards-for-arbitrage">https://www.chatprd.ai/how-i-ai/workflows/build-an-ai-agent-to-find-underpriced-pok-mon-cards-for-arbitrage</a></p><p><strong>&#8627; </strong>Automate Software Development with an AI Agent Manager using OpenAI Symphony and Linear: <a href="https://www.chatprd.ai/how-i-ai/workflows/automate-software-development-with-an-ai-agent-manager-using-openai-symphony-and-linear">https://www.chatprd.ai/how-i-ai/workflows/automate-software-development-with-an-ai-agent-manager-using-openai-symphony-and-linear</a></p></div><div><hr></div><p>If you&#8217;re enjoying these episodes, reply and let me know what you&#8217;d love to learn more about: AI workflows, hiring, growth, product strategy&#8212;anything.</p><p>Catch you next week,<br>Lenny</p><p><em>P.S. Want every new episode delivered the moment it drops? Hit &#8220;Follow&#8221; on your favorite podcast app.</em></p>]]></content:encoded></item><item><title><![CDATA[How I run autonomous coding agents from my phone with OpenAI Symphony + Linear | Alessio Fanelli (Kernel Labs)]]></title><description><![CDATA[Watch now | &#127897;&#65039; Alessio Fanelli shows his Symphony + Linear setup for running parallel coding agents from his phone, then demos Codex hunting underpriced Pok&#233;mon cards in real time]]></description><link>https://www.lennysnewsletter.com/p/how-i-run-autonomous-coding-agents</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-i-run-autonomous-coding-agents</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Mon, 06 Jul 2026 12:03:37 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/204741760/fa419fe415be23d13043d845b6036ad8.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-KtmaWUVdnx4" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;KtmaWUVdnx4&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/KtmaWUVdnx4?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong><span>Alessio Fanelli</span></strong><span>, founder of Kernel Labs and co-host of Latent Space podcast, walks us through two very different AI workflows: (1) a fully autonomous coding setup using OpenAI Symphony + Linear, where Linear acts as a state machine and Symphony manages agents through the whole dev lifecycle with zero babysitting; (2) Codex with browser access searching eBay for underpriced Pok&#233;mon cards&#8212;autonomously browsing, extracting PSA certificate numbers, and flagging deals on $10K&#8211;$20K cards for his San Carlos card shop, Merlin Games.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/KtmaWUVdnx4">YouTube</a>, <a href="https://open.spotify.com/episode/6BinWqjJkC9Slq53nVNqeC">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/how-i-run-autonomous-coding-agents-from-my-phone-with/id1809663079?i=1000775636075">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>Why &#8220;agent manager&#8221; is a better mental model than &#8220;agent prompter&#8221;</span></p></li><li><p><span>Why local Mac Minis don&#8217;t scale, and what a cloud VPS unlocks</span></p></li><li><p><span>How to wire Symphony and Linear together as an agent state machine</span></p></li><li><p><span>How to track token costs per task (and what 221 million tokens buys you)</span></p></li><li><p><span>What Glimpse does, and why better agent senses extend autonomous runs</span></p></li><li><p><span>Why your CLAUDE.md probably needs a full purge, not more instructions</span></p></li><li><p><span>How Codex scouts underpriced $10K Pok&#233;mon cards on eBay at scale</span></p></li><li><p><span>The new category of small business that AI just made possible</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oX79!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oX79!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!oX79!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!oX79!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!oX79!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oX79!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:19337,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/204741760?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oX79!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!oX79!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!oX79!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!oX79!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc024cc9a-84f1-4564-81bd-fb5ff998259a_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://firecrawl.dev/?utm_source=newsletter&amp;utm_medium=partner&amp;utm_campaign=how_i_ai"><span>Firecrawl</span></a></strong><span>&#8212;Power AI agents with clean web data</span></p><p><strong><a href="https://atlassian.com/howiai"><span>Jira Product Discovery</span></a></strong><span>&#8212;Prioritize with insights, build with confidence</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4">00:00</a>) Intro</p><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4&amp;t=144s">02:24</a>) Prompter vs. agent manager</p><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4&amp;t=271s">04:31</a>) Live demo: Symphony + Linear</p><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4&amp;t=571s">09:31</a>) Setting up Symphony</p><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4&amp;t=855s">14:15</a>) Purging your skills files</p><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4&amp;t=1086s">18:06</a>) The benefits of this system</p><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4&amp;t=1150s">19:10</a>) Demo: Using Codex to hunt for Pok&#233;mon cards</p><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4&amp;t=1457s">24:17</a>) The benefit of AI for small businesses</p><p>(<a href="https://www.youtube.com/watch?v=KtmaWUVdnx4&amp;t=1703s">28:23</a>) Lightning round</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; OpenAI Codex: </span><a href="https://openai.com/codex"><span>https://openai.com/codex</span></a></p><p><span>&#8226; OpenAI Symphony (open-source framework): </span><a href="https://github.com/openai/symphony"><span>https://github.com/openai/symphony</span></a></p><p><span>&#8226; Linear (project management/agent state machine): </span><a href="https://linear.app"><span>https://linear.app</span></a></p><p><span>&#8226; PSA (Professional Sports Authenticator) grading: </span><a href="https://www.psacard.com"><span>https://www.psacard.com</span></a></p><p><span>&#8226; TCGplayer (card pricing): </span><a href="https://www.tcgplayer.com"><span>https://www.tcgplayer.com</span></a></p><p><span>&#8226; eBay (used for card price scouting): </span><a href="https://www.ebay.com"><span>https://www.ebay.com</span></a></p><p></p><h3>Other references:</h3><p><span>&#8226; Meta Ray-Ban glasses: </span><a href="https://www.ray-ban.com/usa/ray-ban-meta-smart-glasses"><span>https://www.ray-ban.com/usa/ray-ban-meta-smart-glasses</span></a></p><p><span>&#8226; </span><em><span>The Monk and the Riddle</span></em><span> by Randy Komisar: </span><a href="https://www.amazon.com/Monk-Riddle-Creating-Making-Living/dp/1578516447/ref=sr_1_1">https://www.amazon.com/Monk-Riddle-Creating-Making-Living/dp/1578516447/ref=sr_1_1</a></p><p><span>&#8226; </span><em><span>The Divine Comedy</span></em><span> by Dante Alighieri: </span><a href="https://www.amazon.com/dp/0451208633">https://www.amazon.com/dp/0451208633</a></p><p><span>&#8226; AS Roma (football club Alessio and Claire are both fans of): </span><a href="https://www.asroma.com/en"><span>https://www.asroma.com/en</span></a></p><p></p><h3>Where to find <span>Alessio </span>Fanelli<span>:</span></h3><p><span>X: </span><a href="https://x.com/FanaHOVA"><span>https://x.com/FanaHOVA</span></a></p><p><span>Latent Space podcast: </span><a href="https://www.latent.space/"><span>https://www.latent.space/</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[Sonnet 5 review: I ran 64 generations to find out if it's worth it]]></title><description><![CDATA[Watch now | &#127897; I built the How I AI Bench live using Claude Code, ran 5 frontier models through 64 blind prototype generations, PRDs, and agent voice tests, and the results surprised even me]]></description><link>https://www.lennysnewsletter.com/p/sonnet-5-review-i-ran-64-generations</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/sonnet-5-review-i-ran-64-generations</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Tue, 30 Jun 2026 23:22:23 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/204353515/928109049545fa1bddb70bcf87affe5e.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-yJ-1LB2hF-Q" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;yJ-1LB2hF-Q&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/yJ-1LB2hF-Q?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><span>I&#8217;ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn&#8217;t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/yJ-1LB2hF-Q">YouTube</a>, <a href="https://open.spotify.com/episode/6uNXadMo6yQlD1GejPSXpy">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/sonnet-5-review-i-ran-64-generations-to-find-out-if/id1809663079?i=1000774925171">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>What Anthropic claims Sonnet 5 improves over Sonnet 4.6, and where the benchmark data actually backs that up</span></p></li><li><p><span>How I built the How I AI Bench in under 45 minutes using Claude Code, starting from my own stored session history</span></p></li><li><p><span>Why I combined human vibe scoring (70%) with LLM as judge scoring (30%) instead of trusting either alone</span></p></li><li><p><span>How to set up a local HTML scoring page so you can rate AI outputs on gut feel and export those scores as JSON</span></p></li><li><p><span>Which model I recommend for PRDs, which for complex prototypes, and which for chatting with an agent daily</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nBz4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nBz4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!nBz4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!nBz4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!nBz4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nBz4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:25957,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/204353515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nBz4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!nBz4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!nBz4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!nBz4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff53d9bc9-15aa-47f7-942a-2d095fa18c69_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://runwayml.com/howIAI"><span>Runway</span></a></strong><span>&#8212;The creative AI platform for images, video and more</span></p><p><strong><a href="https://hyperagent.com/howiai"><span>Hyperagent</span></a></strong><span>&#8212;Deploy fleets of agents that handle real work</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q">00:00</a>) Sonnet 5 is out</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=115s">01:55</a>) What Anthropic claims</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=242s">04:02</a>) Why I&#8217;m done with one-off vibe checks</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=305s">05:05</a>) Building the How I AI Bench live with Claude Code</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=462s">07:42</a>) The scoring system</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=643s">10:43</a>) Agent voice eval</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=717s">11:57</a>) Quick recap</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=838s">13:58</a>) Results: The How I AI index leaderboard</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=1281s">21:21</a>) What I&#8217;m improving for the next run</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=1336s">22:16</a>) Generating a Claire-weighted index</p><p>(<a href="https://www.youtube.com/watch?v=yJ-1LB2hF-Q&amp;t=1433s">23:53</a>) Model-by-task recommendations</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; Claude Sonnet 5: </span><a href="https://www.anthropic.com/news/claude-sonnet-5"><span>https://www.anthropic.com/news/claude-sonnet-5</span></a></p><p><span>&#8226; Claude Opus 4.8: </span><a href="https://www.anthropic.com/news/claude-opus-4-8"><span>https://www.anthropic.com/news/claude-opus-4-8</span></a></p><p><span>&#8226; GPT-5.5 (OpenAI): </span><a href="https://openai.com/index/introducing-gpt-5-5/"><span>https://openai.com/index/introducing-gpt-5-5/</span></a></p><p><span>&#8226; Gemini 3 Pro (Google DeepMind): </span><a href="https://deepmind.google/models/gemini/pro/"><span>https://deepmind.google/models/gemini/pro/</span></a></p><p><span>&#8226; Cursor: </span><a href="https://www.cursor.com/"><span>https://www.cursor.com/</span></a></p><p></p><h3>Other references:</h3><p><span>&#8226; SWE-bench Pro (agentic coding benchmark referenced): </span><a href="https://www.swebench.com/"><span>https://www.swebench.com/</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[🎙️ How I AI: GLM-5.2 review & How Gusto built a new product line with Claude Code]]></title><description><![CDATA[Your weekly listens from How I AI, part of the Lenny's Podcast Network]]></description><link>https://www.lennysnewsletter.com/p/how-i-ai-glm-52-review-and-how-gusto</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-i-ai-glm-52-review-and-how-gusto</guid><dc:creator><![CDATA[Lenny Rachitsky]]></dc:creator><pubDate>Mon, 29 Jun 2026 15:02:35 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c3a5f605-f314-4b28-8398-cc9a0d30559f_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gWeJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" width="1456" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:76503,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/177292431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h3>GLM-5.2: why I&#8217;m replacing Opus in Claude Code with this new model</h3><div id="youtube2-ZoBfQZ5utQk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;ZoBfQZ5utQk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/ZoBfQZ5utQk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/ZoBfQZ5utQk">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/4fnfOEDbeM7X0DpYdLkJU0">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/glm-5-2-why-im-replacing-opus-in-claude-code-with-this/id1809663079?i=1000774026946">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!a_tm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a_tm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png 424w, https://substackcdn.com/image/fetch/$s_!a_tm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png 848w, https://substackcdn.com/image/fetch/$s_!a_tm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png 1272w, https://substackcdn.com/image/fetch/$s_!a_tm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a_tm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png" width="262" height="60.57381615598886" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:166,&quot;width&quot;:718,&quot;resizeWidth&quot;:262,&quot;bytes&quot;:27595,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/203730644?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!a_tm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png 424w, https://substackcdn.com/image/fetch/$s_!a_tm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png 848w, https://substackcdn.com/image/fetch/$s_!a_tm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png 1272w, https://substackcdn.com/image/fetch/$s_!a_tm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad7fc184-9e1c-4219-bcf3-f52ccfe07061_718x166.png 1456w" sizes="100vw"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><ul><li><p><strong><a href="https://mercury.com/"><span>Mercury</span></a></strong>&#8212;Radically different banking, loved by over 300K entrepreneurs</p></li></ul></blockquote><p>Claire tests GLM-5.2, the new open-weight model from Z.ai, inside her actual ChatPRD codebase. She runs it through codebase audits, UI redesigns, and a 45-minute autonomous bug-hunting task in Cursor and Claude Code, and breaks down where it surprised her, where it struggled, and why it may be good enough to replace Opus for some coding workflows.</p><h4>Biggest takeaways:</h4><ol><li><p><strong>Open-weight models are no longer a hobbyist curiosity&#8212;they are production-grade alternatives.</strong> GLM-5.2, built by Beijing-based Z.ai, benchmarks near Claude Opus 4.8 and above GPT-5.5 on SWE Bench Pro, with a million-token context window and full support for reasoning mode, function calling, structured output, and context caching. The decision is no longer about capability ceilings but, instead, about cost, control, and vendor dependency. Claire&#8217;s live testing confirmed it: this is not a toy.</p></li><li><p><strong>Self-hosting changes the vendor power dynamic in ways that matter at scale.</strong> Open-weight means the trained model weights are publicly available, letting teams run inference on their own hardware, fine-tune on proprietary data, and route around any single provider&#8217;s API terms. When frontier labs change pricing or policy, teams using open-weight models can switch inference providers without touching a line of application code. The key: you&#8217;re not locked in.</p></li><li><p><strong>Getting GLM-5.2 running in Cursor took 30 minutes, and Claire documented the undocumented part.</strong> Route your API key through Open Router, override the OpenAI base URL in Cursor&#8217;s settings to <code>openrouter.ai/api/v1/cursor</code> (the <code>/cursor</code> suffix isn&#8217;t documented anywhere), and add <code>z-ai/glm-5.2</code> as a custom model. Claude Code requires two environment variable changes and one edit to <code>claude/settings.json</code>. Total time: under an hour, once you have the exact strings.</p></li><li><p><strong>The 45-minute autonomous task revealed both the ceiling and the floor.</strong> Claire gave GLM-5.2 a single prompt inside Claude Code: pull the last 72 hours of Sentry errors and Vercel logs, then build a prioritized bug-fix plan. Over 45 minutes, it ran MCP tool calls, authenticated into external services, and produced a dark-mode engineering canvas with 20 Sentry errors, five Vercel log signals, and 14 planned fixes, including two P0s Claire hadn&#8217;t spotted through normal monitoring. The model surfaced signal-to-noise issues in their error pipeline that weren&#8217;t showing up elsewhere.</p></li><li><p><strong>It hit a wall with React, then recovered.</strong> During the long-running task, GLM-5.2 struggled with TypeScript compilation errors before eventually producing clean React output. Claire&#8217;s read: HTML and CSS generation is reliable; React under agentic, multi-step pressure is shakier. For teams whose codebase is primarily React (she estimates it covers 98% of her own use), this is the friction point to test before committing the model to critical paths.</p></li><li><p><strong>The cost math is striking: $3.36 for 6 million tokens, including the full 45-minute agentic session.</strong> A 72% cache rate helped, but even at full price, open-weight inference through Open Router sits well below Opus or GPT-5.5 rates for equivalent coding capability. For agents accumulating long context windows over extended sessions (the exact workload where frontier model costs compound fastest), open-weight alternatives offer a structurally different cost curve.</p></li><li><p><strong>Claire&#8217;s recommendation: put GLM-5.2 in rotation, not in the spotlight.</strong> She&#8217;s keeping it in Cursor for frontend and design work, and in Claude Code for long-running agentic tasks, alongside closed frontier models rather than as a replacement. The constraint she&#8217;s watching: can it handle her React-heavy workload at the same consistency she gets from Composer? If it can, the cost-and-control argument gets much harder to ignore.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>GLM 5.2: A Live Review of an Opus-Level Open-Weights Model: <a href="https://www.chatprd.ai/how-i-ai/glm-5-2-review-open-weights-model">https://www.chatprd.ai/how-i-ai/glm-5-2-review-open-weights-model</a></p><p>&#8627; How to Deploy an Autonomous AI Agent for Bug Triage and Prioritization: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-deploy-an-autonomous-ai-agent-for-bug-triage-and-prioritization">https://www.chatprd.ai/how-i-ai/workflows/how-to-deploy-an-autonomous-ai-agent-for-bug-triage-and-prioritization</a></p><p>&#8627; How to Perform an AI-Powered Codebase Audit and Architecture Visualization: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-perform-an-ai-powered-codebase-audit-and-architecture-visualization">https://www.chatprd.ai/how-i-ai/workflows/how-to-perform-an-ai-powered-codebase-audit-and-architecture-visualization</a></p><p>&#8627; How to Configure the Open-Weight GLM 5.2 Model in Cursor: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-configure-the-open-weight-glm-5-2-model-in-cursor">https://www.chatprd.ai/how-i-ai/workflows/how-to-configure-the-open-weight-glm-5-2-model-in-cursor</a></p></div><h3><span>No Figma. No Jira. No docs. How Gusto built a new product line with Claude Code | Eddie Kim (CTO)</span></h3><div id="youtube2-5FKBkUCaLa8" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;5FKBkUCaLa8&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/5FKBkUCaLa8?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/5FKBkUCaLa8">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/74S7FEy0h9TfRuUpfiNbBP">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/no-figma-no-jira-no-docs-how-gusto-built-a-new/id1809663079?i=1000774685029">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QRUh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QRUh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!QRUh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!QRUh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!QRUh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QRUh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:28391,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/203730644?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QRUh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!QRUh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!QRUh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!QRUh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb5897fb1-897a-4eb8-90c5-c539c8a630b7_1600x114.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><ul><li><p><strong><a href="https://magicpatterns.com/howiai"><span>Magic Patterns</span></a></strong><span>&#8212;Prototypes that look like your product</span></p></li><li><p><strong><a href="https://atlassian.com/howiai"><span>Jira Product Discovery</span></a></strong><span>&#8212;Prioritize with insights, build with confidence</span></p></li></ul></blockquote><p><strong>Eddie Kim</strong> is the co-founder and CTO of Gusto. In this episode, he shares how a five-person team used Claude Code, a permanent Zoom room, and almost none of the usual product process&#8212;no PM, no Figma, no Jira, no long specs&#8212;to build Gusto Cofounder from scratch in just 10 weeks.</p><h4>Biggest takeaways:</h4><ol><li><p><strong>A five-person team with no process can outship a large team with full process, if AI handles the engineering.</strong> Eddie&#8217;s product launched at Gusto&#8217;s tier-one level after 10 weeks, starting from zero code. The constraint wasn&#8217;t a liability&#8212;it was the design. When AI does the building, coordination overhead doesn&#8217;t scale the engineering; it just slows it down. The key: strip process to what the team actually needs, then let AI fill the gap.</p></li><li><p><strong>&#8220;Zero code to tier-one launch&#8221; is now a viable founding path.</strong> The team reached a production milestone at Gusto without a line of pre-existing code. This flips the assumption that early teams spend months on infrastructure before shipping anything real. With Claude Code as the primary builder, the initial sprint becomes about direction and judgment, not typing. It compresses the time between idea validation and real user contact from months to weeks.</p></li><li><p><strong>No meetings, no Jira, no text threads. It shipped anyway.</strong> The team had no standup cadence, no ticket system, no async thread to resolve blockers. What replaced all of that: shared context held inside the AI loop. When the model carries state and the team is small and aligned, human coordination overhead becomes optional.</p></li><li><p><strong>The technical stack for a production AI agent is shockingly minimal.</strong> The entire agent loop ran on Cloudflare Workers with the Vercel AI SDK. Nothing else. No proprietary orchestration layer, no third-party agent framework. Everything else was built in-house. Teams often over-architect before they&#8217;ve proven anything; Eddie&#8217;s stack is evidence that infrastructure minimalism accelerates the path to learning what the agent actually needs to do.</p></li><li><p><strong>Building agents is not as complicated as the community makes it sound.</strong> An agent is an AI SDK running somewhere in the cloud, able to look up files and call tools. That&#8217;s the full definition. The complexity people fear (state management, orchestration, reliability) is solvable with the same judgment calls any backend system requires. Eddie&#8217;s team shipped one at production quality in 10 weeks without specialist AI infrastructure experience.</p></li><li><p><strong>The &#8220;permanent Zoom&#8221; model of AI development changes how teams think about context.</strong> Claude Code running in a persistent loop means the model has continuous access to the codebase&#8217;s current state. That&#8217;s closer to having an engineer who never closes their laptop than a chat interface you query on demand. For small teams, this is the equivalent of a senior engineer who is always available, always current, and never needs onboarding after a break.</p></li><li><p><strong>The lesson for founding teams isn&#8217;t &#8220;use Claude Code.&#8221; It&#8217;s &#8220;design your process for AI as a team member.&#8221;</strong> Most early teams graft AI tools onto a human-scaled workflow: standups, tickets, PRs reviewed by three people. Eddie&#8217;s team treated the AI as a primary contributor from day one and built their coordination model around that assumption. The result: a workflow that gets faster as the AI improves, not one that merely offloads tasks to it.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How Gusto Built a New Product Line in 10 Weeks with Claude Code, No Jira, and No Docs: <a href="https://www.chatprd.ai/how-i-ai/how-gusto-built-a-new-product-line-in-10-weeks-with-claude-code-no-jira-and-no-docs">https://www.chatprd.ai/how-i-ai/how-gusto-built-a-new-product-line-in-10-weeks-with-claude-code-no-jira-and-no-docs</a></p><p>&#8627; How to Build a New AI Product in 10 Weeks Using the &#8216;No-Process&#8217; Method: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-new-ai-product-in-10-weeks-using-the-no-process-method">https://www.chatprd.ai/how-i-ai/workflows/how-to-build-a-new-ai-product-in-10-weeks-using-the-no-process-method</a></p><p>&#8627; How to Fix Bugs Using an AI-Powered Test-Driven Development (TDD) Workflow: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-fix-bugs-using-an-ai-powered-test-driven-development-tdd-workflow">https://www.chatprd.ai/how-i-ai/workflows/how-to-fix-bugs-using-an-ai-powered-test-driven-development-tdd-workflow</a></p></div><div><hr></div><p>If you&#8217;re enjoying these episodes, reply and let me know what you&#8217;d love to learn more about: AI workflows, hiring, growth, product strategy&#8212;anything.</p><p>Catch you next week,<br>Lenny</p><p><em>P.S. Want every new episode delivered the moment it drops? Hit &#8220;Follow&#8221; on your favorite podcast app.</em></p>]]></content:encoded></item><item><title><![CDATA[No Figma. No Jira. No docs. How Gusto built a new product line with Claude Code | Eddie Kim (CTO)]]></title><description><![CDATA[Watch now | &#127897;&#65039; Eddie Kim, the CTO of Gusto, shows how a 5-person team shipped a full AI product line in 10 weeks using Claude Code, a perma-Zoom, and zero documentation]]></description><link>https://www.lennysnewsletter.com/p/no-figma-no-jira-no-docs-how-gusto</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/no-figma-no-jira-no-docs-how-gusto</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Mon, 29 Jun 2026 12:03:38 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/203589481/846f1cae45d76ef325a01c8acbcd5f61.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-5FKBkUCaLa8" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;5FKBkUCaLa8&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/5FKBkUCaLa8?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong><span>Eddie Kim</span></strong><span> is the co-founder and CTO of the payroll and HR platform </span>Gusto, which<span> just crossed $1 billion in revenue and serves more than 500,000 small businesses. Recently he did something most CTOs don&#8217;t: he went back to writing code. With three other engineers and one designer, Eddie built Gusto Cofounder, a net-new AI product, from zero code to a tier-one launch in 10 weeks. He walks through how that team actually worked, why they threw out nearly every process, and how anyone can copy the approach.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/5FKBkUCaLa8">YouTube</a>, <a href="https://open.spotify.com/episode/74S7FEy0h9TfRuUpfiNbBP">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/no-figma-no-jira-no-docs-how-gusto-built-a-new/id1809663079?i=1000774685029">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>The trash-can method: how to write, review, and delete a full PR as a product decision instead of a planning doc</span></p></li><li><p><span>The two-tool agent stack behind Gusto Cofounder</span></p></li><li><p><span>The exact &#8220;perma-Zoom&#8221; setup that replaced standups, retros, and Slack threads for 10 weeks</span></p></li><li><p><span>How a designer with no engineering background hit the 94th percentile for shipping code</span></p></li><li><p><span>The eval-first workflow Eddie uses to fix real customer bugs with Claude Code</span></p></li><li><p><span>How a non-technical leader can prototype an idea to win buy-in, then carry it all the way to production-quality code</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9xGW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9xGW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!9xGW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!9xGW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!9xGW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9xGW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/af4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:28391,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/203589481?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!9xGW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!9xGW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!9xGW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!9xGW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faf4f709f-eb54-4b2e-bb13-369602a6e7ba_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://magicpatterns.com/howiai"><span>Magic Patterns</span></a></strong><span>&#8212;Prototypes that look like your product</span></p><p><strong><a href="https://atlassian.com/howiai"><span>Jira Product Discovery</span></a></strong><span>&#8212;Prioritize with insights, build with confidence</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8">00:00</a>) Intro: five people, 10 weeks</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=158s">02:38</a>) The origins of Cofounder</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=512s">08:32</a>) Inside the 10-week build process</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=770s">12:50</a>) Building with no PMs</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=878s">14:38</a>) The &#8220;trash can&#8221; method</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=1035s">17:15</a>) The stack architecture</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=1150s">19:10</a>) Shipping to production from day one</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=1323s">22:03</a>) How a designer became a top engineer</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=1745s">29:05</a>) Demo: Cofounder over text and Slack</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=1905s">31:45</a>) Demo: running a real payroll</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=2186s">36:26</a>) Live coding with evals in Claude Code</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=2379s">39:39</a>) Recap: prototype, small team, permission</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=2597s">43:17</a>) Lightning round</p><p>(<a href="https://www.youtube.com/watch?v=5FKBkUCaLa8&amp;t=2924s">48:44</a>) Where to find Eddie and Cofounder</p><p></p><h3>Tools referenced:</h3><p><span>&#8226; Gusto Cofounder (early access/waitlist): </span><a href="https://gusto.com/cofounder"><span>https://gusto.com/cofounder</span></a></p><p><span>&#8226; Claude Code (Anthropic): </span><a href="https://claude.ai/code"><span>https://claude.ai/code</span></a></p><p><span>&#8226; Cloudflare Workers: </span><a href="https://workers.cloudflare.com/"><span>https://workers.cloudflare.com/</span></a></p><p><span>&#8226; Vercel AI SDK: </span><a href="https://sdk.vercel.ai/"><span>https://sdk.vercel.ai/</span></a></p><p><span>&#8226; DX (engineering analytics): </span><a href="https://getdx.com/"><span>https://getdx.com/</span></a></p><p><span>&#8226; Wispr Flow (voice-to-text): </span><a href="https://wisprflow.ai"><span>https://wisprflow.ai</span></a></p><p><span>&#8226; OpenClaw: </span><a href="https://openclaw.ai/"><span>https://openclaw.ai/</span></a></p><p></p><h3>Other references:</h3><p><span>&#8226; Gusto (the main product, &#8220;Gusto Classic&#8221;): </span><a href="https://gusto.com"><span>https://gusto.com</span></a></p><p><span>&#8226; </span>Mindbody<span> (referenced as customer data source): </span><a href="https://www.mindbodyonline.com/"><span>https://www.mindbodyonline.com/</span></a></p><p></p><h3>Where to find <span>Eddie Kim:</span></h3><p><span>LinkedIn: </span><a href="https://www.linkedin.com/in/edawerd/"><span>https://www.linkedin.com/in/edawerd/</span></a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[GLM 5.2: why I’m replacing Opus in Claude Code with this new model]]></title><description><![CDATA[Watch now | &#127897;&#65039;I ran GLM-5.2, the open-weight model from Z.AI, through codebase audits, UI redesigns, and a 45-minute autonomous bug-hunting task in Cursor and Claude Code, and it cost me $3.36]]></description><link>https://www.lennysnewsletter.com/p/glm-52-why-im-replacing-opus-in-claude</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/glm-52-why-im-replacing-opus-in-claude</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Wed, 24 Jun 2026 12:03:52 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/203168823/b822175a80b54f3ee429cb674704290e.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-ZoBfQZ5utQk" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;ZoBfQZ5utQk&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/ZoBfQZ5utQk?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><span>I put GLM 5.2, the open-weight coding model from Z.AI, through four real tasks inside my actual codebase: a codebase architecture audit, a UI redesign, and a 45-minute autonomous bug-hunting session pulling from Sentry and Vercel logs. Total cost: $3.36 for roughly 6 million tokens, a prioritized bug-fix dashboard I&#8217;m actually shipping from, and a landing page redesign that matched Chat PRD&#8217;s design system on the first try.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/ZoBfQZ5utQk">YouTube</a>, <a href="https://open.spotify.com/episode/4fnfOEDbeM7X0DpYdLkJU0">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/glm-5-2-why-im-replacing-opus-in-claude-code-with-this/id1809663079?i=1000774026946">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>What &#8220;open-weight&#8221; actually means and why it matters for cost and vendor independence</span></p></li><li><p><span>How to connect GLM 5.2 to Cursor and Claude Code</span></p></li><li><p><span>How it performs on codebase exploration and autonomous architecture summarization in a real production Next.js app</span></p></li><li><p><span>Whether GLM 5.2 can match an existing design system</span></p></li><li><p><span>How the model handles a 45-minute long-running autonomous task</span></p></li><li><p><span>Where GLM 5.2 stumbled </span></p></li><li><p><span>The actual cost breakdown</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hRK9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hRK9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png 424w, https://substackcdn.com/image/fetch/$s_!hRK9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png 848w, https://substackcdn.com/image/fetch/$s_!hRK9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png 1272w, https://substackcdn.com/image/fetch/$s_!hRK9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hRK9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png" width="316" height="73.05849582172702" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:166,&quot;width&quot;:718,&quot;resizeWidth&quot;:316,&quot;bytes&quot;:27595,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/203168823?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hRK9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png 424w, https://substackcdn.com/image/fetch/$s_!hRK9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png 848w, https://substackcdn.com/image/fetch/$s_!hRK9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png 1272w, https://substackcdn.com/image/fetch/$s_!hRK9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b2d9882-6d21-4e4e-859d-df89a8652a1f_718x166.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://mercury.com/"><span>Mercury</span></a></strong><span>&#8212;Radically different banking loved by over 300K entrepreneurs</span></p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk">00:00</a>) What open-weight models are and why GLM 5.2 is worth testing</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=98s">01:38</a>) GLM 5.2 model overview</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=242s">04:02</a>) Capabilities and benchmark results</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=362s">06:02</a>) How to set up GLM 5.2 in Cursor</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=517s">08:37</a>) How to set up GLM 5.2 in Claude Code</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=664s">11:04</a>) Live test 1: codebase exploration and architecture audit on ChatPRD</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=763s">12:43</a>) Live test 2: generating an HTML architecture and roadmap page</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=997s">16:37</a>) Live test 3: redesigning the How I AI landing page in Cursor</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=1257s">20:57</a>) Live test 4: 45-minute autonomous task, pulling Sentry errors and Vercel logs</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=1355s">22:35</a>) Where it struggled</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=1429s">23:49</a>) My verdict on the output</p><p>(<a href="https://www.youtube.com/watch?v=ZoBfQZ5utQk&amp;t=1523s">25:23</a>) Cost breakdown</p><p></p><h3>Tools referenced:</h3><ul><li><p><span>z.ai: </span><a href="https://z.ai"><span>https://z.ai</span></a></p></li><li><p><span>GLM 5.2: </span><a href="https://z.ai/blog/glm-5.2"><span>https://z.ai/blog/glm-5.2</span></a></p></li><li><p><span>OpenRouter: </span><a href="https://openrouter.ai"><span>https://openrouter.ai</span></a></p></li><li><p><span>Cursor: </span><a href="https://cursor.com"><span>https://cursor.com</span></a></p></li><li><p><span>Claude Code: </span><a href="https://docs.anthropic.com/en/docs/claude-code"><span>https://docs.anthropic.com/en/docs/claude-code</span></a></p></li><li><p><span>Sentry: </span><a href="https://sentry.io"><span>https://sentry.io</span></a></p></li><li><p><span>Vercel: </span><a href="https://vercel.com"><span>https://vercel.com</span></a></p><p></p></li></ul><h3>Other references:</h3><ul><li><p><span>SWE-Bench Pro leaderboard (coding benchmark scores referenced in episode): </span><a href="https://www.swebench.com"><span>https://www.swebench.com</span></a></p></li><li><p><span>Frontier Suite and Post-Train Bench (additional benchmarks cited): </span><a href="https://scale.com/leaderboard"><span>https://scale.com/leaderboard</span></a></p></li><li><p><span>Use Claude Code with OpenRouter: </span><a href="https://openrouter.ai/docs/cookbook/coding-agents/claude-code-integration"><span>https://openrouter.ai/docs/cookbook/coding-agents/claude-code-integration</span></a></p></li></ul><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[🎙️ How I AI: How to write AI agent loops in Claude Code and Codex + How Claude Mythos found a 15-year-old bug in Mozilla Firefox]]></title><description><![CDATA[Your weekly listens from How I AI, part of the Lenny&#8217;s Podcast Network]]></description><link>https://www.lennysnewsletter.com/p/how-i-ai-how-to-write-ai-agent-loops</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-i-ai-how-to-write-ai-agent-loops</guid><dc:creator><![CDATA[Lenny Rachitsky]]></dc:creator><pubDate>Mon, 22 Jun 2026 15:02:37 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a0360823-01e5-4060-a794-d51675b9befb_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gWeJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" width="1456" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:76503,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/177292431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h3><span>How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex</span></h3><div id="youtube2-JoXbk2fm7jM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;JoXbk2fm7jM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/JoXbk2fm7jM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/JoXbk2fm7jM">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/43kcUzpSfJExbkIotdUbBp">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/how-to-design-ai-agent-loops-schedules-goals-and/id1809663079?i=1000773109920">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hnh5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hnh5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!hnh5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!hnh5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!hnh5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hnh5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:20114,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/202479892?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hnh5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!hnh5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!hnh5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!hnh5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa30d9a22-40aa-4d45-a24b-526a1d2989cc_1600x114.png 1456w" sizes="100vw"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><ul><li><p><strong><a href="https://workos.com/?utm_source=lennys_howiai&amp;utm_medium=podcast&amp;utm_campaign=q22025"><span data-color="rgb(17, 85, 204)" style="color: rgb(17, 85, 204);">WorkOS</span></a></strong><span>&#8212;Make your app enterprise-ready today</span></p></li><li><p><strong><a href="https://runwayml.com/howIAI"><span data-color="rgb(17, 85, 204)" style="color: rgb(17, 85, 204);">Runway</span></a></strong><span>&#8212;The creative AI platform for images, video and more</span></p></li></ul></blockquote><p>In this hands-on tutorial, Claire explains the difference between heartbeats, crons, hooks, and goal-based loops, then builds real automations in Claude Code and Codex, including a daily PR-review loop and a weekly skills loop that spawns its own subagents. If you&#8217;ve heard &#8220;loop engineering&#8221; and wondered what it actually means, this is the beginner-friendly breakdown.</p><h4>Biggest takeaways:</h4><ol><li><p><strong><span>A loop is just a prompt that fires itself, nothing more exotic than that.</span></strong><span> The reason &#8220;loops&#8221; sound intimidating is that the hype cycle turned a basic automation concept into something mystical. Heartbeats, crons, and webhooks have been around forever. What&#8217;s new is pointing them at an AI agent instead of a batch job.</span></p></li><li><p><strong><span>Goals are the most powerful loop type, and the one most people get wrong.</span></strong><span> A goal loop sets an outcome and runs an agent against it until the outcome is validated or the agent gets stuck. It doesn&#8217;t stop on a timer; it stops when the work is actually done. Fuzzy success criteria means the agent loops forever, burning tokens, so my advice is to let Codex write its own goals, using OpenAI&#8217;s goal-writing guide as a starting point.</span></p></li><li><p><strong><span>Think about loops the way you think about onboarding an employee.</span></strong><span> Define the job: what they check, how often, what output you want, and who to contact when something&#8217;s wrong. &#8220;Every Friday at 10 a.m., review all merged PRs and identify skills our agents are missing&#8221; is a job description. It&#8217;s also a loop prompt.</span></p></li><li><p><strong><span>Your agent can have its own agents. This is where loops get truly powerful.</span></strong><span> The PR-review loop Claire built in Claude Code doesn&#8217;t just check PR status; it spins off dedicated subagents to babysit individual PRs until all merge checks are green. The skills loop in Codex identifies gaps and immediately spawns subagents to validate each new skill using a goal loop.</span></p></li><li><p><strong><span>Loops get expensive if you don&#8217;t write them carefully.</span></strong><span> If the success criteria is vague or the validation threshold is too thin, the agent will keep running and keep charging without meaningful progress. Monitor both cost and output quality from day one.</span></p></li><li><p><strong><span>The morning briefing in Claude Cowork is a perfect loop starter.</span></strong><span> A scheduled task that fires every morning, checks your calendar and email, and sends a summary to Slack is already a fully functional loop. No code required. From there, scaling up to PR reviews or skills identification in Claude Code or Codex is a natural next step.</span></p></li><li><p><strong><span>The power move is loops that generate their own subagent loops.</span></strong><span> In the Codex demo, Claire&#8217;s weekly automation spawned two named subagents that each ran their own goal loops to validate skills in real time. The ceiling on loop-based automation is basically &#8220;how well can you define the job?&#8221; not &#8220;how complex is the engineering?&#8221;</span></p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How I AI: Designing AI Agent Loops in Claude Code and Codex: <a href="https://www.chatprd.ai/how-i-ai/how-i-ai-designing-ai-agent-loops-in-claude-code-and-codex">https://www.chatprd.ai/how-i-ai/how-i-ai-designing-ai-agent-loops-in-claude-code-and-codex</a><br>&#8627; Build a Self-Improving AI to Generate Agent Skills in Codex: <a href="https://www.chatprd.ai/how-i-ai/workflows/build-a-self-improving-ai-to-generate-agent-skills-in-codex">https://www.chatprd.ai/how-i-ai/workflows/build-a-self-improving-ai-to-generate-agent-skills-in-codex</a><br>&#8627; Automate Daily Pull Request Reviews with a Claude Code Agent: <a href="https://www.chatprd.ai/how-i-ai/workflows/automate-daily-pull-request-reviews-with-a-claude-code-agent">https://www.chatprd.ai/how-i-ai/workflows/automate-daily-pull-request-reviews-with-a-claude-code-agent</a></p></div><h3><span>How Claude Mythos found a 15-year-old bug in Mozilla Firefox | Brian Grinstead</span></h3><div id="youtube2-Idjt53tTv2U" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;Idjt53tTv2U&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/Idjt53tTv2U?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/Idjt53tTv2U">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/4MNilflycEO8KNZRM3gtWc">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/how-claude-mythos-found-a-15-year-old-bug-in/id1809663079?i=1000773719910">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7lYV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7lYV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!7lYV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!7lYV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!7lYV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7lYV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:30703,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/202479892?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7lYV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!7lYV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!7lYV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!7lYV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5ad75402-a7c6-49f9-a482-9333373fde5e_1600x114.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><ul><li><p><strong><a href="https://workos.com/?utm_source=lennys_howiai&amp;utm_medium=podcast&amp;utm_campaign=q22025"><span data-color="rgb(17, 85, 204)" style="color: rgb(17, 85, 204);">WorkOS</span></a></strong><span>&#8212;Make your app enterprise-ready today</span></p></li><li><p><strong><a href="https://www.metaview.ai/home/how-i-ai"><span data-color="rgb(17, 85, 204)" style="color: rgb(17, 85, 204);">Metaview</span></a></strong><span>&#8212;The agentic recruiting platform for winning teams</span></p></li></ul></blockquote><p><strong>Brian Grinstead</strong>, distinguished engineer at Mozilla, breaks down how his team used AI agents to ship 423 Firefox security fixes in one month. He explains why the real unlock wasn&#8217;t just a better model, but the custom harness around it: scoring files, running goal loops, verifying bugs with subagents, and keeping humans in the review process. It&#8217;s a tactical look at how to point agents at a massive codebase and get fixes you can actually ship.</p><h4>Biggest takeaways:</h4><ol><li><p><strong><span>The Firefox security bug spike wasn&#8217;t just about the model; it was the harness too. </span></strong><span>While everyone focused on Mythos, the real story is that Firefox built a custom harness that gives AI agents the right tools to find, verify, and fix bugs. Brian says this is simpler than it looks: &#8220;It&#8217;s actually a reasonably simple wrapper around it. You just need to give it access to the right tools for the job.&#8221;</span></p></li><li><p><strong><span>Agents are relentless in a way humans can&#8217;t be. </span></strong><span>Agents will try 14, 15, 20 different approaches to trigger a bug without getting tired or losing focus. Brian found bugs that required the agent to try 14 times before succeeding. As Brian notes, &#8220;Cognitive energy declines over time in a way that agents don&#8217;t.&#8221;</span></p></li><li><p><strong><span>The verification loop is what eliminates false positives. </span></strong><span>Firefox uses a two-stage verification process: first, the agent must trigger an actual crash in their fuzzing build (a crystal-clear signal), and second, a verifier subagent checks that the bug report makes sense and doesn&#8217;t involve test-only configurations. By the time a bug reaches human engineers, there are almost no false positives.</span></p></li><li><p><strong><span>Agents get laser-focused on the specific task and miss the bigger pictu</span></strong><span>re. When the patching agent fixed a bug, it would often patch just the one vulnerable location. Human engineers would then look at the fix and say, &#8220;This is right, but we should also check three other similar places in the codebase.&#8221;</span></p></li><li><p><strong><span>Prioritization is essential when you have millions of lines of code.</span></strong><span> Firefox built a simple LLM judge that scores each file on two dimensions: likelihood of a memory safety issue, and ease of access from a webpage. Brian says this is &#8220;very, very simple&#8221; and anyone can replicate it.</span></p></li><li><p><strong><span>The harness can be built in an afternoon using vendor SDKs. </span></strong><span>Firefox started with Claude&#8217;s agent SDK, which is essentially a wrapper around Claude Code CLI that streams JSON and provides programmatic hooks. Brian&#8217;s advice: use the vendor-provided harnesses (Claude agent SDK, OpenAI agent SDK) rather than third-party frameworks, because the models are likely post-trained to work best with their own infrastructure.</span></p></li><li><p><strong><span>You should run multiple models and harnesses for security work.</span></strong><span> Because attackers will use whatever model and technique finds bugs, defenders need to scan with multiple approaches. Different models and harnesses spike on different strengths and will identify different vulnerabilities.</span></p></li><li><p><strong><span>This approach works for more than security&#8212;performance, tech debt, and UX are all viable targets.</span></strong><span> The same pattern applies: score and prioritize areas of your codebase, give the agent a constrained goal with verification criteria, and plug the results into your existing pipeline. Brian says they&#8217;re doing active work on performance optimization using the same harness structure.</span></p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p>How Mozilla Fixed 500 Security Bugs with Claude Mythos: <a href="https://www.chatprd.ai/how-i-ai/how-mozilla-fixed-500-security-bugs-with-mythos">https://www.chatprd.ai/how-i-ai/how-mozilla-fixed-500-security-bugs-with-mythos</a><br>&#8627; Create an AI-Powered Patch and Verification Loop for Security Bugs: <a href="https://www.chatprd.ai/how-i-ai/workflows/create-an-ai-powered-patch-and-verification-loop-for-security-bugs">https://www.chatprd.ai/how-i-ai/workflows/create-an-ai-powered-patch-and-verification-loop-for-security-bugs</a><br>&#8627; Use an LLM as a Security Judge to Prioritize Codebase Analysis: <a href="https://www.chatprd.ai/how-i-ai/workflows/use-an-llm-as-a-security-judge-to-prioritize-codebase-analysis">https://www.chatprd.ai/how-i-ai/workflows/use-an-llm-as-a-security-judge-to-prioritize-codebase-analysis</a><br>&#8627; Build an AI Agentic Harness for Automated Security Bug Hunting: <a href="https://www.chatprd.ai/how-i-ai/workflows/build-an-ai-agentic-harness-for-automated-security-bug-hunting">https://www.chatprd.ai/how-i-ai/workflows/build-an-ai-agentic-harness-for-automated-security-bug-hunting</a></p></div><div><hr></div><p>If you&#8217;re enjoying these episodes, reply and let me know what you&#8217;d love to learn more about: AI workflows, hiring, growth, product strategy&#8212;anything.</p><p>Catch you next week,<br>Lenny</p><p><em>P.S. Want every new episode delivered the moment it drops? Hit &#8220;Follow&#8221; on your favorite podcast app.</em></p>]]></content:encoded></item><item><title><![CDATA[How Claude Mythos found a 15-year-old bug in Mozilla Firefox | Brian Grinstead]]></title><description><![CDATA[Watch now | &#127897;&#65039;423 security fixes in one month: Brian Grinstead (Mozilla) shows the goal-loop harness behind it, and why the model was only half the story]]></description><link>https://www.lennysnewsletter.com/p/how-claude-mythos-found-a-15-year</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-claude-mythos-found-a-15-year</guid><dc:creator><![CDATA[Claire Vo]]></dc:creator><pubDate>Mon, 22 Jun 2026 12:03:06 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/202480416/565d4d9f8648f6fc863051748270dd17.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-Idjt53tTv2U" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;Idjt53tTv2U&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/Idjt53tTv2U?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong><span>Brian Grinstead </span></strong><span>is a distinguished engineer at Mozilla, where he&#8217;s worked on Firefox and the web platform since 2013 (he joined to help launch Firefox DevTools). Recently he and his team pointed an agentic bug-finding pipeline at Firefox&#8212;a codebase with tens of thousands of files and tens of millions of lines of code&#8212;and shipped a record month of security fixes. The viral chart everyone saw gave the credit to Anthropic&#8217;s new Mythos model. Brian&#8217;s take is that the harness and pipeline did just as much of the work, and he walks through exactly how it runs and how anyone can build a starter version.</span></p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/Idjt53tTv2U">YouTube</a>, <a href="https://open.spotify.com/episode/4MNilflycEO8KNZRM3gtWc">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/how-claude-mythos-found-a-15-year-old-bug-in/id1809663079?i=1000773719910">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p><span>How to build a basic bug-finding harness by running Claude Code or Codex with one prompt and the -p flag, no SDK required</span></p></li><li><p><span>Why pointing an agent at a whole codebase fails, and how an LLM judge can score and rank files before you spend any compute</span></p></li><li><p><span>How a verifier subagent kills false positives by catching the agent when it cheats</span></p></li><li><p><span>The goal-loop pattern: give an agent a tightly scoped problem, a clear pass/fail signal, and let it retry far past the point a human would quit</span></p></li><li><p><span>Why teams that already invested in fuzzing, CI, and dev tooling are so far ahead</span></p></li><li><p><span>How to weigh model versus harness, and why Brian splits the credit close to 50-50</span></p></li><li><p><span>How a non-engineer can reuse the same score, verify, and fix the loop for design quality, conversion rate, or tech debt</span></p></li><li><p><span>Why AI-generated patches still can&#8217;t ship on their own, and where humans stay in the loop</span></p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TuKR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TuKR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!TuKR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!TuKR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!TuKR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TuKR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:30703,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/202480416?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TuKR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!TuKR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!TuKR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!TuKR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7deca321-4db7-4855-8a92-3d86293a4582_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://workos.com/?utm_source=lennys_howiai&amp;utm_medium=podcast&amp;utm_campaign=q22025">WorkOS</a></strong>&#8212;Make your app enterprise-ready today</p><p><strong><a href="https://www.metaview.ai/home/how-i-ai">Metaview</a></strong>&#8212;The agentic recruiting platform for winning teams</p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U">00:00</a>) Introduction to Brian Grinstead</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=163s">02:43</a>) The viral chart: Firefox Security Bug Fixes by Month</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=332s">05:32</a>) How the custom harness works</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=622s">10:22</a>) Goal loops and guardrails</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=885s">14:45</a>) How they built it</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=1015s">16:55</a>) Real bugs, including a 15-year-old one</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=1380s">23:00</a>) Open-sourcing it</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=1586s">26:26</a>) Why humans still review every fix</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=1950s">32:30</a>) Live demo and prioritizing files</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=2418s">40:18</a>) Mobilizing the team and recap</p><p>(<a href="https://www.youtube.com/watch?v=Idjt53tTv2U&amp;t=2553s">42:33</a>) Lightning round</p><p></p><h3>Tools referenced:</h3><p>&#8226; Claude Code: <a href="https://claude.ai/code">https://claude.ai/code</a></p><p>&#8226; Claude Agent SDK: <a href="https://code.claude.com/docs/en/agent-sdk/overview">https://code.claude.com/docs/en/agent-sdk/overview</a></p><p>&#8226; Codex: <a href="https://openai.com/index/openai-codex/">https://openai.com/index/openai-codex/</a></p><p>&#8226; OpenAI Agent SDK: <a href="https://developers.openai.com/api/docs/guides/agents">https://developers.openai.com/api/docs/guides/agents</a></p><p>&#8226; VS Code: <a href="https://code.visualstudio.com/">https://code.visualstudio.com/</a></p><p>&#8226; Docker: <a href="https://www.docker.com/">https://www.docker.com/</a></p><p>&#8226; Firefox: <a href="https://www.mozilla.org/firefox/">https://www.mozilla.org/firefox/</a></p><p>&#8226; Address Sanitizer: <a href="https://github.com/google/sanitizers">https://github.com/google/sanitizers</a></p><p>&#8226; RLBox: <a href="https://rlbox.dev/">https://rlbox.dev/</a></p><p></p><h3>Other references:</h3><p>&#8226; Mozilla Bug Bounty Program: <a href="https://www.mozilla.org/security/bug-bounty/">https://www.mozilla.org/security/bug-bounty/</a></p><p>&#8226; Mozilla GitHub: <a href="https://github.com/mozilla">https://github.com/mozilla</a></p><p></p><h3>Where to find <span>Brian Grinstead:</span></h3><p>LinkedIn: <a href="https://www.linkedin.com/in/bgrins/">https://www.linkedin.com/in/bgrins/</a></p><p>GitHub: <a href="https://github.com/bgrins">https://github.com/bgrins</a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[How to design AI agent loops: schedules, goals, and subagents in Claude Code and Codex]]></title><description><![CDATA[Watch now | &#127897;&#65039;Every loop type explained: heartbeats, crons, goals, and subagents. Plus two live Claude Code and Codex builds that run autonomously so you never manually babysit a PR again]]></description><link>https://www.lennysnewsletter.com/p/how-to-design-ai-agent-loops-schedules</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-to-design-ai-agent-loops-schedules</guid><pubDate>Wed, 17 Jun 2026 12:04:04 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/202288504/bcaf736dba1dcc855294378d2aa12db1.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<div id="youtube2-JoXbk2fm7jM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;JoXbk2fm7jM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/JoXbk2fm7jM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p>I break down every loop type from scratch&#8212;what a heartbeat, cron, hook, and goal loop actually are, when each one fits, and the five things any effective loop needs before it touches production. Then I build two live loops: a daily aging-PR reviewer in Claude Code that schedules itself at 10:15 a.m. and spins off its own subagents, and a weekly skills-identification loop in Codex that spawns goal-based subagents to validate its own output in real time.</p><div class="pullquote"><p><strong>Listen or watch on <a href="https://youtu.be/JoXbk2fm7jM">YouTube</a>, <a href="https://open.spotify.com/episode/43kcUzpSfJExbkIotdUbBp">Spotify</a>, or <a href="https://podcasts.apple.com/us/podcast/how-to-design-ai-agent-loops-schedules-goals-and/id1809663079?i=1000773109920">Apple Podcasts</a></strong></p></div><h3>What you&#8217;ll learn:</h3><ol><li><p>The plain-English definition of a loop&#8212;and why it&#8217;s just an automated prompt, not a scary new paradigm</p></li><li><p>The four loop types (heartbeat, cron, hook, and goal) and when each one actually fits your workflow</p></li><li><p>How to think about loop design using the &#8220;onboarding an employee&#8221; mental model</p></li><li><p>The five things every effective loop needs: work trees, skills, plugins/connectors, subagents, and state tracking</p></li><li><p>How to build a scheduled PR-review routine in Claude Code that babysits aging PRs and alerts your team</p></li><li><p>How to set up a weekly skills-identification automation in Codex that spawns its own validating subagents</p></li><li><p>Why goal-based loops are the hardest to write well&#8212;and where most people burn tokens for nothing</p></li><li><p>The two warning signs that your loop is going to get expensive before it gets useful</p></li></ol><div><hr></div><h3>Brought to you by:</h3><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6NSK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6NSK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!6NSK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!6NSK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!6NSK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6NSK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:20114,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/202288504?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6NSK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!6NSK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!6NSK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!6NSK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8c9e213-4cdf-4f70-a58c-066fab0ff72c_1600x114.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><p><strong><a href="https://workos.com/?utm_source=lennys_howiai&amp;utm_medium=podcast&amp;utm_campaign=q22025">WorkOS</a></strong>&#8212;Make your app enterprise-ready today</p><p><strong><a href="https://runwayml.com/howIAI">Runway</a></strong>&#8212;The creative AI platform for images, video, and more</p><p></p><h3>In this episode, we cover:</h3><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM">00:00</a>) Prompts are out and loops are in</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=150s">02:30</a>) Defining a loop</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=183s">03:03</a>) The four ways to automate a prompt: heartbeat, cron, hooks, and goals</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=363s">06:03</a>) Five things every effective loop needs</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=566s">09:26</a>) The &#8220;onboarding an employee&#8221; framework for designing loops</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=718s">11:58</a>) Live build #1: Daily aging PR loop in Claude Code</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=1028s">17:08</a>) Subagents inside loops</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=1140s">19:00</a>) Live build #2: Weekly skills identification loop in Codex</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=1377s">22:57</a>) Watching subagents spin up in real time</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=1528s">25:28</a>) Warning signals around loops</p><p>(<a href="https://www.youtube.com/watch?v=JoXbk2fm7jM&amp;t=1651s">27:31</a>) What listeners are doing with loops</p><p></p><h3>Tools referenced:</h3><p>&#8226; Claude Code: <a href="https://claude.ai/code">https://claude.ai/code</a></p><p>&#8226; Codex: <a href="https://chatgpt.com/codex">https://chatgpt.com/codex</a></p><p>&#8226; OpenClaw: <a href="https://openclaw.ai/">https://openclaw.ai/</a></p><p></p><h3>Other references:</h3><p>&#8226; Claire&#8217;s article &#8220;Why OpenClaw Feels Alive Even Though It&#8217;s Not&#8221;: <a href="https://x.com/clairevo/article/2017741569521271175">https://x.com/clairevo/article/2017741569521271175</a></p><p>&#8226; Addy Osmani&#8217;s article on loop engineering: <a href="https://addyosmani.com/blog/loop-engineering/">https://addyosmani.com/blog/loop-engineering/</a></p><p>&#8226; Using Goals in Codex: <a href="https://developers.openai.com/cookbook/examples/codex/using_goals_in_codex">https://developers.openai.com/cookbook/examples/codex/using_goals_in_codex</a></p><p></p><h3>Where to find Claire Vo:</h3><p>ChatPRD: <a href="https://www.chatprd.ai/">https://www.chatprd.ai/</a></p><p>Website: <a href="https://clairevo.com/">https://clairevo.com/</a></p><p>LinkedIn: <a href="https://www.linkedin.com/in/clairevo/">https://www.linkedin.com/in/clairevo/</a></p><p>X: <a href="https://x.com/clairevo">https://x.com/clairevo</a></p><p></p><p>Production and marketing by <a href="https://penname.co/">https://penname.co/</a>. For inquiries about sponsoring the podcast, email jordan@penname.co.</p>]]></content:encoded></item><item><title><![CDATA[🎙️ How I AI: Claude Fable 5 review & How Braintrust uses AI agents, evals, and CI to ship better software]]></title><description><![CDATA[Your weekly listens from How I AI, part of the Lenny's Podcast Network]]></description><link>https://www.lennysnewsletter.com/p/how-i-ai-claude-fable-5-review-and</link><guid isPermaLink="false">https://www.lennysnewsletter.com/p/how-i-ai-claude-fable-5-review-and</guid><dc:creator><![CDATA[Lenny Rachitsky]]></dc:creator><pubDate>Mon, 15 Jun 2026 15:01:32 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1c21ead7-9311-453f-9dce-9afa916e3466_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gWeJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png" width="1456" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:76503,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/177292431?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!gWeJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 424w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 848w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1272w, https://substackcdn.com/image/fetch/$s_!gWeJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361d81ef-7faf-4d8e-8028-5d5e03432a9a_2329x551.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h3>Claude Fable 5 review: what the new Mythos model gets right (and very wrong)</h3><div id="youtube2-IREnr4I89Ho" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;IREnr4I89Ho&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/IREnr4I89Ho?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/IREnr4I89Ho">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/2b6KxnlVcSeVFKQzKAPjFd">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/claude-fable-5-review-what-the-new-mythos-model-gets/id1809663079?i=1000771908698">Apple Podcasts</a></strong></p></div><p>Claire puts Claude Fable 5, Anthropic&#8217;s first generally available Mythos-class model, through a series of real-world tests: product specs, agent workflows, design tasks, vision tasks, and multi-agent orchestration. She breaks down what Anthropic is claiming, where the model genuinely feels like a leap forward, and where it surprisingly falls short. </p><h4>Biggest takeaways:</h4><ol><li><p><strong>Fable 5 is Anthropic&#8217;s first &#8220;Mythos-class&#8221; model to reach general availability, and it&#8217;s crushing benchmarks across the board. </strong>It hit 80% on SWBench Pro, significantly outperforming Opus 4.8, GPT-4.5, and Gemini 3.1 Pro. Claire found the model excels in specific areas while falling short in others that matter for everyday product work.</p></li><li><p><strong>The model is expensive by design: $10 per million input tokens and $50 per million output tokens. </strong>That&#8217;s a new tier above Opus, and it consumes tokens at roughly twice the rate of other models. You need to be strategic about when to deploy this level of intelligence versus using cheaper models like Sonnet or Opus for simpler tasks.</p></li><li><p><strong>Fable 5 works like a &#8220;seasoned engineer&#8221;&#8212;which is both its superpower and its Achilles&#8217; heel. </strong>It&#8217;s thorough, autonomous, and will investigate every corner of a problem to be 120% sure it&#8217;s shipping the right thing. Sometimes you need a model that&#8217;s a little less thorough, a little &#8220;dumber,&#8221; to actually ship something useful quickly.</p></li><li><p><strong>The model is exceptionally good at vision tasks, particularly document formatting and PDF parsing. </strong>Claire tested it on creating handwriting worksheets for her 7-year-old and found it dramatically outperformed Opus 4.8&#8212;better spacing, clearer layout, appropriate white space. This extends to other vision tasks where you want something to look good or need to parse complex documents.</p></li><li><p><strong>The writing is nearly unreadable for specs and PRDs. </strong>Claire found that Fable 5 produces extremely detailed, technically complete documents that are almost impossible to parse. It gets wrapped around the axle on details, creates big blocks of dense paragraphs with internal references, and makes it hard to see the forest for the trees.</p></li><li><p><strong>Design output is shockingly bad, at least for one-shot design tasks. </strong>When Claire asked Fable to design a skills registry, it produced fundamentally terrible design: gray, black, red, simple outlines. This was a real surprise given the model&#8217;s benchmark performance.</p></li><li><p><strong>The model is conservative on execution and takes &#8220;minimal&#8221; very literally. </strong>When Claire asked it to ship an MVP that would deliver customer value, Fable produced something extremely narrow and not actually that useful. This conservatism may stem from the safety guardrails built into the model.</p></li><li><p><strong>Fable 5 includes specific safeguards for cybersecurity, biology, chemistry, and distillation tasks.</strong> Instead of blocking you entirely, it uses a new &#8220;fallback&#8221; concept&#8212;if you get classified into one of these categories, it gracefully falls back to Opus 4.8. Anthropic reports that 95% of sessions don&#8217;t hit a fallback, and they maintain a 30-day retention policy solely to catch misuse.</p></li><li><p><strong>Multi-agent orchestration is technically possible but not yet reliable.</strong> Claire tested the dynamic workflows and subagent capabilities extensively and had some successful multi-agent runs, but also encountered frequent stalls and errors. She walked away from her laptop and came back to find subagents had stalled after about three hours.</p></li><li><p><strong>The key insight: match model intelligence to task complexity.</strong> Claire recommends using it for hard technical problems where extreme detail matters, long-horizon work, and vision tasks. But for front-end work, strategy, specs, and design, other models in the ecosystem will serve you better and cost less.</p></li><li><p><strong>This is &#8220;baby Mythos,&#8221; not the full Mythos model. </strong>Fable 5 has guardrails that the unrestricted Mythos model (available only to Project Glasswing partners) doesn&#8217;t have. The underlying model is the same, but Fable is tuned for safety and general availability.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog from this episode:</h4><p><strong>How I AI: My Honest Review of Claude Fable 5:</strong> <a href="https://www.chatprd.ai/how-i-ai/claude-fable-5-review">https://www.chatprd.ai/how-i-ai/claude-fable-5-review</a></p></div><h3>How Braintrust uses AI agents, evals, and CI to ship better software | Ankur Goyal</h3><div id="youtube2-QE_1hRLsehM" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;QE_1hRLsehM&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/QE_1hRLsehM?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><div class="pullquote"><p>Listen now on <strong><a href="https://youtu.be/QE_1hRLsehM">YouTube</a> &#8226; <a href="https://open.spotify.com/episode/6jZjRjDBNC3QIgyzD17MDU">Spotify</a> &#8226; <a href="https://podcasts.apple.com/us/podcast/how-braintrust-uses-ai-agents-evals-and-ci-to-ship/id1809663079?i=1000772794077">Apple Podcasts</a></strong></p></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NNhG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NNhG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!NNhG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!NNhG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!NNhG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NNhG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png" width="1456" height="104" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:104,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:37219,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.lennysnewsletter.com/i/201229065?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NNhG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png 424w, https://substackcdn.com/image/fetch/$s_!NNhG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png 848w, https://substackcdn.com/image/fetch/$s_!NNhG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png 1272w, https://substackcdn.com/image/fetch/$s_!NNhG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab377b2-8000-417b-9ad4-79cf02a5d8f5_1600x114.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><blockquote><p><strong>Brought to you by:</strong></p><ul><li><p><strong><a href="https://www.getguru.com/?utm_source=howi_ai_podcast&amp;utm_medium=podcast&amp;utm_campaign=q1">Guru</a></strong>&#8212;The AI layer of truth</p></li><li><p><strong><a href="https://withpersona.com/lp/howiai">Persona</a></strong>&#8212;Trusted identity verification for any use case</p></li></ul></blockquote><p>Claire sits down with <strong>Ankur Goyal</strong>, the founder and CEO of Braintrust, to unpack how top engineering teams are using AI agents, evals, and CI to ship better software faster. They get into why agents are now capable of tackling hard infrastructure problems, how to decide what work sits &#8220;below the agent line,&#8221; and why evals are quickly becoming the modern version of a PRD. Ankur&#8217;s core message: the best teams won&#8217;t just use AI to write more code; they&#8217;ll build the feedback loops, benchmarks, and systems that let AI improve the quality of the product itself.</p><h4>Biggest takeaways:</h4><ol><li><p><strong>There&#8217;s no staff engineer running as many rigorous benchmarks as someone using an agent.</strong> Ankur viscerally disagrees with engineers who say AI can&#8217;t handle complicated problems. While models might not be perfect at writing highly concurrent code, they excel at running exhaustive experiments&#8212;testing every column store format, every execution engine, every optimization strategy. The baseline of rigor you get from agents is incredible, and there&#8217;s simply no excuse anymore to skip benchmarks because they&#8217;re tedious.</p></li><li><p><strong>The agent line keeps going up&#8212;and you need to identify what&#8217;s below it.</strong> Many interactions, decisions, and directions that feel like they need human judgment actually fit &#8220;below the agent line.&#8221; If you took the information from a meeting and gave it to an agent, would it solve the same problem? Increasingly, the answer is yes. The best teams push this line higher by building smart skills and integrations that expand what agents can handle autonomously.</p></li><li><p><strong>Practical quality beats theoretical quality every time.</strong> In theory, a human engineer with infinite time and focus might produce better code than an AI agent. In practice, humans lose context over days, have decaying attention spans on hard-but-tedious problems, and skip benchmarks they know they should run. AI agents maintain consistent focus, run every test, and can work on problems continuously for days or weeks. The practical quality of AI-assisted engineering is higher because of sustained rigor, not because the code is theoretically better.</p></li><li><p><strong>You can now bite off much harder technical problems than before.</strong> Companies historically avoid major infrastructure changes because the cost of testing alternatives is prohibitively high and the unknown unknowns are risky. With AI agents, you can exhaustively test six different database solutions, run thousands of benchmarks on production-scale data, and make informed decisions about platform shifts that would have been impossible before. The business case for deep technical work becomes much easier when agents do the heavy lifting.</p></li><li><p><strong>Run four to six foreground agents simultaneously&#8212;that&#8217;s the human concurrency limit. </strong>Ankur runs different agents working on different problems. This matches the personal concurrency limit most people can manage; you can&#8217;t effectively context switch between more than that. Some agents run locally, and others run remotely on cloud infrastructure with production-scale data. The key is isolation: each agent has its own environment, ports, and services.</p></li><li><p><strong>Evals are the modern PRD&#8212;they define </strong><em><strong>what</strong></em><strong> success looks like, not </strong><em><strong>how</strong></em><strong> to achieve it. </strong>Machine learning shifts programming from defining implementation details to defining success criteria. Just like the best PRDs include user stories and examples, the best evals include concrete test cases and scoring functions. The difference is that evals quantify success in ways that can be automatically measured and improved. This lets you focus on outcomes while AI figures out the implementation.</p></li><li><p><strong>Build a feedback loop that automatically turns real-world data into evals. </strong>For AI product teams, the #1 engineering priority isn&#8217;t prompt engineering or picking an agent framework&#8212;it&#8217;s building a pipeline that summons real-world data and converts it into evals. This is the same principle as investing in CI for traditional software: you&#8217;re building the platform that lets agents do the work engineers used to do manually. Without this feedback loop, you&#8217;re stuck in whack-a-mole mode, fixing individual cases without systematic improvement.</p></li><li><p><strong>Quantify your designer&#8217;s taste so it scales across your product.</strong> Ankur runs hundreds of evals to improve things quantitatively, then asks David (their tastemaker designer) for a vibe check every few days. When David destroys his work, Ankur captures the feedback (&#8220;David thinks it&#8217;s OK to show both languages as long as . . .&#8221;) and improves the scoring functions to encode David&#8217;s palette. This doesn&#8217;t replace David; it amplifies him. They&#8217;re able to apply David&#8217;s quality bar to more things than he could ever review manually.</p></li><li><p><strong>Product building is now carving, not constructing. It&#8217;s extremely fast to create something with too many features, too many buttons, and too much code.</strong> The hard part is removing stuff. When customers complain, Braintrust removes the thing causing confusion 90% of the time, making the system work better by eliminating complexity. This is the opposite of traditional product development, where you carefully add features one by one.</p></li><li><p><strong>Invest in CI to earn the ability to move faster&#8212;it&#8217;s the platform for AI-powered engineering.</strong> Every engineer is now building a platform upon which agents do the work engineers used to do manually. For traditional software, that platform is CI. If you feel constrained by velocity, don&#8217;t ship crappy stuff faster. Instead, pause and improve CI so you earn the ability to move faster safely. The same principle applies to AI products: build the eval pipeline first, then let agents optimize within that system.</p></li><li><p><strong>When agents fail, close the session and improve the evals&#8212;don&#8217;t yell or bribe.</strong> Ankur&#8217;s back-pocket strategy is remarkably disciplined: he doesn&#8217;t try to prompt his way out of problems. He closes the session, improves the evaluation criteria or success metrics, and starts fresh. Sometimes this means hand-writing code to better understand the problem (like when he spent a weekend hand-writing a 3,000-line eval that had become trash through vibe coding). The solution is always better evals, not better prompting.</p></li></ol><div class="callout-block" data-callout="true"><h4>Blog and detailed workflow walkthroughs from this episode:</h4><p><strong>Blog: </strong>Ankur Goyal&#8217;s Playbook for Agent-Driven Benchmarking and AI Evals<a href="https://www.chatprd.ai/how-i-ai/ankur-goyals-playbook-for-agent-driven-benchmarking-and-ai-evals"> https://www.chatprd.ai/how-i-ai/ankur-goyals-playbook-for-agent-driven-benchmarking-and-ai-evals</a></p><p><strong>Workflows:</strong></p><p>&#8627; How to Scale Expert Judgment in AI Systems with a Human Feedback Loop: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-scale-expert-judgment-in-ai-systems-with-a-human-feedback-loop">https://www.chatprd.ai/how-i-ai/workflows/how-to-scale-expert-judgment-in-ai-systems-with-a-human-feedback-loop</a></p><p>&#8627; How to Use AI Coding Agents for Exhaustive Infrastructure Benchmarking: <a href="https://www.chatprd.ai/how-i-ai/workflows/how-to-use-ai-coding-agents-for-exhaustive-infrastructure-benchmarking">https://www.chatprd.ai/how-i-ai/workflows/how-to-use-ai-coding-agents-for-exhaustive-infrastructure-benchmarking</a></p></div><div><hr></div><p>If you&#8217;re enjoying these episodes, reply and let me know what you&#8217;d love to learn more about: AI workflows, hiring, growth, product strategy&#8212;anything.</p><p>Catch you next week,<br>Lenny</p><p><em>P.S. Want every new episode delivered the moment it drops? Hit &#8220;Follow&#8221; on your favorite podcast app.</em></p>]]></content:encoded></item></channel></rss>