← Collected sources
BUILDERS · EDITED DIGEST

Builders’ Picks | 2026-09-23

2026-09-23 · Historical edition

X / Twitter

Boris Cherny, Anthropic Claude Code Team

Boris used Opus 5.5 with Lean to formally verify the Claude Agent SDK, producing 16 PRs that fixed different bugs and race conditions from just a few short prompts. He said TLA+ is also well suited to this kind of work, and that he sometimes combines Lean and TLA+ to investigate data flow, concurrency, and state management issues. Although he is not familiar with either language, Claude's performance allowed him to use formal modeling to find defects that manual inspection might miss. This led him to ask whether formal verification could be the future of programming, or at least of finding bugs. In another test, the team had Opus 5.5 and Fable 5.1 each port HAProxy from C to Rust, and both passed almost all HAProxy tests. Opus 5.5 took 9.5 hours and Fable 5.1 took 12 hours, with the former also costing 51% less. Boris said Opus 5.5 had been his daily go-to model for the past few weeks.

Thibault Sottiaux, OpenAI Codex and ChatGPT Teams

Thibault announced the release of GPT-6 Sol and Luna, saying both models had significantly improved overall capabilities. He singled out their writing abilities and broader improvements to the experience that become apparent through actual use. At the same time, API prices for both models will be permanently reduced by 50%, making more use cases economically viable. He said efficiency improvements would also make subscribers' usage allowances go further. OpenAI will additionally give every Plus, Pro, and Business account a one-time usage reset that can be saved for later. He described the team's direction as improving efficiency and intelligence together, emphasizing that capabilities from top-tier models can be used to improve products at other tiers.

Cat Wu, Anthropic Claude Code and Cowork Teams

Cat announced that Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans. She has made it her daily go-to model, particularly appreciating its clear communication and ability to write in her own style. Products default to medium reasoning effort; she said intelligence at this setting is comparable to Fable 5.1, but faster. Compared with Opus 5, Opus 5.5 provides 25% more usage under the same limits. This update affects the default model, reasoning effort, and usage efficiency, directly shaping subscribers' everyday experience. Cat encouraged users to give Opus 5.5 more challenging tasks and report back on its performance.

Thariq, Anthropic Claude Code Team

Thariq believes the right response to more capable models is not to ship 10 times as many features to production. He advocates spending more time understanding users, running experiments, building prototypes, and learning about areas one does not yet understand. The aim of these investments is to make what ultimately ships genuinely effective, rather than simply increasing the number of features. In his own use, workflows have become an important part of how he uses Claude. The development he values is that intelligence approaching Fable's level is now more affordable for workflows. Together, these two views point to a shared concern: the value of model progress needs to be realized through more extensive exploration and affordable practical use.

Guillermo Rauch, Vercel CEO

Guillermo suggested that software can live on through versions users generate and deploy themselves, using Google Reader as an example: if you love a product, you can build your own version. He also released a new round of Next.js evaluations, in which Opus 5.5, GPT 6 Sol, and Fable 5.1 all scored 97%, while Grok 4.7 scored 94%. Beyond the scores, he specifically noted that Grok was 2 to 7 times cheaper, making price another factor worth considering alongside these results. On design, he praised the aesthetic sensibility of Anthropic's product launches. He explained that what originally drew him to headless web architecture was the possibility of pages taking all kinds of distinctive, imaginative forms. In his view, AI further lowers the barrier to bringing these designs to life, and developers should continue exploring the boundaries of web design.

Alex Albert, Anthropic Research Team

Alex shared an attempt to reconstruct a historical streetscape using Opus 5.5 and Blender, with the goal of recreating San Francisco's Market Street before the 1906 earthquake. His prompt to Claude Tag specified the afternoon of April 17, 1906, covering Market Street from the Ferry Building to Fifth Street and naming landmarks such as the Palace Hotel and Call Building. Before modeling began, the model was required to create a source file based on Sanborn fire insurance maps from 1899 to 1905, the Miles Brothers' street footage, historical photographs, and USGS topographic data. Each building had to have records of its footprint, height, facade materials, occupants, and the source and confidence level for each fact. All modeling had to use Blender Python, without downloading premade meshes, textures, or HDRIs. The prompt also required reusable generators for building facades, roofs, windows, streetlights, and vehicles, followed by assembly of the street based on the source file so that buildings could be traced back to the data. The final deliverables included a 10-second video moving along the street. Alex credited the attempt to Opus 5.5's improved 3D modeling and visual capabilities.

Aaron Levie, Box CEO

Aaron connected Opus 5.5's price reduction with the 50% drop in token prices for GPT-6 Sol and Luna, arguing that lower costs per task would significantly expand the range of viable agent use cases—an instance of Jevons paradox applied to agents. Examples he listed included processing enterprise data, scanning code for security issues, analyzing logs, and using multiple agents within workflows. After testing complex unstructured data tasks with Box Agent, Box reported that, compared with Opus 5, Opus 5.5 used 63% fewer tokens, was 42% less verbose, and was 30% faster. In financial due diligence tests, task accuracy improved by 39%; the model found every calculation error in the acquisition target's pricing tool on every run, while total token usage fell by 82%. In cloud cost analysis, task accuracy improved by 65%; the model correctly selected the basis for calculating retention and kept units consistent, halving the time required and reducing token usage by 70%. In customer account analysis, task accuracy improved by 17%; the model inferred a seniority mix not specified in the contract from a roster of 18 people and calculated average satisfaction from each customer's responses, rather than being misled by a perfect score on a single item. In malaria rapid diagnostic test data analysis, task accuracy improved by 15%; the model discovered that the standard deviations of two groups differed by more than 100 times and, after reanalyzing the data, concluded that the dry-season difference did not hold up. Aaron said customers would soon be able to use Opus 5.5 to build AI Agents in Box AI Studio.

Garry Tan, Y Combinator CEO

Garry shared his experience using Capy, saying it let him submit PRs significantly faster than using Codex or Claude Code alone. This was a personal impression, with no specific speedup ratio or evaluation data provided. In another post, he advocated teaching more people to write prompts and make full use of AI. He wants people to see how AI can help them pursue their ambitions in different fields. Alongside tool use, he called for being more proactive in finding and solving problems, both one's own and those of others.

Nikunj Kothari, FPV Ventures Partner

Nikunj cautioned readers to treat headlines about large funding rounds carefully, because even top investors are assembling a striking number of SPVs, or special purpose vehicles. He argued that some rounds that appear to be backed by large funds actually draw a substantial share of their capital from SPVs, a structure the headlines may not reveal. Add valuations set in tranches and inconsistent revenue definitions, and public reporting struggles to give a complete picture of a company's actual situation. His central point was that not all evidence that can truly validate performance is publicly available. He also published Part 3 of A Walk In The Park with Todd Saunders, covering the move from Google to software for the flooring industry, an experience labeled a "$10 million pivot," and life after an exit. The discussion's contents also included vertical versus horizontal software, co-founders, and why he raised venture capital again. The video was shot on an iPhone in Westfield, New Jersey, and edited in Descript.

Peter Steinberger, OpenClaw and OpenAI

Peter shared an experience tracing an application crash to an underlying defect. After upgrading to macOS 27, he encountered occasional ChatGPT crashes. Astra subsequently found a bug in libuv that had existed for approximately 14 years. The update specified the operating system version, the troubleshooting tool, and the component containing the defect. However, the source material did not explain the bug's technical cause or state whether a fix had been merged or released.

Aditya Agarwal, SPC General Partner and Bevel Health Co-founder

Aditya placed today's discussions of model safety and alignment in the context of self-driving technology, noting that some of the early major safety debates around AI and machine learning took place in that field. He recalled that what struck him most on his first Waymo ride was that the technology could handle the many complexities of real roads. During a conversation with Dmitri Dolgov at SPC, what interested him most was the large-scale evaluation and testing infrastructure Waymo had built to establish confidence in releases. That infrastructure serves a very concrete task: enabling a 2-tonne robot to navigate urban environments at 30 miles per hour. He has hosted more than 50 Minus One fireside chats, and the Waymo conversation was the first event his 9-year-old child had actively asked to attend, bringing home the appeal of building products for the physical world. For AI products themselves, he also reminded developers not to lose sight of delight and fun in the pursuit of productivity, citing his own experience using Sentience.