X / Twitter
Boris Cherny, Anthropic Claude Code Team
Boris used Opus 5.5 with Lean to formally verify the Claude Agent SDK, producing 16 PRs that fixed different bugs and race conditions from just a few short prompts. He said TLA+ is also well suited to this kind of work, and that he sometimes combines Lean and TLA+ to investigate data flow, concurrency, and state management issues. Although he is not familiar with either language, Claude's performance allowed him to use formal modeling to find defects that manual inspection might miss. This led him to ask whether formal verification could be the future of programming, or at least of finding bugs. In another test, the team had Opus 5.5 and Fable 5.1 each port HAProxy from C to Rust, and both passed almost all HAProxy tests. Opus 5.5 took 9.5 hours and Fable 5.1 took 12 hours, with the former also costing 51% less. Boris said Opus 5.5 had been his daily go-to model for the past few weeks.
https://x.com/bcherny/status/2102543349102338309
https://x.com/bcherny/status/2102439069053747549
Thibault Sottiaux, OpenAI Codex and ChatGPT Teams
Thibault announced the release of GPT-6 Sol and Luna, saying both models had significantly improved overall capabilities. He singled out their writing abilities and broader improvements to the experience that become apparent through actual use. At the same time, API prices for both models will be permanently reduced by 50%, making more use cases economically viable. He said efficiency improvements would also make subscribers' usage allowances go further. OpenAI will additionally give every Plus, Pro, and Business account a one-time usage reset that can be saved for later. He described the team's direction as improving efficiency and intelligence together, emphasizing that capabilities from top-tier models can be used to improve products at other tiers.
https://x.com/thsottiaux/status/2102463847714247142
https://x.com/thsottiaux/status/2102440619616682120
Cat Wu, Anthropic Claude Code and Cowork Teams
Cat announced that Opus 5.5 is now the default model in Claude Code and the Claude app, including Cowork, for Pro, Max, and Team plans. She has made it her daily go-to model, particularly appreciating its clear communication and ability to write in her own style. Products default to medium reasoning effort; she said intelligence at this setting is comparable to Fable 5.1, but faster. Compared with Opus 5, Opus 5.5 provides 25% more usage under the same limits. This update affects the default model, reasoning effort, and usage efficiency, directly shaping subscribers' everyday experience. Cat encouraged users to give Opus 5.5 more challenging tasks and report back on its performance.
https://x.com/_catwu/status/2102437713781944397
Thariq, Anthropic Claude Code Team
Thariq believes the right response to more capable models is not to ship 10 times as many features to production. He advocates spending more time understanding users, running experiments, building prototypes, and learning about areas one does not yet understand. The aim of these investments is to make what ultimately ships genuinely effective, rather than simply increasing the number of features. In his own use, workflows have become an important part of how he uses Claude. The development he values is that intelligence approaching Fable's level is now more affordable for workflows. Together, these two views point to a shared concern: the value of model progress needs to be realized through more extensive exploration and affordable practical use.
https://x.com/trq212/status/2102548686303854790
https://x.com/trq212/status/2102477527688388752
Guillermo Rauch, Vercel CEO
Guillermo suggested that software can live on through versions users generate and deploy themselves, using Google Reader as an example: if you love a product, you can build your own version. He also released a new round of Next.js evaluations, in which Opus 5.5, GPT 6 Sol, and Fable 5.1 all scored 97%, while Grok 4.7 scored 94%. Beyond the scores, he specifically noted that Grok was 2 to 7 times cheaper, making price another factor worth considering alongside these results. On design, he praised the aesthetic sensibility of Anthropic's product launches. He explained that what originally drew him to headless web architecture was the possibility of pages taking all kinds of distinctive, imaginative forms. In his view, AI further lowers the barrier to bringing these designs to life, and developers should continue exploring the boundaries of web design.
https://x.com/rauchg/status/2102594015669756323
https://x.com/rauchg/status/2102519097770885231
https://x.com/rauchg/status/2102438365455167883
Alex Albert, Anthropic Research Team
Alex shared an attempt to reconstruct a historical streetscape using Opus 5.5 and Blender, with the goal of recreating San Francisco's Market Street before the 1906 earthquake. His prompt to Claude Tag specified the afternoon of April 17, 1906, covering Market Street from the Ferry Building to Fifth Street and naming landmarks such as the Palace Hotel and Call Building. Before modeling began, the model was required to create a source file based on Sanborn fire insurance maps from 1899 to 1905, the Miles Brothers' street footage, historical photographs, and USGS topographic data. Each building had to have records of its footprint, height, facade materials, occupants, and the source and confidence level for each fact. All modeling had to use Blender Python, without downloading premade meshes, textures, or HDRIs. The prompt also required reusable generators for building facades, roofs, windows, streetlights, and vehicles, followed by assembly of the street based on the source file so that buildings could be traced back to the data. The final deliverables included a 10-second video moving along the street. Alex credited the attempt to Opus 5.5's improved 3D modeling and visual capabilities.
https://x.com/alexalbert__/status/2102466524934271381
https://x.com/alexalbert__/status/2102466523164274839
Aaron Levie, Box CEO
Aaron connected Opus 5.5's price reduction with the 50% drop in token prices for GPT-6 Sol and Luna, arguing that lower costs per task would significantly expand the range of viable agent use cases—an instance of Jevons paradox applied to agents. Examples he listed included processing enterprise data, scanning code for security issues, analyzing logs, and using multiple agents within workflows. After testing complex unstructured data tasks with Box Agent, Box reported that, compared with Opus 5, Opus 5.5 used 63% fewer tokens, was 42% less verbose, and was 30% faster. In financial due diligence tests, task accuracy improved by 39%; the model found every calculation error in the acquisition target's pricing tool on every run, while total token usage fell by 82%. In cloud cost analysis, task accuracy improved by 65%; the model correctly selected the basis for calculating retention and kept units consistent, halving the time required and reducing token usage by 70%. In customer account analysis, task accuracy improved by 17%; the model inferred a seniority mix not specified in the contract from a roster of 18 people and calculated average satisfaction from each customer's responses, rather than being misled by a perfect score on a single item. In malaria rapid diagnostic test data analysis, task accuracy improved by 15%; the model discovered that the standard deviations of two groups differed by more than 100 times and, after reanalyzing the data, concluded that the dry-season difference did not hold up. Aaron said customers would soon be able to use Opus 5.5 to build AI Agents in Box AI Studio.
https://x.com/levie/status/2102477253070430322
https://x.com/levie/status/2102448415775051790
Garry Tan, Y Combinator CEO
Garry shared his experience using Capy, saying it let him submit PRs significantly faster than using Codex or Claude Code alone. This was a personal impression, with no specific speedup ratio or evaluation data provided. In another post, he advocated teaching more people to write prompts and make full use of AI. He wants people to see how AI can help them pursue their ambitions in different fields. Alongside tool use, he called for being more proactive in finding and solving problems, both one's own and those of others.
https://x.com/garrytan/status/2102544711647129902
https://x.com/garrytan/status/2102501556348440983
Nikunj Kothari, FPV Ventures Partner
Nikunj cautioned readers to treat headlines about large funding rounds carefully, because even top investors are assembling a striking number of SPVs, or special purpose vehicles. He argued that some rounds that appear to be backed by large funds actually draw a substantial share of their capital from SPVs, a structure the headlines may not reveal. Add valuations set in tranches and inconsistent revenue definitions, and public reporting struggles to give a complete picture of a company's actual situation. His central point was that not all evidence that can truly validate performance is publicly available. He also published Part 3 of A Walk In The Park with Todd Saunders, covering the move from Google to software for the flooring industry, an experience labeled a "$10 million pivot," and life after an exit. The discussion's contents also included vertical versus horizontal software, co-founders, and why he raised venture capital again. The video was shot on an iPhone in Westfield, New Jersey, and edited in Descript.
https://x.com/nikunj/status/2102534909076349291
https://x.com/nikunj/status/2102395699895677325
Peter Steinberger, OpenClaw and OpenAI
Peter shared an experience tracing an application crash to an underlying defect. After upgrading to macOS 27, he encountered occasional ChatGPT crashes. Astra subsequently found a bug in libuv that had existed for approximately 14 years. The update specified the operating system version, the troubleshooting tool, and the component containing the defect. However, the source material did not explain the bug's technical cause or state whether a fix had been merged or released.
https://x.com/steipete/status/2102501642176528743
Aditya Agarwal, SPC General Partner and Bevel Health Co-founder
Aditya placed today's discussions of model safety and alignment in the context of self-driving technology, noting that some of the early major safety debates around AI and machine learning took place in that field. He recalled that what struck him most on his first Waymo ride was that the technology could handle the many complexities of real roads. During a conversation with Dmitri Dolgov at SPC, what interested him most was the large-scale evaluation and testing infrastructure Waymo had built to establish confidence in releases. That infrastructure serves a very concrete task: enabling a 2-tonne robot to navigate urban environments at 30 miles per hour. He has hosted more than 50 Minus One fireside chats, and the Waymo conversation was the first event his 9-year-old child had actively asked to attend, bringing home the appeal of building products for the physical world. For AI products themselves, he also reminded developers not to lose sight of delight and fun in the pursuit of productivity, citing his own experience using Sentience.
https://x.com/adityaag/status/2102498288658526614