← Collected sources
BUILDERS · EDITED DIGEST

Builders’ Picks | 2026-09-10

2026-09-10 · Historical edition

X / Twitter

OpenAI Codex and ChatGPT team member Thibault Sottiaux

ChatGPT Work and Codex experienced an issue that day in which banked resets did not fully take effect. Users who used a reset during the affected period will receive another reset and an apology email. Thibault Sottiaux also mentioned that discussing names for new models with researchers is one of the pleasures of the job, because the proposed names are often very entertaining. These updates show both how the team compensates users for failures in its usage mechanisms and a lighter side of the internal naming process before model releases. He said he was happy to be building with the team and believed there were still many directions worth exploring.

AI tutorial creator Peter Yang

Peter Yang shared a way to test a new model’s thinking capabilities: ask it to candidly analyze blind spots that users have not yet recognized but that are holding them back. The specific procedure is to paste an image containing the full prompt directly into ChatGPT. He said Astra’s feedback was particularly sharp, illustrating that this kind of test focuses not just on answering factual questions but on whether a model can form targeted judgments from limited material. He also asked when a related feature would support Work and Remote Codex threads on mobile. Based on his observations, it currently appears to be available only in Chat threads. These two areas of interest concern the quality of model judgment and the consistency of product capabilities across platforms, respectively.

Vercel CEO Guillermo Rauch

Guillermo Rauch announced AI Benchmarks, an event jointly hosted by Benchmark and Vercel. He believes that building benchmarks for models and helping the industry assess capabilities, truthfulness, and efficiency is one of the most important categories in this generation of software. The event is aimed at people interested in model evaluation and its industry implications, and includes a registration link. He also noted that over the past 6 months Vercel had made 8 price reductions, fee removals, or billing structure improvements, and offered discounts on 16 models. He said further changes would follow. These actions show Vercel advancing both the model evaluation ecosystem and cost optimization for AI infrastructure.

Box CEO Aaron Levie

Aaron Levie believes the practical uses of coding agents will extend far beyond what most people previously expected. As the cost of code falls, agents can help teams enter new tool categories, develop systems for companies that previously could not afford custom software, and support cyber risk protection, life sciences research automation, complex data workflows, and upgrades to legacy infrastructure. He believes software will become more useful and engineers will be able to accomplish much more, potentially resulting in a need for more engineers, not fewer. Regarding the gap between AI capabilities and their impact on GDP, he believes the key is that AI will diffuse more slowly than people expect. Companies still need to prepare data, redesign processes, manage change, and coordinate new workflows across their organizations. Real-world timelines for customer replies, project permits, and drug development will not instantly disappear as models become more capable. Many valuable everyday AI use cases may not even increase GDP in the short term. He characterized the theme of the next decade as AI diffusion and believes the biggest opportunity for builders lies in connecting superintelligence with real-world workflows.

Every CEO Dan Shipper

Dan Shipper praised the design of a research report on automation but pointed out two assumptions that substantially affect its conclusions. The first is that jobs can be broken down into tasks. The second is that the human tasks created by automation always represent only a fraction of the tasks automated away, effectively assuming that automation will ultimately reduce human labor. Based on his own experience, he argued the opposite: automation often creates several times more new work for people, so any honest analysis of AI’s economic impact must take job creation seriously. Task lists do make it possible to calculate “how much of a job can be automated,” but they cannot fully describe the job itself. New tasks are not necessarily just items managers add to a list; they can also arise from new observations and actions people make within the same role as tools and environments change. Drawing on Wittgenstein and Heidegger, he described work as a way of observing and caring about the world: editors and nurses, for example, notice different problems. In his words, “What you care about shapes what you notice, what is worth doing, and what counts as doing it well.” If a model cannot explain how new tasks arise, it leaves out a key variable in analyzing automation and employment.

Anthropic AI assistant Claude

Claude Marketplace has added CrowdStrike, Cursor, FactoryAI, Gamma, and Vercel. Enterprises can now use their Anthropic spend commitments to purchase more Claude-powered products and agents. This arrangement extends enterprises’ committed Anthropic spending to third-party ecosystem products, reducing budget friction when procuring these tools. Claude also invited teams building products on Claude Platform to apply for a Marketplace listing. A listing makes it easier for Anthropic customers to discover and purchase those products. For developers, Marketplace therefore serves not only as a showcase but also as a distribution channel into enterprise procurement processes.