← Collected sources
BUILDERS · EDITED DIGEST

Builders’ Picks | 2026-09-05

2026-09-05 · Historical edition

X / Twitter

AI Builder Swyx

Swyx is preparing a major report on the frontier models Astra and Fable, and on AEO. During his research, he found that Claude showed a strong preference for recommending Latent Space when asked about the best AI newsletter or podcast. He was surprised by the result, because model recommendations can be influenced by training data, retrieval sources, or context. Feedback from Ric Mac prompted him to check again for a memory leak. The case also illustrates why, when evaluating how models present brands and content sources, it is necessary to first rule out biases introduced by the testing process itself.

OpenAI Codex and ChatGPT team member Thibault Sottiaux

Thibault Sottiaux said that before Astra became widely available, it may have been the team’s greatest competitive advantage. Its productivity gains were so substantial that the team adjusted some product plans. Work originally expected to ship in the middle of next year is now planned for release at DevDay, around 6 months earlier. Astra’s rollout to users was completed ahead of schedule. The team also decided to perform a full banked reset for Plus, Pro, and Business users, with completion planned by the end of the day. Users who created an account or upgraded before 8 p.m. Pacific Time that day would also receive the reset. Together, the two updates send a clear signal: Astra is not just an upgrade in model capabilities; it has already directly changed OpenAI’s internal development speed and release cadence.

AI tutorial creator Peter Yang

Peter Yang suggested that OpenAI release an Apple Watch app. The core interaction would be speaking a task directly into a Codex thread and having Codex return the result by voice. Such an interface would let users continue working with existing task threads without carrying a phone. He believes it could also reduce reliance on phone screens, extending Codex beyond desktop and mobile interfaces to a lighter wearable device. Peter also suggested that this experience might be what the rumored OpenAI device aims to provide. The concept focuses not on shrinking the full ChatGPT interface onto a watch, but on providing an easy voice entry point to continuously running agent workflows.

Meta AI Senior Director Madhu Guru

Madhu Guru advised people seeking to improve their AI product skills to choose a personal or work process they know very well and try to automate it completely with AI. They can start with AI products they already know, but will typically need to experiment with several approaches along the way. The exercise forces a builder to define an excellent end-to-end experience, rather than simply optimize one model call within it. It also reveals where MCP and other tools should enter the process, and which decisions still need to keep a human in the loop. After completing the automation, they must establish an evaluation method to determine whether the results are truly reliable and useful. As their practice develops, they can gradually expand the process’s objectives and complexity. Madhu believes completing one such project teaches more than a month spent reading AI product methodologies, because the real learning comes from building things oneself.

Vercel CEO Guillermo Rauch

Guillermo Rauch is highly optimistic about WebMCP because agents must work on existing WWW infrastructure rather than wait for the entire internet to be rebuilt for them. He compared this to Tesla FSD adapting to real roads, traffic lights, and potholes, emphasizing that WebMCP’s value lies in using the existing web environment more efficiently. One concrete scenario is a Next.js development page directly exposing debugging tools to the agent testing that tab. The agent can then obtain page-level context without first sifting through large volumes of server logs to find problems. Developers also do not need to separately find and configure an MCP server, because the web page itself can provide its own agent tools. He believes that under this architecture, fx and agent-browser can form a complete web development stack while retaining debugging depth. Guillermo also believes AI software factories will eventually produce bug-free software that continuously improves itself. Both views point toward bringing agents closer to real software runtime environments and embedding observation, debugging, and iteration capabilities directly into systems.

Y Combinator CEO Garry Tan

Garry Tan used AsideAI to configure a complete integration between OpenClaw and Slack, an experience that contrasted sharply with the previous process. The original setup took around 2 hours, while AsideAI’s harness completed it in under 3 minutes. It also provided full integrations, browser integration, and fairly comprehensive default access controls. Garry has made the AsideAI browser GStack’s preferred remote session browser. For AI agents that need to access web pages and credentials as the user would, he considers AsideAI the best solution he has tried. Beyond browser and credential management, he particularly praised its integrations, harness, and memory system. The case shows that the practical efficiency of agent tools depends not only on model capabilities but also on whether authentication, permission boundaries, browser control, and long-term memory can be integrated into a coherent workflow.

Builder Zara Zhang

Zara Zhang argued that most people rate AI’s writing ability more highly than its actual quality warrants. This assessment reminds builders not to equate linguistic fluency directly with writing quality. AI-generated text may sound sufficiently natural on the surface while still requiring independent judgment about its ideas, structure, and facts. For people using AI to produce content, the key question is not whether a model can generate a complete piece of text, but whether the finished work meets actual publication standards. Her view shifts attention from generation speed back to editorial judgment and quality assessment.

SPC General Partner Aditya Agarwal

Aditya Agarwal said he understands how RL, inference time scaling, and modern data-verification loops work. Even with an understanding of these technical mechanisms, he still finds the capabilities of contemporary models almost “magical.” This feeling reveals the gap between understanding individual principles and explaining the behavior of an entire system. Model capabilities emerge from the combined effects of multiple training, inference, and verification stages; knowing each component does not make their combined effects intuitive. His reflection reminds AI builders to maintain a sense of awe at capability jumps emerging from complex systems, even as they analyze the underlying mechanisms.

OpenAI’s Sam Altman

Sam Altman announced that GPT-6 Astra was available to all Pro, Enterprise, and Business Premium users in Work and Codex. Astra was also available through the API. At launch, OpenAI said access would gradually expand to Plus and Business users next. A subsequent update confirmed that Astra had reached all Plus and Business users. The two messages show the rollout progressing from higher-tier plans, enterprise plans, and the API to a broader base of paying users. Sam’s direct message to users was that they could now start building with Astra.

Official blogs

Claude in Chrome is generally available

Claude in Chrome is now generally available to all paid Claude plans and can perform some browser actions autonomously, without requiring users to approve each one. It can use users’ existing signed-in sessions to read pages, enter text, click links, navigate between pages, and fill out forms, covering internal dashboards, legacy systems, and vendor portals without connectors. Users can still turn off automatic approval in settings to return to manual confirmation.

The core risk for browser agents is prompt injection: attackers hide malicious instructions in web pages, emails, or form fields to induce models to perform actions users did not request. Claude builds multiple layers of defense through training on a continually expanding set of attack examples, probes that scan tool results, and a safety classifier that checks task intent before execution. In the current, stronger evaluation, without additional protections, attacks that reached the model had a success rate of 17.6% against Opus 4.5 and 3.8% against Opus 5. With probes and the safety classifier enabled, attacks against Sonnet 5, Opus 5, and Mythos 5 were unsuccessful, while the success rate against Fable 5 was 0.3%; manual review classified all successful cases as low severity. The official post also emphasized that these results come from an evolving testing and defense system and do not mean prompt injection risks in browser environments have disappeared.

Claude gets its own browser in Cowork

The Claude Cowork desktop app has added a built-in browser. When a task involves websites, the browser opens in the sidebar, where Claude can navigate pages, read content, click, and type, allowing users to delegate work such as filling out forms, reading dashboard data, or handling portals without connectors directly to it. This capability requires no browser extension and does not share tabs, bookmarks, or passwords from the user’s personal browser by default.

The built-in browser and Claude in Chrome serve different use cases. The former is a work environment separate from the personal browser, suited to web tasks that can be delegated in full, such as research and collecting invoices. The latter is suited to working with pages users have already opened and signed into, such as updating a CRM, handling an inbox, or editing the current document. Users can import sign-in information from Chrome, Edge, or Firefox on a site-by-site basis; banking, email, and SSO sites are excluded by default unless users explicitly select them.

The built-in browser uses the same prompt injection protections and action verification mechanisms as Claude in Chrome, but the official guidance still recommends starting with trusted websites. It will gradually roll out over a week in the Claude desktop app on macOS, Windows, and Linux, with a beta available to Pro, Max, and Team plans. Enterprise administrators can already enable it in organization settings. As long as the desktop app stays open and online, users can also control the browser from the web or mobile.