← Collected sources
BUILDERS · EDITED DIGEST

Builders’ Picks | 2026-07-26

2026-07-26 · Historical edition

X / Twitter

Josh Woodward, Google Labs VP

Josh Woodward demonstrated a practical Gemini Spark use case: users can drop a school-calendar PDF into Gemini and ask it to add all “No School” dates to Google Calendar. The example emphasizes Gemini completing a concrete cross-application task rather than merely serving as a chat interface. Gemini Spark is now available to all Google AI Pro subscribers in the US, with a global expansion coming next. For builders, PDF-to-calendar workflows are representative because they connect unstructured files, natural-language instructions, and personal tools. They also show AI product competition shifting from “how well does it answer?” to “can it reliably operate the software users already have?”

Boris Cherny, Claude Code

Boris Cherny considers Opus 5 a strong model for coding, data analysis, design, biology, and knowledge work, but is even more excited about its progress against prompt injection. He highlighted an important detail buried in the system card: Opus 5 proved difficult to prompt-inject successfully in PI evals and red teaming. More importantly, combining strong model alignment, prompt injection probes, and Claude Code’s Auto Mode brings the success rate of prompt injection attacks down to approximately 0. Builders should watch this direction because as agents gain the ability to operate real tools, prompt injection becomes a product-level security issue rather than just a model evaluation issue. Boris also suggested that Anthropic will share more details about its layered defenses.

Thibault Sottiaux, Codex & ChatGPT

Thibault Sottiaux announced that ChatGPT Work is now available globally on all paid plans, across mobile, web, and desktop. He described it as giving ChatGPT a jetpack, emphasizing the move from one-off questions and answers toward stronger work execution. Although the post did not detail features, “all paid plans” and “already on your phone” indicate broad distribution to existing paying users rather than a small experiment. For builders, the notable shift is that AI workflow capabilities are moving from desktop IDEs or dedicated agent tools into the general ChatGPT entry point. Mobile availability will also change usage contexts, letting users engage with more persistent work threads from their phones.

Peter Yang, AI Tutorial Author

Peter Yang shared a specific workflow: collaborating with Codex through ChatGPT Voice while lying in bed. He noted that using this approach effectively requires remembering the names of all your long-running threads. It is a small but practical observation: once voice becomes the way to control agents, thread naming, context management, and task resumability become central to the experience. Traditional GUIs let users locate tasks visually in lists; voice mode depends more on accurately naming the relevant context. For builders, this points to a product detail: naming, retrieving, and conversationally invoking long-running tasks may determine whether voice agents are genuinely usable.

Madhu Guru, Meta AI Sr Director

Madhu Guru sees an enormous opportunity over the next few years in connecting messy real workflows to foundation models and making them excel in particular domains. This requires more than prompt tuning: it means understanding how work happens, designing evals, improving models through post-training, and establishing continuous feedback loops. He defines this as the critical path from general-purpose models to domain experts. Today, this capability remains concentrated in a handful of labs, leaving considerable room for outside builders who master the approach. He also added a sharper observation: some people should be required to use AI for writing. Taken together, the posts focus on whether AI can become embedded in production workflows rather than remain limited to isolated content generation.

Cat Wu, Claude Code

Cat Wu emphasized that Claude Opus 5 is well suited to long-running autonomous work and invited users to try it and provide feedback. Her perspective directly concerns Claude Code: maintaining goals, context, and execution quality over extended periods is key to a coding agent’s transition from assistant to actual executor. No specific benchmarks were discussed, but “long-running autonomous work” itself indicates that Opus 5’s selling point extends beyond single-turn problem-solving. For builders, long-task capabilities require attention to checkpoints, recovery, permissions, and failure handling rather than just individual answers. Her remarks also echo Anthropic’s recent overall narrative around Claude Code and agentic work.

Thariq, Claude Code

Thariq said Anthropic removed approximately 80% of Claude Code’s system prompt for its latest model and summarized lessons about system prompts, skills, and Claude.MDs. The number matters because it suggests stronger models may no longer need the lengthy, granular behavioral constraints of earlier models, and may instead benefit from simpler instructions. He also recommended pairing Opus 5 with Fable for planning, brainstorming, or fixing the hardest bugs. For builders, this means a model upgrade is not necessarily a seamless replacement: prompts, skills, and project instruction files may all need redesigning. Thariq added that the material was also published on the Claude Blog, indicating engineering experience Anthropic is systematically documenting rather than merely a personal impression.

Amjad Masad, Replit CEO

Amjad Masad focused on policy positions regarding open-weight models, publicly asking whether Anthropic would sign the relevant initiative and suggesting employees ask leadership whether it supports banning open-weight models. This is a question of the AI ecosystem’s industrial structure rather than model technology: restrictions on open weights directly affect the options available to startups, developers, and downstream products. He also noted that anyone who has not used Replit in a while is in for a major surprise. Together, the posts combine public pressure on open-model policy with a hint at Replit’s product progress. For builders, Amjad’s concern is whether application-layer tools can continue iterating rapidly in an open ecosystem.

Alex Albert, Anthropic Research

Alex Albert shared some of his favorite charts from the Opus 5 release, emphasizing the team’s extensive work on token efficiency across domains. The focus is not simply raising the intelligence bar but completing tasks with fewer tokens while maintaining or improving capabilities. He said Opus 5 feels smooth to use and that he prefers it to Fable 5 for many coding tasks. Token efficiency is a practical product metric for builders because it affects latency, cost, and stability on long-context tasks. Alex’s observation also shows that model experience is determined not only by peak scores but by smooth execution and cost structures across tasks.

Aaron Levie, Box CEO

Aaron Levie discussed both open-weight AI and Claude Opus 5’s performance on enterprise document tasks. He believes open weights accelerate AI’s diffusion through the economy, offer more choices for different customer needs, lower costs for some workloads, and facilitate highly specialized post-training. Box signed the letter supporting open weights, and he emphasized that open versus closed is not a zero-sum war: strong open-weight models move the whole industry forward. In Box’s internal tests, Claude Opus 5 significantly outperformed Opus 4.8 on the Box AI Agent Complex Work Eval, particularly on complex end-to-end enterprise document tasks. Examples include improvements of 17% in due diligence, 30% in life sciences, 12% in legal, 19% in technology, and 13% in healthcare. His conclusion was that Opus 5’s reasoning, analytical, and data-processing capabilities will be powerful for enterprise agentic use cases, with relevant agents soon buildable in Box AI Studio.

Garry Tan, Y Combinator CEO

Garry Tan shared a rehearsal at Chase Center the day before YC Startup School 2026 and welcomed future founders to San Francisco. More notable was his assessment of how quickly AI productivity will spread: macroeconomic productivity gains require managers and CEOs to approve radically different staffing and workflow plans. He believes enterprise leadership has not truly done this yet, so the change may take 10 years rather than 2. This matters for builders because it shifts the adoption bottleneck from model capabilities to organizational authorization and workflow redesign. In other words, available tools do not automatically produce macroeconomic productivity gains unless management is willing to change staffing and process design.

Matt Turck, FirstMark VC

Matt Turck noted an underappreciated irony: the world’s top AI researchers are building recursive auto-research to discover systems that may replace their own jobs. He also called it a major week for model routing, mentioning rumors of Stripe acquiring OpenRouter for $10B, Cursor Router launching on Wednesday, and Runway Router launching yesterday. Databricks, Vercel, Cloudflare, Dataiku, AWS, and Google also have routers, although implementations behind the same term vary substantially. For builders, model routing is moving from an infrastructure detail to a product capability because models’ differing costs, speeds, abilities, and availability require dynamic orchestration. Matt’s observation cautions against treating routers as a single category: they may address different problems in payments, development tools, video generation, cloud platforms, and enterprise data stacks.

Zara Zhang, Builder

Zara Zhang said speed is now the first thing she wants from any model because intelligence is already good enough. She considers waiting 1 to 5 minutes per task the worst interval: too short to enter deep work, too long to do anything but stare at the screen, so users end up scrolling X. She also suggested that bringing agents into chat groups and meetings turns chat logs and meeting transcripts into PRDs. This is an illuminating observation: work culture previously favored strong writers, but verbal communicators now gain equal opportunities because agents do not care whether information originated orally or in writing. She added that San Francisco’s main benefit is that when you attempt something crazy, people consider it admirable rather than crazy. Overall, she is interested in how agents change work rhythms, organizational memory, and startup culture.

Dan Shipper, Every CEO

Dan Shipper offered a cautious Day 0 assessment of Claude Opus 5: it is hard to love. He and the Every team tested coding, writing, knowledge work, and internal agents, finding that Opus 5 argued with instructions, stopped before completing work, and worked poorly with their existing skills and plugins. After they removed old skills and started fresh, performance improved substantially, even showing flashes of brilliance. His central warning is that Opus 5 may break backward compatibility, so teams using existing workflows should be careful. He also said medium or low effort worked better, because more thinking time could increase annoying behavior. His conclusion was that Opus 5 has a Fable-like personality without Fable’s highest ceiling, leaving it in an awkward middle ground in his daily workflow.

Sam Altman

Sam Altman said he wants the US to win in both open-source and proprietary AI models and welcomed progress on that front. Though brief, the post echoes the day’s broader discussion of open-weight models. Its emphasis is competition on both tracks rather than choosing between open and closed. For builders, this position suggests that applications may benefit simultaneously from open models’ customizability and closed frontier models’ peak capabilities. It also connects with Aaron Levie’s, Amjad Masad’s, and others’ discussions of open weights as part of the same industry trend.

Claude

Claude officially announced that Opus 5 is available on all paid plans and through the Claude API at the same price as Opus 4.8. It is the default model for Claude Max and the strongest model on Claude Pro. A Fast mode is also available at approximately 2.5 times the default speed. Claude emphasized that Opus 5 is stronger than Opus 4.8 on cybersecurity tasks but still significantly behind Mythos 5 in exploit development. The safety policy aims to let developers identify and fix software vulnerabilities while blocking high-risk uses. According to automated behavioral audits, Opus 5 is Claude’s most aligned model to date, with the lowest rates of reckless or deceptive behavior relative to other models and the strongest adherence to the Claude Constitution. For builders, these details cover pricing, availability, speed modes, safety boundaries, and behavioral alignment together.

Official Blogs

Anthropic Engineering

Anthropic’s engineering post “How we contain Claude across products” discusses a central change: a year ago, the company would simply refuse to give Claude enough access to affect internal services, but today that access is routine and substantially improves developer productivity. The article separates risk into the likelihood of failure and the damage a failure could cause. As models and safeguards improve, the former is decreasing, but the more systems agents can access, the larger their theoretical blast radius. Anthropic argues that agent safety cannot rely solely on humans in the loop: Claude Code telemetry shows users approve approximately 93% of permission prompts, and more approvals create more oversight fatigue. It therefore promotes Claude Code auto mode, using automated safe approvals to reduce approval fatigue, while acknowledging that probabilistic defenses can never be 100% effective. The article turns to containment: limiting what agents can access, not merely supervising what they do, through process sandboxes, VMs, filesystem boundaries, and egress controls. Anthropic divides agent risks into user misuse, model misbehavior, and external attackers, noting that external sources such as MCP servers, third-party plugins, and web search tools can bring uncontrolled content into context. A key example is that an audited connector does not imply audited data: a GitHub connector may pass security checks yet load a compromised README for the model. The practical lesson for builders is to design permissions, environmental isolation, external-content boundaries, and model-level defenses together, rather than rely solely on user confirmations or the assumption that a model “probably won’t misbehave.”

Podcasts

No Priors — Building an Autonomous Delivery Experience with DoorDash Co-Founders Andy Fang and Stanley Tang

Key takeaway: DoorDash is rebuilding food delivery from an app where “people order and people deliver” into a local fulfillment network driven jointly by agentic commerce, robotics, and a multimodal fleet.

Andy Fang and Stanley Tang are DoorDash co-founders. The material notes that DoorDash has invested in robotics for around eight years, and its network involves 9,000,000 Dashers and 3,000,000,000 deliveries annually. What matters is that they are not discussing adding AI to a search box for smarter recommendations, but letting AI directly change how consumers discover restaurants, buy groceries, replenish supplies, and organize group orders.

The first substantive signal comes from Ask DoorDash. DoorDash initially saw promise in voice as an entry point, but what actually launched was a natural conversational experience: users can express vague needs directly rather than optimize keywords. In restaurant use cases, 50% of journeys using Ask DoorDash lead users to restaurants they have never ordered from before; in groceries, basket size is approximately 40% larger. This suggests AI is not merely a small conversion-boosting component: it may unlock latent demand previously suppressed by search and filtering friction.

The second signal is that agents need more real context. Users can photograph their refrigerator and ask DoorDash to restock it, plan meals with dietary constraints, or prepare ingredients for a weekend family pasta dinner. Andy offered a more specific example: an office pantry shelf, where a camera seeing supplies running low could trigger a DoorDash replenishment request. The key is not “ordering through chat,” but agents gradually participating in real-world inventory, preferences, and routine tasks.

The third signal is DoorDash’s robotics timescale. Stanley said they began studying autonomy and robotics in 2018, before the direction’s potential was obvious. He recalled that DoorDash itself began as an experiment: politoidelivery.com in a Stanford dorm room, with eight PDF menus and a Google Voice phone number. Robotics likewise did not begin with a commitment to building hardware in-house. It started as a skunkworks project involving him and half an engineer’s time, exploring, partnering, and learning before deciding which infrastructure and operational gaps they needed to fill themselves.

The fourth signal is that automation does not necessarily mean fewer human delivery workers. Stanley’s counterintuitive view is: “Ten years from now, we may have more Dashers delivering, not fewer.” His reasoning is that DoorDash is still growing—the material cites 25% year-over-year growth—with ambitions to grow 5x or 10x further. Relying solely on human supply is unrealistic, but expanded demand will not be covered solely by robots either. A multimodal fleet combining human Dashers, DOT, robotics, drones, Waymos, and sidewalk robots is more likely.

The lesson for builders is direct: agentic commerce means redesigning how demand is expressed, context is accessed, actions are automatically triggered, fulfillment is delivered, and costs are structured, rather than replacing an existing app with a chat UI. DoorDash particularly illustrates how, once AI products enter real-world transactions and logistics, the real barriers to entry expand from model capabilities to network data, operational systems, supply-side capabilities, and patience for long-term experimentation.