← Collected sources
BUILDERS · EDITED DIGEST

Builders’ Picks | 2026-07-22

2026-07-22 · Historical edition

X / Twitter

Swyx, AI Builder

Swyx focused on an analysis of trajectory similarity in the RLM paper, looking beyond individual benchmark scores to the gray area between training and evaluation where a model may appear not to have cheated yet be very close to the test distribution. He noted that frontier models can approach a target benchmark by training on “test lookalikes” without directly training on its test set, producing impressive scores. The problem is that open-weight releases generally do not include complete datasets or rlenvs, making it difficult for outsiders to know whether a model trained on data resembling Temu Tbench. He mentioned Alex and Omar’s attempt to apply standard NLP distance metrics to hidden trajectories to examine potential similarities between training and test tasks. He presented this as preliminary exploration, not a final solution. The observation also supports a more positive conclusion: RLMs may genuinely generalize to unseen tasks sharing latent structure. For builders, the reminder is to look beyond leaderboards and ask about data, environments, trajectories, and generalization boundaries.

Peter Yang, AI Tutorial Author

Peter Yang’s first point was direct: banning Chinese models could repeat the “self-own” of banning Chinese EVs. Rather than compare individual models, he warned that excluding competitors from the US market might forfeit broader progress driven by cost, speed, and product pressure. His second post was practical: have one agent perform a task and another review it against a rubric. The example concerned video shorts, where “is this a good short video?” has no deterministic answer that can be checked like a unit test. The approach he cited uses a separate verification agent to read the rubric, review the video, and provide feedback, reducing the self-preferential bias of models checking their own output. The key is separating creator and reviewer, rather than simply adding another model, so an agent does not judge its own work too leniently. This is a transferable evaluation pattern for AI workflows, especially creative, content, and design tasks without a single correct answer.

Madhu Guru, Senior Director at Meta AI

Madhu Guru described the current AI product opportunity as “the best time to have product sense.” He believes the path to AGI is paved with economically valuable tasks rather than abstract accumulation of capabilities. Enterprise AI is therefore one of the most important frontiers, containing many real tasks with budgets and measurable value. He also recast debates about web3 and crypto tokenomics four years ago in terms of today’s new AI tokenomics: open versus closed weights, inference costs, and model routing. The analogy is interesting because tokenomics here concerns the economics of model deployment and task allocation, not token issuance. For AI builders, the central questions become when to use closed frontier models, when to use open weights, and when to switch models to reduce costs. His overall focus is the connection between product judgment and economic value: implementation depends on valuable tasks, manageable costs, and intelligent routing.

Guillermo Rauch, Vercel CEO

Guillermo Rauch’s “big lesson” from AI is that everything is code. Here, code means not narrowly software source but structures that can be described, generated, combined, and automated. In his formulation, a slide deck is code, design is code, a promotional video is code, and Excel automation is code. This reflects AI tools turning more creative and business outputs into executable, iterative representations. For builders, programming ability increasingly affects presentations, design, video, spreadsheets, and operational automation, rather than traditional engineering alone. It also explains coding agents’ expanding boundaries: agents process structured intent, not just repository files. His conclusion is somewhat extreme but directionally clear: products and organizations in the AI era will increasingly manage noncode assets as code.

Aaron Levie, Box CEO

Aaron Levie examined Cursor’s research on multi-model agentic systems. He considers these systems clearly the future, because not every step needs a frontier model’s highest intelligence. Cursor’s research shows that using a frontier model for planning and orchestration while assigning much execution to a cheaper workhorse model can substantially reduce total project token costs, delivering a 15X cost improvement. The central logic he cited is that large tasks require frontier intelligence at relatively few moments, primarily initial decomposition, design decisions, and important tradeoffs. Once the frontier planner converts an ambiguous problem into detailed, explicit instructions, the cheaper model simply follows them. Aaron believes this will become a core design pattern for complex agents because much token consumption happens in phases that do not require the same intelligence threshold. Real application-layer differentiation will come from domain understanding and effective routing across model tiers. High-value fields such as coding, finance, legal services, healthcare, and life sciences are particularly well suited, as lower costs make previously unaffordable workloads deployable.

Zara Zhang, Builder

Zara Zhang proposed a hiring process for the AI era. The first round is an in-person interview without AI, testing domain knowledge and judgment on the spot. The second requires candidates to use AI to complete a project they could not complete without it. Evaluation should examine not only the result but also chat transcripts with agents, revealing how candidates decompose problems, correct errors, orchestrate tools, and judge output. Another post divided companies into those founded before coding agents and now trying to retrofit them, and those founded afterward. She believes the latter are different from day one: teams are usually under ten people because more are unnecessary; work is organized around projects rather than departments; everyone can close their own loop; and internal meetings are almost nonexistent. Her central point is that AI-native companies do more than adopt tools: they rewrite hiring, organizational structure, and collaboration from the roots.

Nikunj Kothari, FPV Ventures Partner

Nikunj Kothari cautioned many founders from the past 18 months: even if traditional “moats” no longer exist in AI, scale and capital do not automatically become the main defenses. He cited Webvan, Groupon, MySpace, Yahoo, AltaVista, Blockbuster, Nokia, and the zombiecorns of 2021 to show that structural position and capital advantages cannot guarantee lasting success. These companies once appeared strong in resources, scale, or market position, yet lost to rivals with better unique insights or collapsed under their own scale. He advised founders to find a distinctive insight worth pursuing for more than 10 years while remaining disciplined enough not to substitute capital and scale for genuine insight. Another observation was an accelerating trend of VCs joining fast-growing startups: he had seen three more people take that path in the previous week. He compared it to investment bankers or consultants moving into BizOps around 2012, noting “Special projects” as a common title for this cohort. Together, the posts concern talent, capital, and company-building discipline during an AI boom: excitement brings resources but cannot replace long-term judgment.

Podcasts

No Priors — Travel Through the Lens of AI with with Booking.com CEO Glenn Fogel

Key takeaway: AI will not automatically erase the travel industry’s complexity. The real opportunity lies in jointly optimizing customer needs, supply-side partnerships, token economics, and the boundaries of human service.

Glenn Fogel is Booking Holdings’ CEO and a veteran of more than twenty years who joined when the company was still Priceline and valued at only a few hundred million dollars. He saw the internet bubble rise and collapse, and watched the company grow from near delisting and a roughly $6 share price after a reverse split into a travel giant that once approached a $180,000,000,000 market capitalization. This background makes his approach to AI restrained: he does not deny disruption but repeatedly emphasizes that no “moat” offers permanent protection, only continual creation of new services and better fulfillment of customer needs.

His most counterintuitive judgment is that outsiders see travel as easy and assume AI agents can replace platforms such as Booking, while the actual business is far more complex. Travel is not simple product search: it involves travelers, hotels, and other partners, with the platform in the middle of a marketplace solving information, supply, payments, fulfillment, service, and trust problems simultaneously. He explicitly cautioned founders seeking to “knock away these very big players” to understand the business before committing capital.

Booking treats AI as a tool for fulfilling its mission rather than something to defend against. Fogel said it can make information more accessible to travelers and bring partners more demand, ultimately making things “cheaper and better for our customers.” Adoption of Priceline’s agentic tool Penny has doubled monthly over the past few months, but he stressed that absolute usage remains small, so growth rates alone should not exaggerate financial impact. Booking handled $186,000,000,000 in travel last year and more than 1,000,000,000 room nights; at that scale, every new tool must demonstrate real ROI.

Cost is one of his main AI concerns. Penny or customer-service AI must be assessed not only for being “cooler,” but by how many tokens it takes to secure a trip, which models are used, how many exchanges users need, and whether long-term lifetime value improves. He explicitly raised today’s token-economics questions: which model should serve which purpose, when should it be used, and can cheaper tokens handle some work? This aligns with current agent-architecture discussions: frontier models need not cover every step.

Customer service is already producing results. Booking’s cost per service contact has fallen and customer satisfaction has risen. AI can prevent peak-time waits for human support and resolve some standard problems faster. Yet Fogel warned against automation for its own sake, since some customers simply want to speak to a person. His principle is straightforward: what matters is always what the customer wants, not a company’s unilateral desire to hand everything to AI.

On AI’s employment impact, he emphasized upskilling. Booking discusses daily how to make employees AI-literate. Even if some roles cannot be fully preserved, workers should have better career options from mastering new tools. His concern is that if businesses and society fail to discuss the transition honestly, fear will cause people to reject technology that could benefit society. He stated plainly that other regions will not stop because of this fear, specifically noting that China will not start from the premise that “AI is bad.”

Fogel’s most memorable line is: “There is no such thing as a moat.” Put another way, nowhere offers permanent shelter from innovation. The lesson for AI builders is neither that incumbents are safe nor that agent startups will inevitably win. Every side must prove itself amid real business complexity through lower costs, better experiences, higher conversion, fewer cancellations, greater customer success, and knowing when to bring humans back into the process.