← Collected sources
BUILDERS · EDITED DIGEST

Builders’ Picks | 2026-08-22

2026-08-22 · Historical edition

X / Twitter

AI Builder Swyx

Swyx believes calling simulation the new scaling law is more than marketing. As models increasingly automate ML research and AI engineering, the harder barrier will be simulating humans and their feedback. This led him to reconsider why Karpathy, Fei-Fei Li, and others backed the Smallville team, even though Smallville had almost no commercial applications at the time. Simile is now beginning to find PMF among Fortune 100 companies, demonstrating that simulation can enter real business use. He admitted it took him two extra years to recognize the importance of this direction and was glad the evidence had overturned his earlier judgment.

Thibault Sottiaux, OpenAI Codex and ChatGPT Team

Thibault Sottiaux provided an update on the investigation into Codex rate limits. Some users' cache hit rates this week were below the stable levels seen in preceding weeks, potentially causing the same work to consume allowances more quickly. Consistent cache hits are important to efficient Codex usage; the team is investigating the specific cause and plans another update the following day. He also announced that all paid ChatGPT Work and Codex users would receive a banked reset by 8 p.m. Pacific Time. Together, these updates indicate that recent unusually rapid allowance consumption may partly reflect changes in system efficiency, rather than solely increased user workloads.

Peter Yang, Practical AI Tutorial Creator

After trying Instinct, Peter Yang praised its onboarding flow for connecting iMessages, Google Workspace, and MCP. Once MCP is connected, Instinct proactively suggests tasks it can perform, making it more proactive than some comparable assistants. But having all work in a single thread made it difficult for him to advance multiple substantive tasks in parallel. He preferred Instinct for miscellaneous chores and continued choosing ChatGPT Work and Codex for actual work. He later discovered that Instinct indexed and retained emails without permission and provided no way to delete the records, so he explicitly said he would not recommend the product until this was resolved. The case illustrates that smooth onboarding and proactive execution cannot compensate for deficiencies in data control and deletion rights.

Madhu Guru, Senior Director at Meta AI

In part five of his eval series, Madhu Guru argued against compressing complex evaluations into a single aggregate score. Drawing on his experience with early Gemini, he explained how management's pursuit of one number to simplify decisions can produce the "tyranny of averages." A model might improve from 85% to 89% on simple summarization and from 80% to 85% on basic factual questions, yet fall from 70% to 63% on complex financial analysis. An aggregate score obscures regressions in critical frontier use cases, while weighted scores merely give subjective judgments a false sense of mathematical precision. He recommends maintaining a prioritized list of evals and involving people capable of inspecting the details in decisions. The real question is not whether a new model's average score is higher, but where it fails, where it excels, and whether those changes matter to target users.

Thariq, Anthropic Claude Code Team

Thariq shared the ELI5 skill, which has recently seen frequent internal use at Anthropic. Users can enter `/eli5 <需要解释的内容>` to have Claude explain a subject to a reader with no background knowledge through large diagrams, minimal text, and an HTML artifact. He uses it to understand how a module works, why a technical trade-off was made, or the specific cause of an incident. This makes ELI5 useful not just for simplifying concepts but also for codebase exploration and incident reviews. Whether it will become an official plugin remains undecided, but users can currently try it by adding the marketplace with `claude plugin marketplace add anthropics/claude-plugins-community`, then installing it with `claude plugin install eli5@claude-community`.

Aaron Levie, Box CEO

Aaron Levie believes AI is progressing at a pace unlike any other period in technology history. For the same task, model costs are falling while breadth of capability, speed, and depth in vertical domains all improve. As intelligence becomes extremely cheap, the greatest opportunity shifts from simply improving models to driving AI throughout the economy. This diffusion creates favorable conditions for applied AI companies because upstream model innovation and competition continually support their growth. Startups able to rapidly absorb these capabilities and apply them to specific industry problems are in a rare window of opportunity.

Nikunj Kothari, FPV Ventures Partner

Nikunj Kothari used a home automation example to demonstrate the combined value of Claude Code and existing agents. His daughter's preschool posts daily meals on a poorly structured website without a convenient standard interface for parents. He had Claude Code use network requests to locate the underlying API and discovered that it happened not to require authentication. Claude Code then identified the correct data format and connected the results to the household's existing Hermes bot. Home bot now announces breakfast and lunch each morning so the family can decide whether to prepare extra food. The value of this small example lies in a coding agent finding a hidden interface and connecting real-world unstructured information to an existing automation workflow, rather than generating a new app.

Claude, Anthropic's AI Assistant

Anthropic announced a set of security capabilities centered on Claude Security and Mythos 5. Users can have Claude Security scan a GitHub repository, trace data flows across files, and analyze interactions between components. Each finding includes a CWE classification, confidence, severity, and a suggested fix; proposed patches can be opened directly in Claude Code on the web. Mythos operates behind the scanning process and returns only security findings, so defensive teams do not need direct model access. Charges use standard token usage within existing plans. Anthropic is also working with partners to integrate Mythos 5 into security products and services and is providing $35 million in credits for open-source security through the new Defender Advantage Fund. The Cyber Verification Program is also expected to expand in the coming weeks to bring frontier security capabilities to more defenders.

Podcasts

No Priors — What Chess.com Teaches US About Superhuman Capabilities, with CEO Erik Allebest

Key takeaway: With machines having long surpassed humans at chess, Chess.com's continued growth shows that superhuman AI need not eliminate a human activity; it can instead make learning, competition, and measuring one's own abilities more appealing.

In 2005, Chess.com CEO Erik Allebest and a friend bought the chess.com domain at a bankruptcy auction for $56,000, initially intending to build a MySpace-like chess community. Most investors viewed it as an uninvestable niche project, so the team raised no funding. After launching in 2007, it grew on membership revenue, beginning to charge approximately 18 months later and quickly becoming profitable. The company now has approximately 10 million daily active users, 40 million to 50 million monthly active users, and more than 250 million registered users, with projected annual revenue slightly above $200 million and a team of approximately 650.

Growth was not a smooth curve. COVID, The Queen's Gambit, and news events drove the first surge, and the team worried it might be a short-lived craze like treadmills and sourdough. In 2023, schoolchildren joined in large numbers, spurred by short videos, the Mittens bot, and cheating controversies. After traffic peaks subsided, the underlying user base remained substantially larger than before, showing that cultural events can convert temporary attention into lasting habits.

Allebest's advice to founders is almost the opposite of the mainstream startup playbook. Chess.com had no technical co-founder from an elite university, no office, no funding, and no paid acquisition, and chose a market widely considered tiny at the time. In his words: "Stop listening to people telling you what to do. Just go do it." The key is to be clear about what you want to exist in the world, prove that vision with minimal resources, and keep creating something distinctive and valuable to users.

AI is already used for customer support automation, data analysis, product specifications, and agentic development. Chess.com has also built GNS, an internal system with authentication, knowledge, and permission controls, aiming to shorten the cycle from identifying a problem to shipping code. On the product side, it is exploring a portable AI coach that compares users' historical data with players at similar or higher levels and proactively offers feedback based on the past week's games. It will not replace human coaches for those needing deep guidance, but can provide lightweight coaching available at any time.

Chess.com is also preparing to extend these capabilities into poker and other classic games. Its poker rating considers opponents' skill, chips won, and the number of hands played, attempting to measure "how good you really are," rather than who can keep raising or buy more chips. Allebest believes players may eventually care about ratings as much as money because ratings more directly reflect personal skill.

On AGI and ASI, he believes superhuman intelligence will develop more rapidly, but outcomes depend on guardrails, the distribution of power, and social culture. He worries about further concentration of wealth and control while hoping AI helps address climate, education, and political problems. Technology itself does not automatically lead to good or evil; what determines outcomes is how those who control it take responsibility and share the benefits.