X / Twitter
Box CEO Aaron Levie
Levie published detailed results from the Box AI Complex Work Eval, saying Fable represents a major leap in both coding and complex knowledge work. The evaluation had a Box AI Agent using Fable tackle a set of real-world enterprise document challenges, then compared its scores with Opus 4.8. Fable’s three main differentiators were not taking shortcuts in complex reasoning, getting multistep calculations right, and delivering substantially greater consistency across runs. By industry, the largest improvements were in Media & Entertainment (78% vs 61%), Technology (81% vs 73%), Financial Services (89% vs 83%), and Healthcare (66% vs 60%). The concrete examples were compelling: Fable scored 100% on a legal M&A due diligence task, compared with Opus’s 78%. In a media profit forecasting task, Fable correctly recognized that a 20% Argentine tax credit was already included in the source spreadsheet and avoided double counting; Opus applied it again, producing cascading errors that drove its score negative. In a 5-year debt financing forecast, Opus calculated interest on the total credit facility and chose the wrong tax base; Fable scored 83% versus Opus’s 62%. Levie called this a step change in complex analysis, analytical reasoning, and deep domain understanding. Fable will soon be available in Box AI Studio for customers to build agents.
https://x.com/levie/status/2064922814688481678
OpenAI Codex and ChatGPT lead Thibault Sottiaux
Sottiaux confirmed a strong spike in Codex token consumption over the past 48 hours, noting that this was unusual because OpenAI had released nothing new during that time. He also posted a one-sentence product philosophy: “Simplify until there is nothing left to simplify.” Separately, he welcomed new team members Clint and Michael, saying he was excited to contribute to cybersecurity together and accelerate defenders’ capabilities worldwide, with the slogan “It’s time to build.” The Codex team currently appears to be pursuing two tracks: enjoying organic usage growth while expanding its talent in cybersecurity.
https://x.com/thsottiaux/status/2064911328087810308
https://x.com/thsottiaux/status/2064900105032135010
https://x.com/thsottiaux/status/2064869401359417799
Builder Zara Zhang
Zara published three dense observations. First, agency deliverables are shifting from one-off assets to “a folder for an agent,” citing the line “earn with your mind, not your hands.” Second, she developed an organizational design idea: every team should build agents or skills for its cross-functional collaborators. For example, a design team could build a design agent for marketing based on brand guidelines and design patterns, allowing marketers to produce on-brand assets themselves without repeatedly interrupting designers. Crucially, only the designers could create that agent, because it requires their expertise and context. She believes this will push organizations from being structured by function toward being structured by loop. Third, she criticized San Francisco’s startup ecosystem: when she asks founders who their target users are, 90% answer “engineering and product teams, AI-native startups.” It feels as though the same small group is bombarded by a million products, while few build for the other 99% of the world. Together, the posts form a coherent view of organizations and markets in the agent era.
https://x.com/zarazhangrui/status/2064843560248332577
https://x.com/zarazhangrui/status/2064835289559023958
https://x.com/zarazhangrui/status/2064825302359150870
Y Combinator CEO Garry Tan
Garry Tan recommended Nessie, calling it the best way to move the context, memories, and history accumulated in ChatGPT, Perplexity, and Gemini to other memory-enabled destinations. It can also import into OpenClaw/Hermes Agent, and he praised its OpenClaw and MCP server work. This reflects an emerging category: portability of personal memory across AI products is becoming a real need. His other two posts continued his usual San Francisco political commentary, celebrating Aaron Peskin becoming an ordinary citizen, declaring “may common sense rule San Francisco for a hundred years,” and continuing to attack the “performative nonprofit industrial complex,” which he said should be uprooted and defunded.
https://x.com/garrytan/status/2064947145652994510
https://x.com/garrytan/status/2064948068076986657
https://x.com/garrytan/status/2064947547735789715
Former product lead for Google Gemini, Veo, and Nano Banana Madhu Guru
Madhu shared a mistake he repeatedly saw while serving enterprise customers in Gemini’s early days: companies get the quality-cost tradeoff backward, choosing the smallest, cheapest model first. His rule of thumb has two cases. If replacing a traditional ML model with an LLM, starting small is fine because you already know what “good” looks like. But when building something entirely new, start with the strongest model, “think magically” first, and discover what is possible. Once a high-quality, usable application exists, help the customer migrate to smaller models while preserving quality. This is a directly usable decision framework for teams selecting models for AI applications.
https://x.com/realmadhuguru/status/2064794601320481150
Anthropic Claude Code team member Thariq
Thariq explained how he used Fable to edit Fable’s own launch video, recording a dedicated explanation because so many people had asked. The core approach was to have the model write extensive code and make tool calls: call transcription services, process footage with ffmpeg, perform color grading, use Figma MCP, write Remotion UI, and render the output. He never touched a video editor throughout. He also shared the deck used in the video, invited people to study it and ask questions, and included a reference video, noting that he had not had time to use the updated designs. This is a complete example of an agent handling a creative workflow: video editing, traditionally highly interactive, was decomposed entirely into code and tool calls.
https://x.com/trq212/status/2064826394589442448
https://x.com/trq212/status/2064826541947940910
https://x.com/trq212/status/2064828193446740023
Claude’s official account
Two updates are worth noting. First, The Problem Solvers series released a story about Cursor co-founder Michael Truell: he fell in love with programming at age 12, and the company he co-founded grew from 15 to 700 people in two years. More than 60% of the Fortune 500 now uses its AI coding platform to build products. Second, Claude Platform added scheduled deployments and environment variable management in the vault starting today, aimed at developers running production agents on the platform.
https://x.com/claudeai/status/2064757537992249734
https://x.com/claudeai/status/2064741184547795408
AI tutorial author and newsletter host with 150K+ subscribers Peter Yang
Peter’s theme this time was “give yourself permission to be a builder.” He criticized traditional career ladders for pushing everyone into management: the higher people climb, the further they move from building, with their time consumed by product reviews, cross-functional alignment, managing upward, and performance calibration. He knows many builders who spent their best years climbing the wrong ladder. The good news is that things are changing: companies reward builders and ICs more than ever, and even managers are increasingly expected to do frontline work. His advice is that becoming a good builder takes substantial practice; if you are a builder at heart, embrace it rather than abandoning what you do well to become a “leader.” Another post described his own experience: the more he uses Codex, the more ambitious his requests become, to the point of wondering whether he is still not ambitious enough.
https://x.com/petergyang/status/2064799855059616172
https://x.com/petergyang/status/2064748427892945313
Vercel CEO Guillermo Rauch
Rauch previewed next week’s Vercel Ship conference in London, hinting at special announcements. In another post, he defended Silicon Valley: in his angel investing, he receives introductions to all kinds of people, from “two guys and a dog” to a serial founder with five awards, and treats them equally seriously. The future here is open to everyone, he argued, and nowhere is more meritocratic. For people following the Vercel ecosystem, next week’s Ship London is worth watching.
https://x.com/rauchg/status/2064777495422161205
https://x.com/rauchg/status/2064732935484514729
Every CEO Dan Shipper
Dan revisited his prediction on Lenny’s podcast last year: once AI substantially improves each employee’s productivity, reshoring some jobs to the United States to be closer to customers becomes attractive. A current news story supports that judgment. This is a second-order effect worth remembering: AI does not merely replace labor; it can change the logic of labor’s geographic distribution. When output per person is high enough, differences in labor costs matter less, and proximity to customers becomes more valuable.
https://x.com/danshipper/status/2064777216656097445
FirstMark Capital Partner Matt Turck
Turck used a long post to mock the “brutal” schedule of a VC in 2026: start the year at Davos, freeze in Aspen, rush to Upfront, survive Milken, head straight to Paris for the French Open, return to NYC for the Knicks, then SuperReturn in Berlin, Founders Forum in London, the World Cup, Raise AI in Paris, Sun Valley in Idaho, a break in Mykonos, a marathon of Goldman technology conferences, Slush in Finland, and finally NeurIPS far away in Sydney. After a year of “productive thought leadership and value creation,” one is exhausted. Behind the joke is a precise satire of venture capital’s conference and networking culture: a packed calendar does not necessarily correlate with actual value creation.
https://x.com/mattturck/status/2064806681612362113
Google Labs’ official account
Project Genie’s availability is expanding further: starting today, users worldwide on Google AI Ultra 5X, Google’s newest subscription tier, can access it. Google is using its highest subscription tier as a distribution channel for frontier experimental products, making this product line’s commercialization path worth watching.
https://x.com/GoogleLabs/status/2064801929339752527
Google Labs and Gemini App VP Josh Woodward
Josh handled a Gemini outage by first acknowledging that the service was down, the team was working on repairs, and some fixes were already live. He later confirmed full recovery and apologized again. The communication pattern from announcement to resolution was standard and useful for consumer AI teams: acknowledge the problem promptly, provide progress, and close with an apology.
https://x.com/joshwoodward/status/2064762269674918013
https://x.com/joshwoodward/status/2064869366290841716
Anthropic Claude Code team member Boris Cherny
Boris sent greetings from Code with Claude Tokyo. Anthropic’s developer event series is expanding into Asia, with the Japan stop underway.
https://x.com/bcherny/status/2064885111477219664
Podcasts
AI & I by Every — How Anthropic Uses Claude Fable 5 With Mike Krieger
Key takeaway: Fable 5 turns AI from a “pair-programming colleague” into a “teammate trusted with complex tasks overnight.” The real bottleneck has moved from model capabilities to human methods of verification and work organization.
Mike Krieger is head of Anthropic Labs and an Instagram co-founder who has just moved from CPO back to a frontline builder role. On the eve of Fable 5’s launch, he shared firsthand experiences from months of internal use, directly relevant to anyone who uses coding agents heavily.
His first impression was that his skills had become outdated. “I felt like a complete beginner because the way I prompt and break down tasks was already obsolete for this model.” His routine now is to give Claude a complex task before bed, usually completed around two in the morning. He was particularly impressed by its behavior when blocked: if a remote service fails, it builds a scaffold backend itself and keeps going, records that fact, and fixes things when the service returns. Even on a plane, he does not worry about Wi-Fi disconnecting; once contextual instructions are set, he is comfortable delegating.
Workflow changes include more architectural planning conversations upfront, often asking the model to produce an HTML page or a document with diagrams directly to align the team. Concurrent sessions have increased sharply: either one long session with background subagents, or five or six tabs running long tasks. He has begun adjusting effort levels for the first time, reducing small UI changes to medium because Fable’s capability range is much broader than Opus’s. For quick questions, he switches to Sonnet, joking, “This is not a question worthy of Fable.”
His cost framework is multidimensional: beyond the cost of a single turn, consider the total cost of “completion to your satisfaction.” Fable is expensive, but it gets things right the first time, avoiding nine or ten subsequent rounds of “No, that’s not what I meant.” A personal media tracking app he built over a weekend, which supports agents modifying the software itself, used only a little extra allowance. Compare Instagram v1: back then, even as a skilled engineer, he needed five consecutive all-nighters plus Kevin Systrom’s work on filters to launch it. Now he can complete work of a similar scale over a weekend between childcare responsibilities.
He believes software engineering is not dead but redefined. The coding part is largely solved, leaving intent expression, architectural judgment, verification, and accountability for results. Verification is the next frontier he repeatedly emphasized: attach a complete screenshot set or even a video to every PR. During prototyping, he gave Claude ffmpeg and had it inspect its own animation frame by frame for stutters. The user story that moved him most came from an internal recruiting colleague: “For the first time in my life, I feel that what is in my head and what actually exists in the world are right next to each other. I can just make it.”
The lesson for readers is that both the ceiling and floor of model capabilities have risen. The complexity of what nontechnical people can build is increasing quickly, and the truly scarce abilities are clearly expressing intent, designing verification loops, and exercising judgment while taking responsibility for deliverables.