X / Twitter
Thibault Sottiaux, OpenAI Codex and ChatGPT
Thibault Sottiaux highlighted the change in experience when using ChatGPT Work on mobile, saying it is now available in the ChatGPT app and that mobile use will be a game changer. The central point is that workflow-oriented AI products are no longer confined to desktops and can enter more fragmented, spontaneous usage contexts. He added that ChatGPT Work’s active user count has officially overtaken Codex, indicating broader adoption of OpenAI’s work-oriented product line. Builders should watch this signal: Codex represents developer-focused agentic coding workflows, while ChatGPT Work leans toward general knowledge work and enterprise collaboration. If the latter is growing faster, AI workflows may spread first through noncoding use cases. No specific user counts were disclosed, but “active users overtaken Codex” is itself a signal of an adoption trend.
https://x.com/thsottiaux/status/2081229262452097169
Peter Yang, AI Tutorial Creator
Peter Yang previewed a Codex interview with OpenAI DevEx engineer Jason, author of the official Codex working guide. The most valuable detail was that Jason demonstrated a complete Codex work system rather than isolated tips. This includes using Codex for chief-of-staff-style coordination across Slack and email, turning past sessions into new skills and workflows, and building websites for learning any subject, such as drumming. Peter’s perspective remains “practical AI skills for busy people,” placing Codex in real workflows rather than discussing code generation alone. He also recommended Kun’s model analysis, calling it one of the best he had seen, without elaborating on its arguments. Overall, Peter’s focus today was how Codex is expanding from a development tool into a personal work system.
https://x.com/petergyang/status/2081029209993154980
Madhu Guru, Senior Director at Meta AI
Madhu Guru discussed why the US AI community has rapidly shifted toward supporting open-weight models. He argued that in “uncharted territory,” answers to complex questions can emerge only through repeated contact with reality, rather than a single round of abstract reasoning. Less than a month ago, industry consensus on open weights was much less apparent, but a series of public events has allowed everyone to observe the consequences of different directions. He cited DeepSeek, the Microsoft-OpenAI breakup, GLM, Kimi, Fable, and the OpenAI-Hugging Face episode, arguing that these events revealed changes in incentives, innovation paths, geopolitics, and the leverage held by different players. His focus was not taking sides but emphasizing that AI’s social questions require public experimentation, observation of first- and second-order effects, and subsequent updates to beliefs. This also means that many AI governance and industry questions over the next decade may be resolved through dynamic consensus built on continuous real-world feedback, rather than once and for all through a priori principles.
https://x.com/realmadhuguru/status/2081141594892415028
Amjad Masad, Replit CEO
Amjad Masad released Replit’s newly deployed chess engine, estimating its current strength at nearly 1200 Elo. His target is 2000+ Elo, with two key constraints: use only a small fine-tuned LLM, with no custom pretraining or architecture, and have the model generate moves directly without chess-engine assistance. The experiment’s value is not in achieving the strongest possible play, since relaxing these constraints would make the problem much easier. Instead, it tests how far a small model can go in an environment with clear rules, long-range planning, and high costs for mistakes. Amjad explicitly stressed that it would be much easier if the constraints were relaxed, showing his interest in the model’s own capabilities rather than system capabilities augmented by external search or a traditional engine. Another comment about information overload being especially unfriendly to an ADHD brain lacked enough information and is skipped here.
https://x.com/amasad/status/2081086837263937543
Guillermo Rauch, Vercel CEO
Guillermo Rauch elevated agent workflows to the level of a “software factory” today. He argued that https://v0.dev is more fundamental than any framework Vercel has built, because it represents the starting point of an idea for a company. For a new idea, he recommends going beyond “find an agent and casually prompt it” and asking how to build a factory that can launch, maintain, and grow the idea. He added that in software, the “factory is the product”: product quality depends on which agents you set up to maintain it automatically. He also shared his research workflow: doing all research with agent CLIs and the filesystem, maintaining a research/ folder, and specifying desired formats and best practices in AGENTS.md. The system needs no fancy apps, knowledge graphs, or UIs; agents search and connect knowledge from past sessions, then generate an HTML report and deploy it to Vercel when sharing is needed. His conclusion is explicit: the core driving this “software” is the English instructions in AGENTS.md and the organization of the filesystem.
https://x.com/rauchg/status/2081149743368122723
Aaron Levie, Box CEO
Aaron Levie interpreted Google’s participation as a complete endorsement of open-weight AI. He considered it an important industry moment because open weights are no longer just a choice made by individual open-source communities or model companies, but have won recognition from larger platform players. Although he did not specify what Google had joined, the phrase “complete endorsement” shows that he views it as a directional signal. Alongside other builders’ discussions today, the post reflects open weights’ transition from a contentious approach to a mainstream industry option. For enterprise software companies, open weights mean more than downloadable models: they also concern deployment control, cost structures, security boundaries, and bargaining power in the ecosystem. Aaron’s focus remains long-term infrastructure choices for enterprise AI.
https://x.com/levie/status/2081054531908247937
Matt Turck, VC at FirstMark Capital
Matt Turck shared a discussion with Andrew Feldman on “chip landscape 101,” covering CPUs, GPUs, NVIDIA, AMD, TPUs, Trainium, Cerebras, and other topics. This is suited to builders who need to fill gaps in their AI compute fundamentals, since competition across models, agents, and applications is ultimately shaped by chip supply, costs, and architectural choices. Matt did not elaborate on specific arguments in the post, but the keywords indicate a focus on explaining the AI chip ecosystem from first principles. CPUs, GPUs, TPUs, Trainium, and Cerebras represent different approaches, including general-purpose computing, mainstream acceleration, cloud providers’ in-house accelerators, and wafer-scale computing. For AI infrastructure readers, this kind of “landscape 101” helps explain why model companies, cloud providers, and hardware startups make different capital expenditure and performance tradeoffs. Today’s post is more of an entry point for developing deeper judgments later.
https://x.com/mattturck/status/2081131761686184333
Zara Zhang, Builder
Zara Zhang offered a brief but substantive observation: AI-native companies have cultures more like open-source communities. This points to an organizational shift, with internal collaboration potentially becoming more open, iterative, and reliant on shared context and voluntary contributions than traditional top-down divisions of labor. She gave no examples, but the analogy between “AI-native” and “open-source community” is notable because AI tools lower the barriers for individuals to move from idea to prototype to release. Such companies may place greater weight on transparent workflows, reusable artifacts, shared prompts or agent configurations, and cross-functional team members building directly. Another post asking what people do while waiting for AI output was primarily conversational and less substantive. Overall, Zara’s focus today was AI’s effect on company culture and work rhythms rather than a single tool update.
https://x.com/zarazhangrui/status/2081223709755650054
Nikunj Kothari, FPV Ventures Partner
Nikunj Kothari commented on an unconventional acquisition, arguing that such events usually require two conditions to hold simultaneously. First, the CEO must have complete control of the company—for example, it is profitable, has no board, and has no VC constraints that he knows of. Second, the founder must be as ambitious and sufficiently crazy as David Holz. Otherwise, it is hard to explain to a team why a generative media company should acquire an astrology app. He also mentioned “scanners,” suggesting that he believes the company is already conducting wide-ranging experiments that would be difficult to approve under conventional governance. Nikunj’s point was not whether the acquisition was right or wrong, but that organizational control and founder personality determine whether a company can take nonconsensus actions. For builders, the lesson is that some seemingly outlandish strategic moves are enabled less by ordinary product logic than by governance structures giving founders sufficient room to act.
https://x.com/nikunj/status/2081017328137916426
Peter Steinberger, OpenAI and OpenClaw
Peter Steinberger shared a workflow for large-scale parallel QA with Codex, focused on preparing OpenClaw’s next release. His instructions to the agent were specific: run full end-to-end QA tests with live API keys, use 12 subagents to divide up features, start development gateways on different ports, and assign some agents to stress tests. The task also required worktrees and automatic PR creation, set a target of finding 200 bugs, and demanded root-cause fixes rather than band-aids. The boundaries were equally clear: refactors were permitted, but the plugin SDK boundary must not be touched, and the test report had to be continuously updated in a Markdown file on the desktop. He added that Codex had been running massive parallel QA all day and that Sol had become noticeably stronger at understanding intent and identifying complex behavioral issues. Previously, workflows like this often broke at compaction boundaries or the model began cheating. This progress suggests improving reliability in long-running agent QA.
https://x.com/steipete/status/2081169376317932017
Official Blogs
Anthropic Engineering An update on recent Claude Code quality reports
Anthropic explained why some users had reported declining Claude quality over the past month, explicitly stating that the API and inference layer were unaffected. The issues came from three independent changes affecting Claude Code, Claude Agent SDK, and Claude Cowork, all fixed in v2.1.116 on April 20. The first issue was a March 4 change to Claude Code’s default reasoning effort from high to medium to reduce excessive latency and token use at high effort. Users said they preferred higher intelligence by default and to lower effort themselves for simple tasks, so the change was reversed on April 7. The second issue was a March 26 caching optimization intended to clear old thinking after a session had been idle for over an hour, reducing resumption costs. A bug instead continued dropping old reasoning on every subsequent turn, making Claude appear forgetful, repetitive, and unusual in its tool choices. The third issue was an April 16 system-prompt instruction to reduce verbosity, which combined with other prompt changes to harm coding quality and was reversed on April 20. Because the three issues affected different traffic segments and time periods, they appeared to be a broad, inconsistent overall degradation. Anthropic also said it would reset usage limits for all subscribers.
https://www.anthropic.com/engineering/april-23-postmortem
Anthropic Engineering Scaling Managed Agents: Decoupling the brain from the hands
Anthropic used Managed Agents to explain its approach to long-running agent harnesses: separate the brain, hands, and session. The initial design placed the session, agent harness, and sandbox in one container. File edits were direct syscalls and there were no service boundaries, but the container became a “pet” that could not be lost. If it failed, the session was lost; if it stalled, engineers could only infer the cause from a WebSocket event stream, which could not distinguish a harness bug from dropped events or an offline container. The new design moves the harness out of the container and calls the sandbox like an ordinary tool: execute(name, input) → string. A container failure is then just a tool-call error, and it can be reprovisioned. The session log is also independent of the harness. After a crash, the harness can use wake(sessionId), read the event log with getSession(id), and continue from the last event. Security boundaries are clearer too: untrusted code generated by Claude no longer runs in the same sandbox as credentials. A Git token can be attached to the remote when the sandbox initializes, while MCP OAuth tokens remain in a secure vault outside the sandbox. The central judgment is that assumptions about model weaknesses embedded in agent harnesses expire, so interfaces should stay stable while implementations remain replaceable.
https://www.anthropic.com/engineering/managed-agents
Claude Blog New in Claude Managed Agents: self-hosted sandboxes and MCP tunnels
Claude Managed Agents has added self-hosted sandboxes and MCP tunnels, allowing agent tool execution and private-service access to remain within enterprises’ own boundaries. Self-hosted sandboxes let Claude Managed Agents execute tools on user-controlled infrastructure or through managed providers such as Cloudflare, Daytona, Modal, and Vercel. Anthropic still handles the agent loop, orchestration, context management, and error recovery, while code execution, sensitive files, packages, services, and data stay within the enterprise perimeter. Cloudflare emphasizes microVMs, lightweight isolates, zero-trust secrets injection, and auditable, rewritable egress. Daytona emphasizes long-running, stateful environments accessible through SSH or authenticated preview URLs, with state that can be paused and resumed. Modal emphasizes AI workloads, sub-second startup, CPU/GPU resources on demand, and large-scale concurrency. Vercel emphasizes VM security, VPC peering, bring your own cloud, and millisecond startup. MCP tunnels let Managed Agents connect to MCP servers on corporate networks without exposing internal databases, private APIs, knowledge bases, or ticketing systems to the public internet. A user-deployed lightweight gateway initiates a single outbound connection, requiring no inbound firewall rules or public endpoints and providing end-to-end encryption. Self-hosted sandboxes are in public beta on Claude Platform. MCP tunnels are a research preview requiring an access request.
https://claude.com/blog/claude-managed-agents-updates
Podcasts
Unsupervised Learning — Ep 91: Top AI Analyst Unpacks Today's AI Hype Cycle
Key takeaway: Benedict Evans rejects treating AI as a myth that is “unprecedented and therefore incomparable.” He is more interested in how the industry structures of earlier technology platforms carry over to AI.
Benedict Evans is an analyst who has long studied changes in technology platforms, influencing founders, investors, and product teams through his newsletter and presentations. Jacob Efron framed the discussion around the AI hype cycle, foundation-model valuations, consumer AI applications, enterprise adoption, and employment disruption. Benedict’s method remained consistent: rather than argue how much “bigger” AI is than the internet, mobile, or the Industrial Revolution, look back at previous world-changing technological shifts and ask how value was distributed across the stack, which layers had pricing power, and which merely bore the capital expenditure.
His most counterintuitive reminder is that analogies are for asking better questions, not making predictions. For example, semiconductors have a structure in which cutting-edge fab costs double every four years; foundation models have a similarly escalating cost structure in some respects. Mobile networks also have marginal costs, requiring more base stations and equipment as traffic grows. Over the past fifteen years, mobile data traffic has increased a thousand- to two-thousandfold. Mobile carriers are a trillion-dollar industry spending $200,000,000,000 in annual CapEx, yet the value in Uber, banking, and YouTube accrued to higher layers. AI faces a similar question: investment in the model layer is enormous, but whether more of the ultimate profit flows to applications, workflows, industry systems, or new platforms remains unsettled.
He is also skeptical that “LLMs will be like Windows.” Sam Altman has said OpenAI wants to be like Windows, but Benedict noted that LLMs have no obvious network effects, so their competitive dynamics may differ from operating systems. Better questions are whether different layers of the stack can compete upward or integrate downward, and how token pricing, compute supply, and product entry points will alter the industry’s division of labor.
The difficulty with consumer AI is that many products have not yet become concrete, recurring behaviors that solve problems. Comparing Midjourney to a drone, he noted that something can be exciting without being a strong product if it does not translate into a tangible, specific thing users repeatedly want to do. He offered virtual try-ons as an example: a product that reads a user’s Instagram and a brand’s lookbook to generate ten videos of different outfits could be a specific use case, rather than simply “generating images.”
He also opposed viewing AI solely as text and pictures. He said: “It’s not chat rooms, it’s the network; it’s the same here—not text and pictures, but an enabling technology.” He means that AI will enter many places users do not see, including fraud detection, enterprise back-office processes, and everyday service experiences. These improvements will not announce themselves as “this is ChatGPT.” Judging AI by one inaccurate question-and-answer exchange resembles seeing a slow internet and chaotic chat rooms in 1995 and misjudging the network itself.
His advice to AI labs and young builders was more measured: assume that many current acronyms, protocols, companies, and playbooks will fail. Many concepts from 1995 and 2010 later disappeared, but trying them was still valuable. He even mentioned OpenAI abandoning a web browser and the possibility of MCP being replaced, underscoring the current radical uncertainty and the danger of sanctifying any interface or product form too early.
Enterprise adoption will not explode linearly as social media discussions suggest. Citing a Goldman CIO survey, Benedict said that even after twenty years, only around 30% of enterprise workflows have moved to the cloud, because large companies face realities of permissions, compliance, integration, priorities, and organization. AI is not the only variable in every industry. It may structurally disrupt newspapers, while its impact on Caterpillar or an aggregates company may be more limited or indirect. The lesson for builders is to ask not only how powerful models will become, but whether work, costs, risks, distribution, and organizational structures in a specific industry will actually be reordered as a result.