X / Twitter
Swyx, Builder in Intent, Intensity, and the AI Engineering Community
Swyx continued focusing on AEO, or AI Engine Optimization. He argued that anyone not having Codex, Claude, Gemini, or Devin automatically research SEO / AEO improvements weekly is missing an opportunity that remains “free,” should be commoditized, and is still underused. The focus is not traditional search ranking but how products, content, and companies are discovered, cited, and recommended by AI systems. He raised a finer question: if Claude optimizes AEO, will it favor visibility within Claude rather than general visibility across models? This moves AEO from “how does AI see you?” to “which models see you, and does optimization overfit a particular ecosystem?” For builders, it suggests a new automated workflow: agents periodically researching, testing, and updating visibility strategies for AI search and answer systems.
https://x.com/swyx/status/2078244735794413786
Thibault Sottiaux, OpenAI Codex & ChatGPT
Thibault Sottiaux announced a usage-limit reset for paid Codex and ChatGPT Work users, attributing the adjustment to rapid team iteration while supporting fast-growing infrastructure demand. For heavy Codex users, such resets directly affect whether they can continue long tasks, debug agent workflows, and experiment with code over the weekend. He also relayed that GPT-5.6 Sol had been confirmed as an “extremely good model,” without further evaluation details or use cases. Together, the posts signal operational adjustments to both model experience and allowances. The actionable point is that paying users can recheck their available Codex and ChatGPT Work capacity and reschedule tasks paused by limits.
https://x.com/thsottiaux/status/2078320950488297917
Peter Yang, AI Tutorial and Interview Creator
Peter Yang said Codex browser use had “finally been defeated,” but the material gives no mechanism, so this cannot be interpreted as a security bypass, capability failure, or solved test case. More substantively, he argued that staring at a screen managing agents all day is exhausting. He would prefer assigning work outdoors through voice, as if on a phone call, and receiving spoken status reports. This shifts the bottleneck from task completion to how humans manage multiple workstreams with little effort. He also previewed a weekend episode with Thariq about AI video workflows and more practical demonstrations. For AI builders, voice, status reporting, and asynchronous agent management may become important interaction layers in the next wave of workflow products.
https://x.com/petergyang/status/2078276992470794531
Madhu Guru, Senior Director at Meta AI
Madhu Guru challenged the simple narrative that “Kimi will hurt Google.” Many enterprises will obtain Kimi through platforms such as Google Cloud rather than use it directly, he argued, because they still need security, data residency, compliance, and especially chips. Model competition may therefore shift demand between a cloud platform’s revenue streams rather than merely redistribute revenue among model companies. He also said enterprises struggle to move beyond basic chatbots chiefly because they lack people who can build harnesses and evals. Evals must express real use cases, cover offline and online settings, and help select models along the quality-cost-latency curve. Harnesses should be model-independent, handling routing, multi-agent orchestration, context management, tool calling, and memory. He identified talent as the scarcest element: the hard part of frontier systems is not connecting a model but building evaluation, orchestration, and runtime capabilities.
https://x.com/realmadhuguru/status/2078131628262752550
Thariq, Anthropic Claude Code
Thariq’s practical advice was direct: build mockups, schemas, data models, proofs of concept, and other prototypes before consuming large amounts of tokens. Prototypes reveal earlier whether a direction deserves further work, rather than letting a model generate substantial output before discovering it is unwanted. For Claude Code and similar coding-agent workflows, this means reducing uncertainty before scaling generation. A low-fidelity schema or mockup often constrains the output space better than a long prompt and helps humans judge product form, data structures, and interaction logic sooner. Teams can turn this into an agent-use convention: begin complex tasks with a small inspectable sample before moving to large-scale implementation.
https://x.com/trq212/status/2078189833445654714
Amjad Masad, Replit CEO
Amjad Masad reposted and praised a chess-history exploration project, saying the Replit community is “ChessMaxxing.” The material gives no specific features, implementation, or content, so it cannot support a full product review. What can be retained is that the community is building knowledge-exploration or interactive-narrative projects, not only traditional utility apps. For builders, such examples show accessible development platforms combining history, games, and visual exploration. His interest continues Replit’s community-creation focus: helping users quickly turn personal interests into working projects.
https://x.com/amasad/status/2078273728618877326
Guillermo Rauch, Vercel CEO
Guillermo Rauch announced “Sandbox data for downloads,” encouraging people to keep shipping more agents. The material does not provide a full product explanation, but clearly concerns download data and agent-building scenarios. For Vercel Sandbox developers, it suggests improved observability or availability around download data. It also echoes a broader theme: as agents execute more real tasks, sandbox runtimes, data artifacts, and download capabilities must become more reliable. His emphasis is on Sandbox as infrastructure for shipping agents rather than on an individual demo.
https://x.com/rauchg/status/2078305023784620342
Aaron Levie, Box CEO
Aaron Levie discussed falling AI costs across the ecosystem. Cheaper AI creates opportunities for every layer and end customers because the real bottleneck is successful, affordable deployment on actual workloads. Lower costs generally raise overall usage, spreading value across the stack rather than to one model provider. He added an important qualification: even if cheaper, more optimized models consume many tokens, demand for closed frontier models may keep rising. Complex-task orchestration often still needs the strongest model, while bulk execution can use cheaper or specialized models. Efficiency may actually increase frontier spending by making more tasks worthwhile for AI. Margins may face the real pressure; he expects intelligence margins eventually to approach those of the infrastructure stack.
https://x.com/levie/status/2078139206946459853
Zara Zhang, Builder
Zara Zhang’s advice for building in public is to show work already happening inside the product rather than treat content creation as extra work: a short screen recording, an initial prototype, or how user behavior changed a design. Reasoning matters more than production value, she emphasized, making this useful for early builders who need not package content as an expensive launch announcement. She also observed that although many people were uncomfortable with meeting recordings a few years ago, recorded business meetings are now nearly the default, serving not only humans but agents. She sees this as technology changing culture. Together, the posts concern making product work naturally visible and agents changing default organizational communication habits.
https://x.com/zarazhangrui/status/2078086930756202924
Peter Steinberger, OpenClaw + OpenAI
Peter Steinberger shared Codex using browser and computer control to open Chrome, navigate to a GitHub PR, click a comment, and handle the macOS picker to upload an image. He found it both amazing and painful because the whole sequence served merely to upload one picture, yet “GitHub has no API” did not stop the agent from trying. It is a representative case: agents fall back to real GUI interaction when APIs are unavailable, making reliability and costs harder to control. He said he runs Codex in VMs so it does not steal local app focus. Another update described having Codex build an editor after numerous codexbar icon-customization problems. For tool developers, both show agents handling supporting development and GUI operations while “noncode details” such as isolation, focus management, and system pickers become essential to practical usability.
https://x.com/steipete/status/2078318731785359634
Claude, Anthropic AI Assistant
Claude explained subscription access to Claude Fable 5. Starting July 20, it will be included in all Max and Team Premium plans, with usage counted under 50% limits. Pro and Team Standard users can still access Fable through usage credits and will receive a one-time $100 credit. Claude acknowledged unpredictable demand, explaining the phased inclusion in subscriptions and repeated access extensions as more capacity became available. It also said access would be standardized at 50% usage for plans with the heaviest Fable use, making inclusions clearer. Developers and teams should confirm their plan, allowances, and credit arrangements promptly, since these affect whether Fable belongs in daily workflows.
https://x.com/claudeai/status/2078302415804379218
Official Blogs
An update on recent Claude Code quality reports
Anthropic reviewed reports of declining Claude quality over the past month, identifying three independent changes affecting Claude Code, Claude Agent SDK, and Claude Cowork, with the API unaffected. First, Claude Code’s default reasoning effort changed from high to medium on March 4 to reduce long latency and token consumption, but users perceived reduced intelligence. Anthropic reversed this on April 7, making xhigh the default for Opus 4.7 and high for other models. Second, a March 26 caching optimization was meant to clear old thinking once after more than an hour of inactivity, but a bug kept clearing it on every subsequent turn. Claude lost earlier reasoning context, appearing forgetful, repetitive, and odd in tool selection; this was fixed April 10. Third, an April 16 system-prompt change intended to reduce verbosity combined with other prompt changes to harm coding quality and was reversed April 20. Different traffic segments were affected at different times, creating the appearance of broad, inconsistent degradation. Anthropic announced usage-limit resets for all subscribers as of April 23.
https://www.anthropic.com/engineering/april-23-postmortem
Scaling Managed Agents: Decoupling the brain from the hands
Anthropic’s engineering article explains decoupling Managed Agents’ brain, hands, and session. Initially, session, harness, and sandbox shared one container, making file editing direct and boundaries few. But the container became an irreplaceable “pet”: failure could lose the session, and stalls were difficult to debug without accessing user data. The new design makes the session an append-only log, the harness a restartable orchestration loop, and the sandbox a replaceable execution environment. When a container dies, the harness can pass the failure to Claude as a tool-call error and provision a replacement sandbox from a standard recipe if needed. If the harness itself fails, it recovers through the session log. Security boundaries also become clearer: Claude-generated untrusted code no longer shares a container with credentials. A Git token can be attached to the remote during sandbox initialization, while custom-tool OAuth tokens remain in a vault. The central judgment is that harness assumptions become obsolete as models improve, so interfaces should be stable and implementations replaceable.
https://www.anthropic.com/engineering/managed-agents
New in Claude Managed Agents: self-hosted sandboxes and MCP tunnels
Claude Managed Agents added self-hosted sandboxes and MCP tunnels, letting enterprises place tool execution on infrastructure they control. Self-hosted sandboxes keep code execution, sensitive files, packages, services, and data within enterprise boundaries while Anthropic handles the agent loop, orchestration, context management, and error recovery. Enterprises can alternatively choose managed providers such as Cloudflare, Daytona, Modal, or Vercel for compute and isolation. Cloudflare offers microVMs, lightweight isolates, zero-trust secrets injection, and controlled egress. Daytona emphasizes long-running stateful sandboxes with SSH or authenticated preview URLs and pause/resume support. Modal targets AI workloads with custom container runtimes and on-demand CPUs and GPUs. Vercel combines VM security, VPC peering, bring your own cloud, and millisecond startup. MCP tunnels connect Managed Agents to MCP servers on private enterprise networks without public ingress, using an outbound connection from an enterprise-deployed lightweight gateway and supporting both Managed Agents and the Messages API. The key is returning execution, network, audit, and data boundaries to customers while retaining managed agent orchestration.
https://claude.com/blog/claude-managed-agents-updates
Podcasts
The MAD Podcast with Matt Turck — OpenAI’s Compute Chief: We Can’t Build Fast Enough | Sachin Katti
Key takeaway: AI’s next bottleneck is not just model capability but physical infrastructure comprising compute, power, cooling, networking, supply chains, and skilled labor.
Sachin Katti is now OpenAI’s head of industrial compute, having previously been a Stanford professor, serial entrepreneur, and Intel CTO. He leads the underlying compute buildout supporting training and running intelligence rather than model interfaces. This matters because OpenAI is operating compute as an intelligence supply chain: stronger models and more complex use cases make data centers, chips, networks, and power increasingly industrial-scale infrastructure problems.
Katti described the scale directly. Directional figures cited were approximately $50,000,000,000 in OpenAI compute spending this year and around $700,000,000,000 industry-wide. Demand far exceeds supply, he said, with any new compute immediately consumed; the greatest worry is that the physical world moves more slowly than software. One representative remark was: “Whenever you think you have enough compute and can slow down, the outcome always reminds you the hard way that you shouldn’t have.” This is not ordinary cloud expansion, but data centers as “giant factories turning electrons into tokens.”
AI data centers differ from traditional cloud facilities first in density and heat. Katti said AI essentially requires large supercomputers whose extremely hot chips cannot rely on air cooling. Cooling happens not only in data halls but on chips, connecting components, cables, and transformers, because everything handling energy produces heat. Liquid cooling is not new, but has never been deployed at this scale. Innovation therefore centers on reliability, cost, scalability, and more efficient liquids and materials. Better cooling lets chips run hotter, yielding greater memory bandwidth and flops and ultimately more intelligence.
Power is another hard constraint. Initially, companies mostly connected to the grid; OpenAI now also invests in generation and transmission infrastructure. Katti’s principle is that a new data center should not take existing grid power, but invest in additional generation capacity it can consume. This pushes AI companies toward industrial participation: beyond buying cloud resources, they must consider energy, transmission, and construction timelines.
OpenAI is also building its own compute capabilities more actively rather than only obtaining capacity from partners such as Microsoft, Google, Amazon, and Oracle. Katti said OpenAI is usually a tenant or offtaker committing to consume partner-built compute, but the scale required demands more direct participation in securing and building it. He called this a “new muscle” OpenAI is training. For builders, foundation-model companies are becoming hardware-, energy-, and capital-intensive organizations alongside software companies.
The speed of the in-house Jalapeno chip is particularly notable. Design to tape-out took nine months, one of the fastest timelines Katti has seen in his career. Reasons include a team with Google TPU design experience, Broadcom’s record delivering XPU ASICs, and OpenAI’s visibility into likely future model workloads, shortening many design decisions. More importantly, AI is assisting chip design and optimization, reducing time humans spend experimenting and processing data. He believes AI designing systems to train and run the next generation of AI—including chips—is not far away.
Large-scale training also requires new network reliability. Katti introduced MRC, a networking protocol / routing technology for scaling large cluster fabrics. With 100,000 GPUs continuously communicating, links, switches, and NIC cards are so numerous that failures are frequent and all failure modes may be impossible to enumerate. MRC uses multipath spraying: packets are distributed across multiple paths between chips, using whichever succeeds to prevent a single link failure from stopping training. The core goal is to abstract away network failures so training workloads need not care about them, rather than showcase technology.
The final counterintuitive bottleneck is people. Katti specifically identified shortages of electricians, plumbers, and other skilled trades, describing well-paid roles actively recruited by hyperscalers and labs. As AI reshapes knowledge work, physical construction capabilities become more valuable. OpenAI’s guaranteed capacity follows the same framing: effectively guaranteed tokens, locking in a dollar value of intelligence supply for enterprises. He believes intelligence is becoming a critical supply for every digital enterprise and must be managed like any other essential supply chain.