X / Twitter
Swyx
Swyx said he had been running 5.5 Opus and 6 Sol in parallel for Latent Space’s AINews and had observed clear differences in their output. He preferred 5.5 Opus’s reporting style, finding its writing more concise and its editorial judgment more tasteful. Compared with 5 Opus, he felt the new model further reduced vague, formulaic AI prose. This comparison took place within AINews’s actual workflow, with an emphasis on reporting quality and presentation. He has decided to make 5.5 Opus the default model for AINews going forward.
https://x.com/swyx/status/2102650014552182920
Boris Cherny, Anthropic Claude Code Team Member
Boris Cherny explained in detail how Claude finds code defects through a process resembling formal methods. The first step is to model complex state machines or parts of a program prone to race conditions, then look for counterexamples in the model and treat them as potential bugs. Claude then tries to reproduce these issues before fixing the actual code. He stressed that this does not mean the entire codebase has been formally verified; it means modeling, checking, and fixing the most difficult parts. Another post described the engineering work behind speed improvements to the web and Desktop apps over the past few weeks. He recommended that engineers read the specific techniques and lessons in the accompanying blog post to improve their own applications’ performance.
https://x.com/bcherny/status/2102898067133595992
https://x.com/bcherny/status/2102854267782705648
Thibault Sottiaux, OpenAI Codex and ChatGPT Team Member
Thibault Sottiaux previewed some interesting new developments for next Tuesday’s DevDay, including features he believes could change how people work. He described the recent development push as the team’s most ambitious sprint yet and said Astra had enabled the team to build new capabilities in a very short time. Another post described his everyday use of ChatGPT voice. He has grown accustomed to calling ChatGPT directly to discuss work, check email, handle some coding tasks, and manage his calendar. According to his account, this voice interaction can access the entire plugin ecosystem, including plugins built by third parties. Together, the two posts focus on how people interact with ChatGPT in practical work and the range of tools it can access.
https://x.com/thsottiaux/status/2102996313780736363
https://x.com/thsottiaux/status/2102814202117411196
Peter Yang, AI Tutorial and Interview Author
Peter Yang argued that once model capabilities are strong enough, Claude Code’s harness also needs better interaction and execution capabilities to match. He specifically called for high-quality real-time voice, as well as browser and computer use. He added that the latter two were already improving. He summarized his view as “one lab has the best model, the other has the best harness,” without explicitly naming the two labs in the text. His assessment evaluates model quality separately from the product framework and clearly identifies the features he wants Claude Code to keep improving.
https://x.com/petergyang/status/2102952201350254667
Madhu Guru, Senior Director at Meta AI
Madhu Guru disagreed with treating consumers’ attitudes toward “getting things done” as a single preference, drawing on his experience building consumer and small-business products at Google and Meta. He pointed out that browsing clothes is entertainment for some people, while finding a roofer is a headache for most. Even when shopping, users may sometimes want to browse for an hour and at other times simply want the right product delivered as quickly as possible. He listed the cumbersome steps involved in hiring a roofer: searching, reading 10 reviews, making repeated phone calls, coordinating insurance, scheduling inspections, getting quotes, and checking for hidden fees. For such tasks, a better experience may mean requiring almost no effort from the user; for others, it means comparing more options, enjoying the process, and making an informed decision. Drawing a parallel with early consumer skepticism about buying clothes or a $500 television online, he emphasized that technology continually changes how people accomplish tasks. His conclusion was that there is enormous potential demand for agents that can remove these obstacles for consumers.
https://x.com/realmadhuguru/status/2102777931764498536
Thariq, Anthropic Claude Code Team Member
Thariq poked fun at the preparation that is easy to overlook behind showcase posts claiming that “Claude built it in one shot.” In his example, the actual input included a 10k-character prompt, carefully organized ideas, skills, examples, and API keys. His comment reminds readers that the result of a single generation is closely tied to how much context was prepared in advance. He also introduced a new article format he is trying, sharing in detail how the team uses specific prompts and techniques to get work done. The aim is to help readers reproduce the methods, rather than just see the finished results. He invited feedback on whether these details were useful so he could continue refining how they are shared.
https://x.com/trq212/status/2102870353781641416
https://x.com/trq212/status/2102857025206255902
Guillermo Rauch, Vercel CEO
Guillermo Rauch shared his work optimizing Shell startup speed, saying that Opus 5.5 had found improvements other models missed and suggesting that users ask an agent to inspect configurations such as `.zshrc`. He then broke a successful agent down into three parts: the model and harness responsible for reasoning and orchestration; the tools, browser, and computer responsible for actions; and the files holding memory, skills, and code repositories. Putting all these components on a continuously running Mac Mini is straightforward, but controlling costs in the cloud requires decoupling them. He outlined a combination in which Fluid compute hosts the harness, Workflow persists the event log, and Browserbase, Kernel, Sandbox, and just-bash handle different execution tasks. Drives, introduced in this announcement, adds independent storage that can serve as an “external disk” mounted to a Sandbox on demand. For example, during overnight memory consolidation, files can be read and written directly without starting the agent’s entire computer. He said Drives was still at an early stage and argued that separating components can substantially improve security and auditability as well as reduce costs.
https://x.com/rauchg/status/2102947745132924993
https://x.com/rauchg/status/2102820148629614685
Aaron Levie, Box CEO
Aaron Levie endorsed a view of AI and filmmaking: as barriers and costs fall, more films will be made, and studios will be able to take more creative risks. The view he quoted also held that more people would gain opportunities to participate, giving rise to new forms of storytelling. The argument cited animation’s evolution from a niche into a widely popular storytelling form, along with Steven Spielberg, James Cameron, and Peter Jackson’s use of new visual tools. Levie used these examples to emphasize that technology has repeatedly reshaped creative industries over time, bringing more opportunities or new creative methods. In his view, creative ability and taste remain irreplaceable even as technology and media change. The role of new tools is to let more people put those abilities to effective use and explore new ways to tell stories.
https://x.com/levie/status/2102934874470617303
Ryo Lu, Who Has Worked on Design at Cursor, Notion, and Stripe
Ryo Lu questioned the AI industry’s continued worship of efficiency, productivity, and speed, arguing that shipping faster, merging more changes, and managing more agents do not answer whether the output itself deserves to exist. He worried that people are continually compressing work cycles while losing the time to think deeply, recognize their true intentions, and find a purpose. In his view, AI could be a new canvas for creating tools, worlds, poetry, games, interfaces, and films, yet many applications are moving toward low-quality content, marketing funnels, and busywork. By juxtaposing merging 2000 PRs and managing 500 agents with sleeping only 5 hours a day, he asked whom this mode of production ultimately serves. His counterintuitive warning was that AI’s danger may not be making people lazy, but giving them unlimited capacity to execute and keeping them perpetually busy before they have developed intention and taste. He therefore saw discernment as a more worthwhile pursuit, including knowing what not to do and when to stop. His central argument was that tools should make more room for life, imagination, and freedom, rather than turn life into part of a production system.
https://x.com/ryolu_/status/2102933485795369213
Garry Tan, Y Combinator President and CEO
Garry Tan summarized changes in how startups acquire customers as two trends happening at the same time. The first is to build software that agents want to use. In this formulation, agents themselves become users that a product needs to consider. The second is to use agents to make people want to use software, keeping the focus on human needs and choices. By placing the two side by side, he identified two directions: building products for agents and using agents to acquire human customers.
https://x.com/garrytan/status/2102955139875397806
Dan Shipper, Every CEO
Dan Shipper announced that Every would test whether AI agents could plan an event well by holding an in-person gathering. He said Every’s agent had already planned the September gathering, covering the food, drinks menu, and guest list. The venue was Every’s brownstone in Brooklyn, and the post gave the time as “tomorrow at 6 p.m.” He invited attendees to come and rate the agent’s planning. The experiment puts an agent’s planning abilities to the test at a real event; the source material does not yet provide results or participant feedback.
https://x.com/danshipper/status/2102826854357016793
Aditya Agarwal, SPC General Partner and Bevel Health Co-Founder
Aditya Agarwal argued that when assessing a team, the time its members have worked together deserves more attention than headcount alone. He noted that teams that have worked together for at least 3 years often have significantly higher output and greater resilience than newly formed teams. He considered low staff turnover a signal worth watching and also believed that frequent changes in an individual’s career history could provide information. He criticized investors and job seekers for often asking only about team size while overlooking how long the team has worked together. He summarized the point as “the team’s AUC is the right metric,” emphasizing that accumulated time should be considered alongside team size.
https://x.com/adityaag/status/2102782451643040074
Claude, Anthropic’s AI Assistant
Claude’s official account introduced Claude Marketplace, where users can discover tools, agents, and professional services partners. One category consists of connectors and plugins such as Slack and Notion, which expand the work tools Claude can connect to. Another includes agents and products from companies such as Cursor and CrowdStrike, available for purchase through the marketplace. Enterprises can also find services partners such as Accenture and Deloitte to support broader adoption. The official account also invited developers building for Claude to submit their tools, agents, or services, positioning Marketplace as a destination for both discovery and listing.
https://x.com/claudeai/status/2102840851538080172
https://x.com/claudeai/status/2102840855258452037
Official Blogs
Anthropic: How Claude’s Access and Scope of Impact Are Contained Across Products
Anthropic divides the risks of deploying agents into two dimensions: the probability of a failure and the scope of damage a single failure could cause. As model capabilities and access permissions expand, the potential damage can increase even if the probability of error falls. The key engineering challenge is therefore to set hard boundaries on an agent’s scope of impact.
The article explains why case-by-case human approval is insufficient for this task: telemetry shows that users approve approximately 93% of permission requests, and frequent approvals erode attention. Claude Code auto mode eases this fatigue by automatically handling safer approvals, but probabilistic defenses still miss some cases. According to the article, auto mode blocks approximately 83% of overly proactive actions before execution. In Gray Swan’s Agent Red Teaming benchmark, the attack success rate against Claude Opus 4.7 was approximately 0.1% for a single prompt injection attempt, rising to approximately 5–6% across 100 adaptive attempts.
Anthropic therefore emphasizes environmental isolation, including process sandboxes, virtual machines, filesystem boundaries, and outbound network controls. The article distinguishes three sources of risk—user misuse, the model taking harmful actions on its own, and external attackers—and sets up protections across three layers: the execution environment, model defenses, and external content. One specific principle is to keep credentials out of the sandbox from the outset, so that security does not depend on whether the model follows instructions. The article also cautions that a connector passing an audit does not make the data it reads trustworthy. For example, a GitHub connector could still place a README containing an attack into the model’s context.
https://www.anthropic.com/engineering/how-we-contain-claude
Anthropic: A Postmortem on Three Changes Behind Reports of Declining Claude Code Quality
This postmortem on issues in April traces users’ perceived decline in quality to three independent changes affecting Claude Code, Claude Agent SDK, and Claude Cowork; the API was unaffected. The article says all three issues were resolved in v2.1.116 on April 20. These historical fixes should not be mistaken for new events in September.
The first change occurred on March 4, when Claude Code lowered its default reasoning effort from high to medium to reduce the long delays some users experienced. Users said they preferred a higher level of intelligence by default and wanted to reduce effort for simple tasks themselves, so the team reversed the adjustment on April 7. The second change occurred on March 26. The intention was to clear old reasoning records once after a session had been idle for more than an hour, but a bug caused them to be cleared continuously on every subsequent turn, resulting in forgetting, repetition, and unusual tool choices. This issue was fixed on April 10.
The third change occurred on April 16, when a system prompt intended to reduce verbose output combined with other prompt changes to harm coding quality; it was reversed on April 20. The three changes affected different traffic at different times, making overall performance look like a broad and inconsistent degradation. Internal usage and evaluations initially failed to reproduce the problems. The postmortem shows that user experience depends not only on the model itself: default reasoning parameters, the handling of historical context, and system prompts also directly influence coding results.
https://www.anthropic.com/engineering/april-23-postmortem
Anthropic: Scaling Managed Agents by Separating Reasoning, Execution, and Session Records
Anthropic introduced the architecture of Claude Managed Agents: a session is defined as an append-only event log; a harness is the runtime loop that calls Claude and dispatches tool requests; and a sandbox is the environment for executing code and editing files. The three connect through independent interfaces, allowing the underlying implementations to be replaced without changing the entire system at once.
Initially, the team put all these components in the same container. This made file operations direct and system boundaries simple, but it also turned the container into a single point the system could not readily afford to lose or replace. A container failure could cause a session to be lost, while harness errors, dropped events in the event stream, and an offline container could produce similar symptoms. Keeping user data inside the container also made it difficult to enter the environment directly to troubleshoot. When customers wanted to connect their own private networks, this coupling created additional deployment constraints.
After the separation, the harness calls the execution container through a tool interface. Container failures can be passed to Claude as tool errors, and the container can be recreated using a standard configuration. The harness itself can also restart after a crash and resume work from the externally persisted session log. The key security change is to prevent untrusted code generated by Claude from accessing sensitive credentials required for operation—for example, by storing OAuth tokens for custom tools in a secure vault outside the sandbox.
The article also emphasizes that logic in the harness intended to compensate for weaknesses in older models needs continual review. For example, the team once used context resets to address Sonnet 4.5’s tendency to end tasks prematurely as it approached the context limit. The same problem had disappeared in Opus 4.5, making the original workaround unnecessary. The design goal of Managed Agents is to keep interfaces stable while allowing models, harnesses, and execution infrastructure to evolve independently.