X / Twitter
AI Builder Swyx
Swyx advised users to avoid Codex's "locked use" capability for now. The feature currently relies on unstable macOS functionality and has already left him unable to access his macOS Keychain twice this week. Citing information from the Apple Developer Forums, he said this is a known issue. Although moving all related operations to the cloud might avoid local risks, he believes the cloud capabilities are not yet mature. For users who rely on Codex's operating system permissions or credentials, the safer choice at this stage is to avoid the feature until the underlying issue is fixed.
https://x.com/swyx/status/2092492963435946494
Thibault Sottiaux, OpenAI Codex and ChatGPT Product Team
Thibault Sottiaux introduced a product offering for teams and small companies with a usage model close to the Pro $100 plan. It covers all features of ChatGPT, ChatGPT Work, and Codex and can connect to services including Google Workspace, Slack, GitHub, and Microsoft 365. Enterprise management capabilities include secure workspaces, SAML, SSO, MFA, centralized billing, and unified administration. Teams can also view usage analytics and manage costs through spending controls. The plan has no 5-hour usage limit, focusing on connecting an individual high-allowance offering with enterprise collaboration needs. He also said good products take time, summarizing the development timeline as "at least 34 days."
https://x.com/thsottiaux/status/2092345330272780499
https://x.com/thsottiaux/status/2092487667426738179
Peter Yang, Practical AI Tutorial Creator
Peter Yang open-sourced the free AI skill `/fuck-cancer` to help cancer patients and caregivers organize information about diagnoses, treatment, and communication. It generates a continuously updated, practical brief that keeps patient and medical team information together and limits next steps to three clear actions. The document records confirmed facts separately from unresolved questions and explains medical terminology in plain language. Its fifth section is a care log containing recent developments and decisions, helping family members maintain a single source of information. The skill can read documents and background information supplied by users and consult sources such as the National Cancer Institute and ClinicalTrials.gov API when research is needed. Peter uses ChatGPT, Codex, and Claude Code to prepare for discussions and conduct research; results can be saved as local Markdown or synchronized to a Google Doc for family collaboration. He emphasized that the tool aims to help patients ask questions, move testing forward, and cope with the pressure of rapidly accumulating information from doctors, documents, and insurance materials.
https://x.com/petergyang/status/2092249012913258946
https://x.com/petergyang/status/2092311110871617915
Madhu Guru, Senior Director at Meta AI
Madhu Guru published the ninth installment of his eval series, discussing "The Eval Roadmap Problem." His central argument is that many eval failures happen not because the tests are entirely wrong, but because teams treat them as static artifacts and overlook rising user expectations and evolving usage. For a financial research agent, requests might progress from summarizing a 5-page earnings report to comparing the latest 5 reports, then building an investment thesis from 15 documents, and ultimately continuously monitoring a portfolio and issuing proactive alerts. Corresponding evals must move from short context to long-context single-turn questions, multi-turn citations, and document- and line-level citations, and from simple QA to complex synthesis and proactive agents. Teams should identify dimensions of evolution such as turns, document volume, tool use, autonomy, and user-journey coverage, then determine the next phase's P0 evals. Madhu recommends continuously interviewing users, analyzing production traces, looking for behavioral changes, and iterating on failure patterns. If evals remain at week one while users have reached the more complex usage of week three, the gap will eventually show up in product quality and churn metrics.
https://x.com/realmadhuguru/status/2092426017118028266
https://x.com/realmadhuguru/status/2092461206783373758
Cat Wu, Anthropic Claude Code and Cowork Product Team
Cat Wu announced that Claude has unified memory between Chat and Cowork. Users only need to tell Claude once what to remember, and the same background information can apply across both interfaces. This lets project context developed through chat flow directly into Cowork's tasks without repeated explanations. The update came from user feedback and focuses on reducing context loss when switching between products. Cat also invited users to suggest further capabilities they would like added to memory.
https://x.com/_catwu/status/2092337156455051345
Google Labs, Google's AI Experiments Platform
Google Labs launched Play with Putty, an experiment in collaborative vibe coding. It lets multiple people build tools and websites together in real time, extending individual AI coding into a collaborative workflow. The product emphasizes co-creation, allowing participants to work simultaneously on implementing the same idea. Users currently need to join a waitlist and provide feedback to the team. The experiment is currently available only to users aged 18 and over in the United States.
https://x.com/GoogleLabs/status/2092293667688173593
Guillermo Rauch, Vercel CEO
Guillermo Rauch released Run SDK for safely executing agent-generated code in dynamic Code Mode. It uses a lightweight QuickJS security context, suits tasks that do not require a full sandbox, and aims to reduce execution latency and cost. Developers can install it with `npm i run`. Guillermo also announced the general availability of Vercel Connect, which focuses on security when agents connect to services and data. Developers can run `vercel connect create notion` to create a connection and obtain an MCP client that executes queries on behalf of an authenticated user. These capabilities address agent code execution and external service access respectively, forming two key types of infrastructure. Their shared direction is to let agents securely access code, identities, and business data with lower runtime overhead.
https://x.com/rauchg/status/2092382653161107534
https://x.com/rauchg/status/2092352411839193234
Aaron Levie, Box CEO
Aaron Levie believes a substantial gap remains between foundation AI models and actual enterprise workflows, creating significant opportunities for applied AI companies. What enterprises really need is problems solved and outcomes delivered, rather than raw models or isolated agents. Bridging that gap requires products to understand industry context and drive organizational change management. Technically, it also requires a harness that routes between models, connections to critical business systems, and UX that places users and agents in the right workflows. Teams must also master evals for their domains to determine whether agents have actually achieved business goals. Aaron's view is that this full set of implementation capabilities beyond model intelligence holds considerable value, and that now is the window to build defining companies in each critical enterprise domain.
https://x.com/levie/status/2092466424694649066
Nikunj Kothari, FPV Ventures Partner
Nikunj Kothari launched the El Niño situation monitor, bringing real-time developments on El Niño into one place. The site aggregates government information and news and explains impacts and costs by region. It also includes historical records to help users understand the significance of current readings. A glossary and FAQ explain the various indicators and terms, making the material more accessible to nonspecialists. Initially inspired by an episode of Odd Lots, the project was built with ChatGPT Codex and Railway. Interface details were developed using Emil Kowalski's skills and Kasturi's generative loaders, and Nikunj publicly invited user feedback.
https://x.com/nikunj/status/2092383834470002922
https://x.com/nikunj/status/2092384774459674957
Aditya Agarwal, SPC General Partner
Aditya Agarwal believes public opposition to data center expansion is unsurprising. AI's main beneficiaries are still knowledge workers and the highest earners, leaving a gap between large-scale infrastructure construction and direct benefits for ordinary people. He expects public attitudes could change substantially when AI helps find treatments for diseases that affect everyone. He sees promising signs already emerging in this direction, which also helps explain why some of the smartest people are turning their attention to healthcare problems. Aditya also criticized the AI industry's longstanding emphasis on risks and fears without adequately describing a positive, tangible future. To gain broader public support, AI needs to demonstrate the value of infrastructure investment through real outcomes that everyone can see.
https://x.com/adityaag/status/2092290497826173186
Claude, Anthropic's AI Assistant
Claude now uses the same memory across Chat and Claude Cowork. When Cowork receives a task, it can directly use project background, manager preferences, or customer information that the user previously discussed in chat. All saved content appears as a topic list in Settings, where users can read, edit, or delete each item. Memory updates automatically through chats, continuously retaining newly emerging relevant details. Users can also explicitly say "remember this" to ask Claude to record a particular piece of information. Beyond continuity across interfaces, the design gives users visibility into and control over what is remembered.
https://x.com/claudeai/status/2092299704864284888
https://x.com/claudeai/status/2092299707653439497
Podcasts
Training Data — Parallel’s Parag Agrawal: Building a New Web for AI Agents
Key takeaway: Parallel believes search for agents cannot continue relying on human click data. Instead, indexes, ranking, interfaces, and the internet's business model must be redesigned around feedback from agents' tasks.
Parag Agrawal, formerly Twitter's CEO, has now founded Parallel Web Systems to enable agents to search and use the entire Web. His starting point is that agents could eventually use the internet 1,000 times as frequently as humans, making the technology stack traditional search engines built for human browsing, clicks, and keyword input insufficient.
The first key change is the feedback mechanism. In Parag's words: "At Parallel, human click data is a bug. Agents using search to get work done should rely on agent feedback, not human feedback." Traditional search uses vast amounts of clicks and human ratings to judge result quality, a scale advantage young companies struggle to replicate. Large models can now compress information and generate evaluation data more cheaply, giving teams an opportunity to apply model research to indexing and ranking.
The second change is to start with slower agents rather than replicating an entire search engine on day one. Parallel initially launched a search agent that could crawl pages after a query arrived. Deep research tasks allow a minute of waiting, and some products research for as long as ten minutes, leaving room for live crawling and page selection. Parag sees the index as a latency optimization: the team first competed with manual information gathering, then gradually built larger, more complex indexes through customer tasks. Early use cases included insurance underwriting, claims, and sales processes.
The third change is jointly optimizing quality, cost, and latency. Parallel's goal is not to train another enormous foundation model but to compress capabilities into small ranking models that select the most valuable approximately 1,000 tokens from trillions of web pages. Parag said using Parallel Search typically lets agents consume less than half as many tokens while improving accuracy and speed. Less noise means a model can handle more questions within the same context limit or complete the same task at lower cost.
The fourth change is at the level of the internet economy. Advertising works efficiently because limited human attention can translate into clicks, purchases, and subscriptions, subsidizing large amounts of free content through differentiated pricing. When visitors become unidentifiable agents, content providers cannot tell whether a crawl will lead to a subscription or monetize it through existing advertising models, so they may choose to block agents. Fixed-fee data licensing also struggles to keep pace with multiplying AI inference usage; at renewal, content providers may find their share of revenue continuing to fall.
Parag describes three layers of the future: agents using the Web as a search tool, more complex multi-agent systems activating and orchestrating one another, and the Web moving from pull to push. Eventually, agents will stop repeatedly asking "what should I find now?" and instead continuously monitor changes in web pages, satellite imagery, or customer reviews, proactively triggering work when events meet specified conditions. At that point, writing API documentation, content, and data interfaces agents can read will be as important as designing web pages for humans.
https://www.youtube.com/playlist?list=PLOhHNjZItNnMm5tdW61JpnyxeYH5NDDx8