X / Twitter
Swyx Swyx
Swyx said he has been dogfooding an agentic GitHub clone for the past month and that it has become quite useful. The project includes CI/CD built on Workers for Platforms, indicating an attempt to combine code hosting, agent collaboration, and delivery pipelines rather than just code browsing or issue management. He said 3 undisclosed ideas remain to be implemented before launch and explicitly invited interested people to join swyx inc and help shape the roadmap. Another post focused on @poolsideai’s openness: he believes poolside has released not merely a strong Small model, but one that even outperforms @thinkymachines in coding. More importantly, poolside released its full eval dataset, covering 6 public benchmarks, with 4 runs per benchmark and hundreds of turns in each run. Swyx’s point was not simply to praise the model, but to identify open papers, open evaluations, and openly verifiable data as key to building trust for AI coding companies. For builders, the reminder is that beyond capability claims, letting outsiders inspect whether reward hacking occurred is becoming a more persuasive way to compete.
https://x.com/swyx/status/2080500752183960017
https://x.com/swyx/status/2080387171723137440
Thibault Sottiaux, OpenAI Codex & ChatGPT
Thibault Sottiaux first raised a naming question: should ChatGPT Work be renamed ChatGPT Vibe? The question itself offers little detail but suggests that OpenAI or related teams are still exploring the positioning of ChatGPT for work, particularly the naming tradeoff between “serious productivity” and “natural collaboration.” More substantively, he said the ChatGPT desktop app now offers Jarvis-, Samantha-, or TARS-like voice interactions, allowing users to work away from the keyboard. His emphasis on “away from that keyboard” suggests a direction in which voice becomes part of an ongoing workflow rather than simply an alternative input method. Together with Peter Yang’s posts, today’s theme is clear: AI voice interaction is moving from one-off exchanges toward continuous collaboration, background thinking, and multithreaded assistants. For AI builders, desktop voice capabilities may redefine the entry point for agents at work, taking them beyond the chat box.
https://x.com/thsottiaux/status/2080543574211666029
https://x.com/thsottiaux/status/2080408012515340394
Peter Yang, AI Tutorial Creator
Peter Yang is interested in ChatGPT Voice’s next form: he wants to start multiple ChatGPT Voice threads simultaneously, giving himself a whole team that can speak and converse with one another. This pushes voice beyond one-to-one human-model conversation toward a multi-agent meeting room, where users listen to multiple AI roles collaborating rather than merely issue orders. Another post demonstrated a ChatGPT Voice before-and-after; the source material contains no specifics, but shows him continuing practical demonstrations of voice productivity. His perspective suits busy users: voice reduces manual interaction and puts thinking, feedback, and revision into a more natural rhythm, rather than serving as a technical showcase. Several builders discussed voice mode today, suggesting it is moving from an experimental mobile feature into a foundational interaction layer for desktop work and agent workflows. For product designers, the real question will shift from speech recognition accuracy to keeping multithreaded voice collaboration orderly.
https://x.com/petergyang/status/2080508139091427741
https://x.com/petergyang/status/2080505964936241226
Madhu Guru, Senior Director at Meta AI
Madhu Guru summarized AI organization management’s dual challenge in one sentence: good builders understand the jagged frontier of AI models, and good leaders understand the jagged frontier of their team members. “Jagged frontier” means capabilities have uneven boundaries: models or people may excel at some tasks yet fail at seemingly similar ones. A more concrete post arose from his conversation with a public company’s security leader about agent security after the GPT Sol incident. He identified a central challenge: traditional identity and access management was designed for a finite number of employees, but one employee can now launch hundreds of agents, which may spawn further child agents. The questions become whether an agent inherits the initiating employee’s permissions; whether its lifecycle is a task, a ticket, or a week; whether child agents inherit the same permissions; and how auditing works. Madhu’s contribution is translating “agents are powerful” into enterprise governance questions about identity, permissions, lifecycles, inheritance, and audit trails. For enterprise agent teams, these are architectural questions to answer before launch, not compliance patches afterward.
https://x.com/realmadhuguru/status/2080460579966501257
https://x.com/realmadhuguru/status/2080315474093760714
Amjad Masad, Replit CEO
Amjad Masad said his chess autoresearch agent had “earned a PhD in modern LLM fine-tuning.” Though playful, the point is that agents can now conduct sustained research on specialist topics. Another post was more concrete: Viktor first used Replit to disrupt the agency model and make substantial money, then asked why he should not automate the whole agency rather than just coding. Amjad described an agency as “merely an agent loop,” suggesting many service companies’ delivery processes can be broken into repeatable agent workflows. Viktor asked the Replit team for MCP support, and after they built it, he created an autonomous agency. Replit’s value here extends beyond an online IDE to providing outside builders with MCP connections and automated execution. For founders, the signal is that the next wave of service-business automation may turn the entire sales, requirements, development, and delivery chain into agent workflows rather than replace a single tool.
https://x.com/amasad/status/2080512523389005894
https://x.com/amasad/status/2080371567221944657
Guillermo Rauch, Vercel CEO
Guillermo Rauch announced that Python code now automatically starts 2x faster on Vercel. This is practical for AI builders because many demos, API wrappers, background tasks, and notebook-style services rely on Python, and cold starts directly affect user experience. He stressed “automatically,” meaning users benefit without additional changes. Another post said AI Gateway’s product development continues to accelerate and praised the team’s velocity. Though no specific new capabilities were described, Vercel’s overall direction is toward a more complete production path spanning deployment, model access, and AI application infrastructure. Both updates concern making deployed AI applications faster, more reliable, and less configuration-heavy: Python runtime startup addresses the execution entry point, while AI Gateway addresses model calls. For independent builders, infrastructure competition is shifting from whether deployment is possible to whether default performance and AI-native workflows are good enough.
https://x.com/rauchg/status/2080454509508387251
https://x.com/rauchg/status/2080344136625049690
Aaron Levie, Box CEO
Aaron Levie offered a clear view of AI productivity: AI is best understood as a force multiplier in a field you already know, or a tool to accelerate learning a new one. He rejected a third approach: producing output through AI without existing judgment or any intention of developing it. Such results largely become slop and add little economic productivity. He believes experts benefit most because they know how to steer agents back on course, judge output quality, and incorporate results into real work. Experienced engineers, for example, produce more useful work with agents precisely because they know how to guide them. Designers likewise get better results from AI than people without a design eye. Aaron’s conclusion is that stronger tools may make specialization more important, not less, because market expectations for quality rise. For builders, this challenges the idea that “anyone can replace experts”: AI lowers execution barriers but does not automatically supply judgment.
https://x.com/levie/status/2080471989060559336
Garry Tan, Y Combinator CEO
Garry Tan’s two posts concerned physical infrastructure and the AI model ecosystem respectively. He first stated that it is time to build housing in San Francisco, continuing his interest in the SF boom loop and the physical conditions supporting startups. Housing is not a city issue unrelated to technology for AI builders, since talent density, team formation, and startup costs all depend on urban supply. Another post emphasized that open-weight models are very important. The material contains no further argument, but alongside today’s discussions of poolside’s open eval dataset and enterprise agent permissions, it points to a shared theme: AI’s foundational capabilities should not be defined solely by a few closed interfaces. Garry’s concern is expanding founders’ access to both city and model infrastructure. For founders, open models and urban construction are both underlying conditions for building faster.
https://x.com/garrytan/status/2080443154730553402
https://x.com/garrytan/status/2080345524620914897
Matt Turck, VC at FirstMark Capital
Matt Turck first joked about a contrast in fundraising narratives: a profitable bootstrapped business can excite some VCs less than a neo-lab burning hundreds of millions of dollars on compute. The satire targets AI capital-market preferences and reminds founders not to mistake high compute consumption for a high-quality business. He then featured his interview with Cerebras CEO Andrew Feldman on fast inference, AI chips, and the next compute bottleneck. Beginning with “what is a wafer?”, the conversation explored why the chip industry is reorganizing around inference speed. The timeline covers tokens per second per user, GPU/TPU/Trainium/ASIC, Nvidia, Groq, OpenAI, Broadcom, China, power, HBM, CoWoS, 3nm, agent-driven CPU demand, prefill and decode, the CUDA moat, TSMC, and SaaS’s future. Matt brings chip discussions down from broad capital-market excitement to performance metrics builders can understand: speed, memory, supply chains, and data-center power consumption. For AI application developers, inference is no longer merely a cloud-provider detail; it directly determines agent experience, cost structure, and product form.
https://x.com/mattturck/status/2080451010439352711
https://x.com/mattturck/status/2080333711640285549
https://x.com/mattturck/status/2080333707483725876
Nikunj Kothari, FPV Ventures Partner
Nikunj Kothari listed terms overused in technology and gradually losing their signal: neo-something, full stack, fellows, labs, partner, forward deployed, and RL, which is slowly approaching that state. He added a self-deprecating note: his own firm has a fellowship, and his title is partner. The post reminds readers that overused industry language deteriorates from a signal of capability and positioning into packaging. Fundraising and hiring narratives depend on words, but when every company is a lab and every role is forward deployed, audiences struggle to identify real differences. His observation also applies to AI product naming: labels such as neo-lab, agentic, and full-stack AI ultimately dilute credibility without concrete capabilities and evidence of delivery behind them. The lesson for builders is to explain who the customer is, what the workflow is, and where performance or results improve, rather than chase fashionable titles.
https://x.com/nikunj/status/2080293627784212933
Peter Steinberger, OpenAI
Responding to a system limitation or integration issue, Peter Steinberger said his team had seen the same behavior and added code paths that use claude cli directly. His judgment that it is “hard to fight the system” suggests that in environments with multiple models, CLIs, and agent tools, engineering teams sometimes follow existing interfaces rather than force every difference into an abstraction. Although the source lacks the full context of the issue he was answering, it clearly involves direct Claude CLI calls and alternative execution paths. For AI tooling builders, this is a practical engineering signal: when an official CLI becomes a de facto reliable entry point, supporting it directly may deliver faster than pursuing a unified protocol. It also echoes today’s references to MCP, agentic workflows, and Claude Code artifacts: AI development environments are becoming combinations of tool protocols, CLIs, and product interfaces. Useful products often need to accept this mixture and establish reliable paths first.
https://x.com/steipete/status/2080318789980201224
Claude, Anthropic’s AI Assistant
Claude announced updates to voice mode: voice conversations can now use more of the models already available in chat, including Claude Opus and Sonnet. More importantly, Claude can access connected tools such as email and calendars during voice conversations, expanding voice from a chat experience into an executable workflow entry point. Claude also said voice mode supports more languages—including Spanish, French, Hindi, and Japanese—and is available on every plan. The update begins rolling out today in public beta on mobile, desktop, and web. Mobile users can tap the sound wave to start a conversation. Together, the three posts show Anthropic building voice mode into a cross-platform, cross-model interaction layer with tool access, rather than merely a mobile voice assistant. For builders, the next stage of voice agents will be differentiated by model capabilities, tool permissions, context continuity, and multilingual coverage, rather than simply whether they can speak.
https://x.com/claudeai/status/2080376099268169943
https://x.com/claudeai/status/2080376096873177300
https://x.com/claudeai/status/2080376094939603366
Official Blogs
Claude Code now supports artifacts
Claude Code now supports artifacts, turning the progress of a Claude Code session into a shareable visual page that updates continuously. Use cases include PR walkthroughs, system explainers, dashboards, release checklists, incident investigations, service refactors, and analysis of months of data. The central change is that Claude Code can go beyond conclusions in a terminal or conversation and use a session’s full context—including the codebase, connectors, and conversation—to generate a page teammates can open directly. An incident page, for example, can show a failing test, relevant functions, an error spike from monitoring tools, and Claude Code’s root-cause reasoning in one view.
Another key feature is live updating: when Claude Code updates an artifact, an already-open page refreshes in place, letting teammates see the new version at the same URL. Each publication is a new version at the same link, with version history and rollback support; a gallery provides browsing and management of created artifacts. Anthropic said debugging was a common internal test use case: engineers start an incident investigation before standup, and Claude Code publishes an artifact with a timeline, suspect commits, and an error-rate chart, repeatedly updating it as the investigation progresses. Teams can share the same contextual view rather than listen to someone verbally recount what the agent found.
Artifacts are visible only to their author by default. Once ready, authors can share them with teammates or the organization from the page header. Only authenticated organization members can view artifacts; they cannot be published publicly. Administrators can manage access and visibility through an organization-level toggle, role-based scoping, retention policies, and the compliance API. Usage is straightforward: ask for an artifact in a Claude Code session or request a visual task such as a license audit, personal data flow map, authentication findings, Terraform cost drivers, PR walkthrough, signup-form UX variations, service import graph, incident page, or weekly summary of merged PRs. The feature is currently in beta for Claude Team and Enterprise organizations, available in the Claude Code CLI and desktop app, with pages viewable in any browser.
https://claude.com/blog/artifacts-in-claude-code
Podcasts
The MAD Podcast with Matt Turck — The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
Key takeaway: The next round of AI competition will shift from “how intelligent is the model?” to “how many tokens per second can each user get?”, because agents, real-time interactions, and complex tasks amplify waiting time into a product bottleneck.
Andrew Feldman is Cerebras’s co-founder and CEO. Cerebras makes the largest computing chip ever built, described in the material as 58 times larger than a GPU. It has just completed a semiconductor IPO and has a $20,000,000,000-plus deal involving OpenAI. The important part of Feldman’s background is not the marketing headline but his longstanding bet on wafer-scale computing. While much of the industry discussed training around GPUs, he broke the problem down into deeper constraints involving inference, memory, data-center power, and supply chains.
His most important judgment is that around mid-2025, AI moved from “novel but not very useful” to “useful enough to be used frequently.” Training brings a model into existence, but using it requires inference. Once users bring AI into everyday production, speed immediately becomes valuable. Feldman’s metric is explicit: tokens per second per user, the generation speed each user actually experiences from the first token to the last. It matters even more for agentic flows because multiround, multistep tasks compound waiting time.
He offered a useful Netflix analogy: when the internet was slow, Netflix mailed DVDs; when it became fast, Netflix did not mail DVDs more efficiently—it became a movie studio. AI is similar: speed is not a minor optimization but opens new ways of using the product, keeping users longer, bringing them back more often, and letting them tackle harder problems. One pointed quote captures his view: “How big is the market for slow search? How big is the market for dial-up internet? It’s zero.” The implication is that users will not tolerate agents slowly running in the background indefinitely; a sense of real-time responsiveness will become a basic product requirement.
On the chip landscape, Feldman traced the history from CPUs to GPUs and then specialized AI chips such as TPUs, AWS Trainium, Microsoft Maia, Groq, and Cerebras. The essence of an ASIC is not a mysterious acronym but deliberately sacrificing generality for a class of tasks: stronger on some problems, weaker on others. He viewed NVIDIA’s acquisition of Groq as a strong signal that the story of GPUs doing all AI work is no longer complete and that fast inference has become a sufficiently large market. Cerebras’s position is that it was designed from the outset for AI workloads, rather than optimized for one hyperscaler’s internal problem or one lab’s isolated need.
His views on China and infrastructure were specific: China lags in chips but has advantages in grid and power investment, and power is a critical resource for data centers. He also explained that Cerebras does not sell to China for regulatory and geopolitical reasons. His view of local AI and local chips resembles the division of labor in the mobile internet era: do what can be done on a phone or laptop close to the data, while truly compute-intensive work still goes to the cloud and data centers. In other words, edge computing will matter, but data-center compute remains the center of heavy workloads.
The most counterintuitive supply-chain point is that chip manufacturing cannot easily move between factories. Feldman said chip designs follow a particular fab’s rules, so a design for TSMC cannot simply be produced somewhere else. Cerebras’s next generation will still use TSMC, but chips will return to the US for repackaging, assembly, manufacturing, and shipping. Rapid growth creates problems that cannot be summarized simply as “supply-chain pressure”: faulty batches, customs delays, supplier issues, and many other everyday failures require continual improvement in manufacturing throughput.
His final framing is useful for builders: the model you use today will be the worst model you ever use going forward. Capabilities that seem impressive now may look outdated in six months. What matters is not a day’s stock-market movement or an isolated benchmark, but how rising AI usage combines speed, power consumption, memory, supply chains, and product experience to define new application boundaries.