X / Twitter
Google Gemini team VP Josh Woodward
The Thinking Levels experience that users have long complained about is finally rolling out across Gemini platforms. Josh announced that the Gemini App on web, iOS, and Android now lets users choose reasoning intensity themselves. This puts a reasoning tradeoff previously hidden in internal parameters directly in users’ hands, allowing them to switch between lightweight questions and deep analysis as needed instead of being bound to a default setting. Alongside the opening of Google AI Studio, this is an important step in narrowing Gemini’s product-experience gap with OpenAI’s reasoning models. For heavy users, the most direct benefit is fine control over the latency-quality tradeoff in production.
https://x.com/joshwoodward/status/2062025667852812583
OpenAI Codex and ChatGPT lead Thibault Sottiaux
Thibault has posted several signals over the past two days about Codex’s enterprise evolution. He particularly emphasized that the ChatGPT name, “whether you understand it or not, will become synonymous with AI and agents,” effectively declaring that brand narrative and product form will be tightly bound. More practically, Codex business users can now directly host and share websites they generate, and the plugins and skills systems have received major upgrades for more functional roles. A less-noticed change concerns feedback: users can now give agents visual annotations directly on docs, slides, and sheets, much more precise than text-only prompts. Together, these moves clearly aim to take Codex beyond developer tools into broader knowledge-worker use cases.
https://x.com/thsottiaux/status/2062057881424506950
https://x.com/thsottiaux/status/2061876999564791952
Roblox product manager Peter Yang
Peter’s recent discussion of whether SaaS will die in the AI era deserves close reading by founders. He rejects the “SaaS is dead” slogan but offers 3 specific filters: SaaS for a single narrow use case is becoming harder to sell because AI skills can solve the same problem more flexibly; agents such as Codex and Claude Code, which hold users’ long-term context and memory, have an overwhelming knowledge advantage over standalone SaaS websites and chatbots; and on pricing, people will pay hundreds or thousands of dollars for services delivered by people, but compare a $20 SaaS directly with Claude and ChatGPT subscriptions. He also gave the Devin and Windsurf teams their flowers, saying teams that maintain discipline and keep going through ups and downs are being recognized anew by AI-native builders. Another post quoted a memorable line from Matt: “They got so excited about being able to build anything that they launched with no users at all.”
https://x.com/petergyang/status/2061846283263103274
https://x.com/petergyang/status/2061936952400814392
https://x.com/petergyang/status/2062018242789670929
Anthropic Claude Code team member Thariq
Thariq called Claude Code’s newly launched Workflows its biggest capability upgrade since skills and subagents. After extensive use with sidbid, he compiled best practices and examples. He is particularly excited by the nontechnical tasks Workflows unlocks, meaning Claude Code’s target users are no longer only people who code. The thread was also published on Claude’s official blog. Workflows moves Claude Code’s positioning from “an agent that can code” toward “an orchestratable workflow engine.”
https://x.com/trq212/status/2061907538741006796
https://x.com/trq212/status/2061907897928528349
Replit CEO Amjad Masad
Amjad announced a new channel launched jointly with Microsoft: through Microsoft’s new Rayfin SDK, anyone inside an enterprise can build and deploy secure, compliant Fabric data apps on Replit. This is an important extension from a consumer vibe-coding tool into enterprise data platforms. He also reposted the new ViBench benchmark, emphasizing that “traditional SWE benchmarks do not capture app-building ability” and that dedicated app-building evaluations are needed to quantify models’ performance in real development scenarios. He shared a screenshot laying out an entire business structure on a canvas, hinting that Replit is turning business logic as a whole into a canvas AI can operate.
https://x.com/amasad/status/2061893093696434578
https://x.com/amasad/status/2061878314311266552
https://x.com/amasad/status/2062048812345291259
Vercel CEO Guillermo Rauch
Guillermo proposed a “YES-CODE” framework that directly challenges the logic of the no-code era. No-code rests on the assumption that code is expensive and scarce, but coding agents have overturned that assumption: code is now cheap, easy, and abundant. He cited colleague cramforce’s reply to analysts years ago—“Vercel is not a no-code platform, it is a yes-code platform”—as the company’s consistent positioning. His central argument is that Vercel’s mission is to “build the simplest cloud for agents that you never need to graduate from,” without compromising on performance or complexity. On education, he believes the entry point for learning in the AI era is language itself. English previously could not directly produce things and had to be translated into machine instructions; now it can go direct. He also announced that Vercel Sandbox powers Conductor, calling it “an ADE built for coding agents,” and predicted that agents will make remote development mainstream.
https://x.com/rauchg/status/2061934154732974376
https://x.com/rauchg/status/2061862134469062850
https://x.com/rauchg/status/2061809689973944724
Box CEO Aaron Levie
Aaron identified a clear opportunity for the enterprise AI application layer over the next 12 to 24 months: token budgets will become a major operating expense, making model routing an inevitable winner. Different domains have different working patterns, and only players with domain evals can finely optimize cost and performance. He believes most use cases still need frontier-model capabilities in the short term, but as evaluations mature, companies will gradually move suitable workflows to cheaper, smaller models. Crucially, enterprises struggle to scale these routing systems internally, so products that intelligently route workflows to the right model tier will aggregate more demand. This is an early step toward framing Box’s content-layer AI as a routing-layer AI platform.
https://x.com/levie/status/2061974298760495132
Builder Zara Zhang
Zara cited key figures from OpenAI’s latest Codex report: knowledge workers already account for around 20% of Codex users, and their adoption is more than 3 times faster than developers’. The fastest-growing task types are data analysis, up 110% week over week, research at 37%, and knowledge artifacts at 36%. Her own open-source Frontend Slides project has reached 20k GitHub stars and recently added polished templates, web publishing, PDF export, and inline editing. She said she was surprised by how many people had completely replaced PPT with HTML decks and never looked back. This signal supports her long-held thesis: once developer tools become accessible enough, knowledge workers eagerly adopt them.
https://x.com/zarazhangrui/status/2061924300698091760
https://x.com/zarazhangrui/status/2061889286585405790
FPV Ventures Partner Nikunj Kothari
Nikunj wrote a pointed long post for today’s seed and Series A founders. He sees too many pitches treat one item—AI, timing, fundraising, distribution, market, product, or revenue—as the central investment rationale, whereas excellent founders see all of them as necessary conditions rather than the entire business. He noted that the gap from seed to Series A is widening because the bar really is rising. Founders must explain how they build advantages others cannot easily replicate across multiple dimensions, especially in overcrowded markets. He understands the frustration of “the business is strong, so why is fundraising so hard?” but reminds founders that VCs’ opportunity costs are also soaring. Narrative and ambition determine the subtle difference between them and their next round. If they are profitable and do not need money, there is no need to accommodate VCs; but if they are going to have the conversation, they should spend ten more minutes clearly explaining their long-term ambition.
https://x.com/nikunj/status/2062033620773306763
OpenClaw and OpenAI founder Peter Steinberger
Peter has posted enterprise developments over the past two days. He and Omar added observability and verifiable workspaces to OpenClaw, allowing enterprise users to audit agents’ behavioral paths. The bigger news is the implementation of a Microsoft partnership, officially introducing OpenClaw to Microsoft’s enterprise customer base. These two steps take Peter’s self-described “ClawFather” role from early adopter to the center of enterprise AI agent infrastructure. Alongside OpenAI’s own enterprise iterations, the Claw ecosystem is rapidly adding the enterprise essentials of trust, auditability, and deployability.
https://x.com/steipete/status/2061877813053907083
https://x.com/steipete/status/2061874084649025728
Every CEO Dan Shipper
Dan had 3 noteworthy updates over the past two days. He publicly said goodbye to Every’s design lead Lucas, recalling that he joined four years ago selling ads and ultimately built a team that set a design benchmark in AI. Lucas played a key role in creating Every’s visual style. Dan also teased that “the team is freaking out about something worth watching,” without elaborating. On Opus 4.8, he made a candid request for feedback: Every was extremely positive during internal testing, but the external response had been more muted. He guessed that Opus 4.8’s tendency to challenge users’ framing creates a high-variance experience, sometimes producing stunning results and sometimes contradicting users in obviously wrong ways. He wanted the community’s views to help adjust internal evaluations. The inquiry itself reveals his judgment that model evaluation should reflect users’ actual experiences rather than benchmarks.
https://x.com/danshipper/status/2061962774918373592
https://x.com/danshipper/status/2061817375519809665
OpenAI CEO Sam Altman
Sam stated his industry position unusually directly: the United States should lead in AI by continuing to build the best models, ensuring they are safe, and putting cyber tools in the hands of trusted defenders. He said the new Executive Order “gets the balance right,” publicly endorsing this round of White House AI policy. As AI governance narratives become increasingly polarized, OpenAI’s active support for official policy suggests that government-business relations will remain a central arena for leading model companies.
https://x.com/sama/status/2061973280655904815
Anthropic’s official Claude account
Claude launched The Problem Solvers, a series focused on founders solving complex problems with Claude. Its first episode features Legora co-founder and CEO Max Junestrand, who is using Claude to bring one of the oldest professions, interpreting the law, into a new era. His bet is that every model release raises the water level across the industry, and Legora is building boats for everyone else. This is an important editorial move in Anthropic’s effort to expand Claude’s narrative from model API to customer success stories.
https://x.com/claudeai/status/2061829558999912680
https://x.com/claudeai/status/2061829560505655316
Podcasts
Training Data — Knowing What Your Customers Want, All the Time: Listen Labs' Alfred Wahlforss
Key takeaway: Once AGI makes building things cheap, knowing what to build becomes the truly scarce capability. Listen Labs is betting on AI-powered user interviews at scale and behavioral simulation to supply that capability to enterprises.
Guest Alfred Wahlforss is founder and CEO of Listen Labs, which launched a year ago and already serves 20% of the Fortune 500, with customers including Microsoft, Anthropic, Sweetgreen, and NBC. Listen is an AI-first user research platform capable of running thousands of voice interviews simultaneously. As AGI approaches the point of driving execution costs ever lower, his positioning is direct: “The closer we get to AGI, the easier it is to build things. The hard part is knowing what to build.”
Listen’s core mechanism assigns an AI agent to each research project. It first generates an interview guide, then recruits participants from the platform’s audience of 30 million for video interviews, and finally analyzes data and makes recommendations. Interviews use video so AI can capture expressions and ways of speaking, bridging the gap between “what people say” and “what they actually feel.” Alfred cited a counterintuitive finding: giving the same multiple-choice questionnaire to the same person repeatedly produces highly inconsistent answers, but having them respond verbally and think seriously through Listen substantially improves consistency. He stressed that 80% of engineering resources go into the audience, because “every company is driven by a power law.” Sweetgreen, for example, appears to serve everyone, but the customers responsible for 80% of revenue are “the 1% who are urban, high-household-income, mostly female, and know what seed oil is.”
The most contrarian observation was that AI is more popular than human interviewers: “Because AI does not judge and is always interested, respondents open up very candidly, and even involving children becomes possible.” He also said Procter & Gamble’s Tide Pods success originated in customer interviews revealing that “using liquid laundry detergent is really inconvenient,” while Mars’s early market research in the 1950s transformed M&M from a military snack into a treat for children. This is the kind of work Listen wants to redo faster with AI.
On the business model, he said pricing is actually rising: “We have done projects where we could charge customers hundreds of thousands of dollars to interview 20 doctors across 8 countries.” He acknowledged that the marginal cost of each interview will fall and total research volume will rise by two orders of magnitude. The company’s larger bet is simulation: after interviewing someone for an hour, it can predict their answers to certain questions with 95% accuracy and scale a high-fidelity model of a person into a representative sample of thousands. Listen has already run head-to-head comparisons using real data against ChatGPT. In his own case, when choosing a talk title from 100 candidates, the title selected through simulation had twice the conversion rate of the runner-up, while ChatGPT chose incorrectly in the same test. He offered a clear moat formula for vertical AI companies: “Own your proprietary eval; continuously climbing that eval is your moat.” Listen’s eval accuracy was 20% in the GPT-4 era and has now reached 85%, but it immediately designed a harder next-generation eval that started again at 20%.
Finally, he offered a future scenario: “When we reach the one-person billion-dollar company, Listen will remain in that loop alongside coding agents, forming an autonomous organization.” In his view, AI unlocks execution for everyone, but the scarcest step will always be “whose problem should we solve, and which one?” Listen is working to delegate that step to AI itself.