X / Twitter
Thibault Sottiaux, OpenAI Codex & ChatGPT
Thibault Sottiaux shared a real task he dictated directly to ChatGPT Work: go through Twitter DMs and compile every message mentioning early-access ChatGPT Work beta into a spreadsheet. He requested names, DM links, and message content, with line breaks removed so each message fit in one cell. More importantly, he asked ChatGPT Work to design a taxonomy of roughly 8 to 12 labels based on users’ work, such as sales, marketing, and design. Another field would assess the sophistication of their current ChatGPT Work workflows from their descriptions, helping select a beta cohort spanning different levels. The example combines inbox triage, taxonomy design, candidate selection, and spreadsheet creation into a delegable workflow rather than simple extraction. He then said his work has essentially become “delegating to ChatGPT Work,” with dictation especially indispensable. Finally, he explicitly listed use cases: creating and hosting sites, managing email, summarizing large document collections, and generating docs, sheets, and slides. It is already available in the mobile app and on the web, included in Plus, Pro, Business, and Enterprise plans.
https://x.com/thsottiaux/status/2078702412085498087
Peter Yang, AI Tutorials Builder
Peter Yang shared a concrete parent-child AI building example: he and his 8-year-old made a @ChatGPTapp Site to help her learn multiplication tables. Starting from her interest in animals, they used ChatGPT Images to generate the UI and characters. The page added music alongside multiplication practice, making the experience more like a small game. They also created a timed boss level, turning rote memorization into an interactive process with rhythm and challenge. Its value lies in showing how an ordinary family converts a personalized learning need into a playable miniature product, rather than merely demonstrating an AI-generated demo. For builders, it suggests an important ChatGPT Sites use case: easily generating small, customized tools designed around real users’ interests.
https://x.com/petergyang/status/2078638568784994686
Guillermo Rauch, Vercel CEO
Guillermo Rauch discussed several models’ cybersecurity performance based on internal evals. He called Kimi K3 top-tier in cybersecurity and responded to claims on X that Moonshot overfits benchmarks by emphasizing that these were stealth evals, concluding the model has raw IQ. Sol also showed a clear leap in cyber capabilities, at significantly higher cost but with impressive ability. Fable, by contrast, refused too frequently for the team to complete a full run. He specifically noted Sol’s greater willingness to help with defensive cyber hardening, concluding that frontier-level open-weight cybersecurity capabilities have arrived but should be used defensively. In another long post, he criticized “AGI” as an aging term: AI differs enormously from human intelligence and already far exceeds humans on most economically relevant tasks, but that does not mean it is better than humanity. Irreplaceable human qualities, he argued, include caring for others, maintaining identity, expressing distinctive ideas, and creatively using machines rather than outsourcing writing and judgment entirely. He cited AI social replies and templated landing-page copy as counterexamples, believing quality and humanity will win.
https://x.com/rauchg/status/2078647648307880209
Aaron Levie, Box CEO
Aaron Levie argued that the past few months have overturned the assumption that AI ecosystem value will flow only to a handful of frontier labs. He acknowledged that frontier labs will keep advancing models thanks to compute scale, major revenue streams, customers, top researchers, and data pipelines. Meanwhile, a diffusion layer around frontier models is rapidly forming through several viable paths. First are companies helping enterprises and applied AI businesses develop models for specific use cases and run inference, gaining domain-specific performance and cost advantages. Second are applied AI companies responsible for end-user experiences and business-process tools across legal, IT, security, HR, customer support, coding, and other areas. These can act as model-routing layers with deep understanding of enterprise workflows, change management, and data access. Third are new labs specializing in life sciences, financial services, healthcare, and other verticals outside frontier labs’ main focus that require deep domain expertise. Fourth is infrastructure for running and governing agents and models, storing data, security, and orchestration, alongside the many services firms that will help enterprises adopt AI. He believes it is too early to define winning architectures and expects a heterogeneous environment. He also argued that gatekeeping models will not work at scale: the US should advance rapidly while maintaining safety, promote diffusion, build infrastructure, and support US OSS.
https://x.com/levie/status/2078567715544121815
Garry Tan, Y Combinator CEO
Garry Tan’s two posts concerned public governance and accountability in San Francisco. He called for real investment in recovery and treatment, alongside rigorous audits and defunding of harm-reduction organizations that he believes actively encourage, assist, or cause overdose deaths and support drug dealers. His central position prioritizes treatment and recovery over unconditional harm reduction as the response to the city’s drug crisis. He also criticized The Standard, saying it should emphasize journalism more and partisanship less. Demanding city accountability is reasonable, he argued, while attacking accountability misuses journalistic influence. For those following startup ecosystems, these posts reflect his sustained treatment of SF governance, public safety, and urban renewal as foundational issues for founders.
https://x.com/garrytan/status/2078699319209984033
Matt Turck, FirstMark VC
Matt Turck offered a concise satire of the “model layer is commoditizing” narrative. He listed 2024, 2025, and 2026 as years when people repeated the same claim, then concluded directly that the model layer still has not commoditized. This challenges a recurring assumption that open models catching up and falling API prices will turn foundation models into minimally differentiated infrastructure. Rather than explain technical details, he used the passage of time to remind readers that model differences, windows of leadership, and commercial value persist. Another post about Didier Deschamps was primarily a sporting tribute; although it included match and championship numbers, it was less relevant to AI builders.
https://x.com/mattturck/status/2078520552680046920
Zara Zhang, Builder
Zara Zhang recommended that everyone create a personal eval set for AI models: tasks genuinely relevant to everyday work or life. Industry benchmarks are informative but may not show whether a model is useful to you personally. Finding capability boundaries means repeatedly “poking at it” with concrete tasks and discovering limits through engaging experiments. This is practical for builders, since model selection often depends on real workflow performance rather than overall leaderboards. She also said the biggest barrier to enterprise AI adoption is that people who understand AI do not understand the business, and vice versa. Together, the posts bring capabilities back from abstract evaluation to specific use cases while addressing the gap between business context and AI expertise.
https://x.com/zarazhangrui/status/2078666187026911488
Podcasts
Unsupervised Learning — Ep 90: AI Pioneer Jürgen Schmidhuber on the State of AI Today
Key takeaway: Jürgen Schmidhuber is highly optimistic about AI technology itself but deeply pessimistic about large model companies’ current capital spending and business models. True AGI cannot stay behind a screen, and model leadership is difficult to turn into a durable moat.
Jürgen Schmidhuber, described as an AI pioneer by the New York Times, Forbes, and others, has long researched neural networks, meta learning, artificial curiosity, and a formal theory of fun and creativity. His perspective differs from many current AI-company narratives: he does not deny that LLMs can pass the Turing test or that screen-bound AI is already powerful, but argues that “true AI” must enter the physical world. The crucial question is therefore not just stronger chatbots, but robots, real machines, and systems that act and learn in physical environments.
His most counterintuitive judgment is that current data-center and GPU investments may backfire on large companies. Compute per dollar has historically improved roughly 10-fold every five years, he noted, so GPUs bought with enormous capital today may lose most of their economic value within five years. He sees companies’ shift from asset-light software into operating clouds, data centers, power plants, and gas turbines like utilities as a major cause of deteriorating free cash flow. To the claim of unlimited future compute demand, his answer is simple: someone must pay for demand, and a business model cannot stand if providers keep losing money.
On open versus closed source, he believes closed-model companies will struggle to earn lasting profits from a 3- to 6-month lead. Open-source models follow closely, exerting intense pricing pressure. Even if a commercial model sets a benchmark record today, open models may catch up within months, making costly training and data-center debt harder to recoup. Recursive self-improvement may not become an exclusive moat for large companies either, since many important ideas come from small labs, academic labs, and PhD students rather than a few private corporate systems. His sharp formulation was: “Everyone is cooking with the same water.”
His explanation of RSI emphasized its research history. He reviewed meta evolution in 1987, self-referential machines in 1994, and the Godel machine in 2003: before changing its code, a machine must prove the change offers higher expected reward than leaving it unchanged. He acknowledged that this mathematically more general path is difficult in practice. Today, a weaker form of self-modification through neural-network weight changes is more common. Current methods rely on gradient descent and are useful in practice, but remain constrained by gradient descent’s limitations.
His central framing for future AI is the artificial scientist. Today’s LLMs train on the World Wide Web, whose data was preserved because humans found it interesting, making models inherently strongly human-biased. A system closer to how infants learn should generate training data through its own actions, predict consequences, build a world model, and use it to plan the next step. He calls this artificial curiosity: once basic needs are met, the system proactively asks questions, designs experiments, and collects data rather than only answering human questions.
He considers robot hardware the most underestimated bottleneck. Current hardware is far behind the human body, he said, especially because no artificial technology approaches the hand’s combination of strong grasping and fine manipulation. His exact point was: “You can’t have AGI just with things behind a screen.” Superhuman chess systems and text systems that pass the Turing test can exist on screen, but without mastery of the real world, they are not physical AGI. For builders, the next stage includes hardware, world models, autonomous experiments, and closed-loop learning in real environments, not just model leaderboards.