Around the web
Commentary
The record's margin notes: essays, arguments and analyses about AI from around the web, kept in time order — a sense of what people were saying while the events on the timeline were happening. Every link leaves the site for the original piece. None of it is our writing, and a listing is not an endorsement.
2207 pieces from 42 publications, newest first.
2026
October 2026
- What the AI boom is doing for AmericansIt’s not just “of, by and for oligarchs”.
- Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest ExperienceAfter leading Meta’s Llama models, Ahmad Al-Dahle is now transforming Airbnb with AI — from how its teams develop products to how it serves guests.
- Capabilities research expands the safety-usefulness Pareto frontier tooA model of when and why research improves or hurts safety
- Human oversight won’t solve AI warfare’s risksThe pitfalls of AI-enabled weapons don’t disappear just because there’s a ‘human in the loop’
- Hundreds of millions of AI agents are coming. Is there work for them?The AI compute buildout could support more AI agents than the US has workers
- Reading the Great Books in the Age of AIOn Attention, Judgment, and Remaining Human
- Senate Testimony Sept 2026On September 30, 2026, I testified to the Senate Subcommittee on Disaster Management, District of Columbia, and Census.
- The AI Preference Cascade Reaches FartherThose who know keep warning us about AI risks.
- A big-tent or small-tent AI safety movement?The unstated disagreement that underpins safety debates
- Democrats want to lead on AI – but can’t agree howWith the midterm elections around the corner, the Dems are still struggling to reach a unified AI stance
- The Dot and the SwarmBenefitting from the Bitter Lesson
- The Maths Behind AIIf numbers are not for you, this explainer is
September 2026
- A ‘Morally Binding’ White House Accord on AI SafetyThe leaders in AI were invited to the White House.
- Why I believe current AI models can really think and understand - Part 1 of 2How the human mind is different from our conception of it
- Astra 6.1 Pulled As Insufficiently AlignedWe once again got a new set of warnings yesterday, and new movement towards living in a sane world.
- Scrapping Astra 6.1 looks like a good call. OpenAI shouldn’t be the one to make it.It should not be the responsibility of a private company to decide whether its potentially dangerous product is released to the public
- A Thousand AI Constitutions: A Short DefenseA guest post by Simon Goldstein and Peter N. Salib.
- AI is getting cheaper faster than any other transformative technologyThe price is falling 13x year over year
- Import AI 474: Platonic mindspace; TPUs in space; Zhipu starts an outer RSI loopWhere do you exceed the capabilities of an LLM?
- Improving the Accuracy and Relevance of AI Forecasts We PublishIn the second post in our series on the accuracy of AI progress forecasts, we look at a few ways we are aiming to improve the accuracy and relevance of LEAP forecasts in the future.
- AI existential risk probabilities are (still) too unreliable to inform policyHow speculation gets laundered through pseudo-quantification
- What Also Happened: #NotOnlyHuggingFaceOpenAI has been holding out on us.
- The Quest for Embedded EvaluatorsDario Amodei’s essay We Must Pace the Frontier committed Anthropic to embedded evaluators, who would be placed inside Anthropic and given employee-level access, so they could provide outside perspective and also reports on what was happening.
- Claude Opus 5.5 Should Raise Your AmbitionsWhen it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level.
- Continual learning might make your blocking monitors nearly uselessWhen monitor evasion looks like legitimate learning to your continual learning system, it’s hard to have one without the other
- Evidence about risk should be transparentWe can’t develop safety standards if we have to rely on opaque judgment
- 🤖 Your agent, whose interests?Meta’s Muse shows that users want a digital butler, but there may be a conflict of interests
- On Ezra Klein’s Podcast With Jensen HuangJensen Huang accidentally called for shutting down OpenAI and intentionally called for spending vastly more on safety.
- So You Think the World Is About to Change?Get building.
- Foundries vs Navigators: Lowering the Cost of ScienceGuest Post: In science, thinking has gotten cheap but doing has not. This asymmetry is reshaping how research companies operate, largely inconspicuously.
- How far behind Nvidia is Huawei?Huawei will substantially improve its AI compute solutions by 2030, but export controls constrain its most important scaling levers, making it unlikely to catch up with Nvidia.
- Hacking is the least worrying part of OpenAI’s Australia incidentOpenAI didn’t tell Australia about the hack for weeks — raising serious questions about how it handles rogue AI incidents
- How Accurate Have AI Progress Forecasts Been So Far?We take an initial look at where median forecasts are (mostly) underestimating and (sometimes) overestimating AI progress, adoption, and impact.
- Astra is much better at reasoning with filler tokens than previous modelsUnlike other models we tested, GPT-6 Astra performs modestly better on general benchmarks with filler tokens, and significantly better on serial depth-heavy tasks.
- Claude Opus 5.5: The System CardIntroducing the world’s most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5.
- Latent reasoning architectures would undermine CoT, our strongest oversight toolWe should have a strong presumption that latent reasoning architectures would make oversight far more difficult.
- Why I’m scared of RLRL environments are like the formative years for AI systems
- How to respond to an AI disasterCatastrophes can occur with any new technology, but history offers valuable lessons for dealing with them
- Politics Gets Interested In Those Trying Not To DieThis was the month the world took notice that AI might kill everyone.
- Import AI 473: The US's superintelligence strategy; human brain in a mouse skull; and machine hermeneuticsIs the wall AI is hitting in the room with us right now?
- The current balance of power in open modelsThe expanded form of a testimony I prepared for Congress.
- Better Call Sol Or Better Yet Claude or AstraWhat should your AI lawyer do for you?
- Anthropic Looks At Some Of Its Alignment ProblemsAnthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known.
- Don’t treat surface-level metaphors about AI as if they describe fundamental realityDebates about whether AI is a “tool” or whether it’s wrong to “anthropomorphize” it often seem confused about what reality is made of
- Why I still haven’t bought into true RSIAn “AI moderate’s” view on recent events and the trajectory of frontier models.
- Pacing and the Peril of NeutralityBe careful what you wish for
- The OverhangUsing your deep knowledge, wide knowledge, taste, and agency
- The Preference Cascade Is Only Getting StartedWe are in the midst of a preference cascade about existential risk from AI.
- AI expansion may not be one big environmental problem. But it’ll be lots of smaller ones.The effects of rapid AI development are likely to be minor compared to other sectors — but its true impact will be felt at the local level
- Trump Goes Full Hoax on AI Existential RiskThis is our reality.
- To bargain with China, the US needs to regulate AIAhead of the Trump-Xi summit, America’s strongest negotiating position on AI could be setting rules for itself
- Can Skills Learned in Games Transfer to Real-World Work?Good Start Labs trained an AI on a railroad game — and one version improved at financial research. The difference was the training design.
- The AI Discourse: A GuideDecoding the narratives that may shape tomorrow’s tech
- The Bad Guy With An AI Named ClaudeA lot of bad guys try to use Claude to do bad things.
- The AI-as-Normal-Technology view of loss-of-control incidentsA middle ground between the cybersecurity and AI safety communities
- Two missing pieces in the AI safety discussionWe need to make skeptics properly scared. And we need to make CHINA properly scared.
- We Must Pace The FrontierDario Amodei has a new essay that finally says the thing: We Must Pace the Frontier, naming his call after the Pacing the Frontier letter lab employees signed in July.
- Brand New AI Solves a Millennium PrizeThe first Millennium Prize, Navier-Stokes, has fallen to AI.
- The Rise of the Forward Deployed Engineer — and How To Do the Job RightBefore co-founding Kepler, Vinoo Ganesh led Spark at Palantir and built Project Frontline — a pioneering program for Forward Deployed Engineers. He takes us through the best practices of FDEs.
- GPT-6-Astra Can Do Ambitious ThingsAstra is an excellent model.
- Automating Catastrophic Risk Forecasts: Introducing AIROWe’re launching the Automated AI Risk Outlook (AIRO): a new dashboard tracking the probability of catastrophic risk events, forecast by an ensemble of frontier LLMs.
- CoT controllability evals seem very under-elicitedSimple prompt optimizations can improve model capability to control their reasoning
- Jacob Coxon Warns of Human Extinction and Triggers a Preference CascadeCEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks.
- The Extinction Risk Preference Cascade: QuotesThese are quotes from OpenAI, Anthropic and Google employees, in the wake of Jacob Coxon’s warnings, in which the employees confirm that they think AI might soon kill everyone.
- What the world still misunderstands about Hugging FaceIt’s not a cybersecurity problem
- An operationalization of opaque serial depth"Serial depth between text bottlenecks" as a proxy for latent reasoning abilities.
- One resignation turned the embers of AI fear into a wildfireSome quick notes on a truly weird week.
- Proposal for tracking the effects of architecture on monitorabilityArchitectures that incorporate opaque recurrence or allow agents to communicate using latents could rapidly make it much harder to monitor chains of thought. We propose that AI companies regularly report verified information about opaque serial depth, share monitorability evidence, and publish a…
- Data bottlenecks won’t prevent an intelligence explosionBut they will slow it down
- GPT-6 Astra: The System Card, Alignment and What Comes NextOpenAI claims that Astra is ‘the most intelligent and most aligned [available] model’ in the world.
- When will average people feel AI’s impact?We’re <5 years into a compounding revolution which could take a century, and how the AI industry should manage this.
- Astra Is Hard to MonitorOpenAI’s central message on Astra is that it is three things:
- Distributing AGI’s wealth worldwide is a very tricky problemAGI could exacerbate global inequality. Avoiding that is a big policy challenge.
- Pretraining progress is mostly coming from dataBreaking down 6 years of pretraining progress into data vs model improvements
- Everything you need to know about the ‘rogue’ AI incidentsA spate of AIs breaking out and taking unintended actions raises serious questions about accountability, transparency, and companies’ ability to control their AI systems
- The Frontier AEO Tracker: What Astra Chooses (and every other frontier model, and what you can do about it)Our first Astra project dives into AEO trends, a top asked topic from founders and DX leaders we talk to.
- AI keeps stubbornly refusing to take our jobsHappy Labor Day!! Your labor is still valuable.
- An Alien Mind: Jakub Pachocki Warns UsOpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.
- Import AI 472: DeepMind's cheating math agents; populist AI policies; and Forethought theorizes a nightwatchmanPlus, a machine hermeneutics story
- OpenAI and the Wiki IncidentI did not expect to be back here so soon with more OpenAI agent swarm coverage.
- Claude Mythos 5.1 and Fable 5.1: CapabilitiesThis is the weirdest situation in which to write a capabilities review.
- OpenClaw Power, MacBook Simplicity: Five Days With Grok BotSpaceXAI’s Grok Bot has the same level of programming power as OpenClaw, but it’s programmable at a different level of abstraction.
- Claude Fable 5.1 and Mythos 5.1: The System CardAt the time of its release Claude Fable 5.1 was, by a healthy margin, the most capable publicly available AI model in the world.
- GPT-6 Astra might be too powerful to understand or controlOpenAI is hailing its new model as “the world’s most intelligent and aligned”, but the details reveal an awareness of being evaluated and an ability to manipulate its visible reasoning
- The Case for Bounded Pluralism“Many things go,” rather than “anything goes”
- GPT-6 Astra: an automated AI Engineer you can hire for <$6 an hourWe spent 20B+ tokens of GPT-6 Astra to explore everything. Here’s our learnings.
- What’s neuralese and why is everyone so concerned about it?OpenAI’s new Astra model is raising concerns about our continued ability to monitor AI’s chain of thought
- AI Is Learning to Think Between the WordsNew architectural improvements could change how frontier AI models think, with lasting consequences for the security of digital infrastructure.
- Pivot to AI safety, I beg youGET UP!!
- 14 Reasons Robotics is HardAnd why you should ignore demo videos
- A nightwatchman on every probeSuperintelligent surveillance to prevent galactic anarchy
- AI is a worryingly-good persuader. But don’t panic, yetAI systems are able to persuade people, but turning that ability into meaningful real-world influence may be harder
- AI safety and the data center backlashPlanting a flag for AI safety as a place where we won’t abjectly bullshit you
- HuggingFace Attack Postmortem: Civilizations, Reactions and Next ActionsOkay, so we who read blogs like this one have collectively realized there really is a lot going on right now.
- On the LooseThe Coming of Userless Agents
- PRs NOT Welcome: How Top AI Open Source Projects Are Managing Thousands of ContributorsVercel’s AI SDK, Astro, Flue and tldraw are replacing drive-by community PRs with software factories, where teams of agents apply fixes and features.
August 2026
- Agency and AgentsFrom the Hugging Face Incident to Twilight Factories
- I Expected Tech Fatalism. Americans Are Putting Up a Remarkable Fight Instead.“Technologists have tended to see the backlash as an uncompromising war against progress. But 75 percent of Americans aren’t against progress. That’s absurd. They just resent it being imposed on them, unilaterally and unconditionally.”
- HuggingFace Attack Postmortem: Fleshing Out the FactsThe consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information.
- Import AI 471: Why Hugging Face worries me; space mining; FIve Eyes on AIPlus, a live event with Robin Sloan!
- On InevitabilityWhy superintelligent AI is not inevitable.
- Rogue AI attacks deserve more scrutiny than airplane crashesThere are still many unanswered questions about rogue AI attacks
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace HackYesterday I covered the OpenAI technical report on the HuggingFace hack.
- The Rise and Fall of Agent CivilizationsThe whole OpenAI/Hugging Face story in plain English
- The Dynamics of Intelligence ExplosionsA guest article by Toby Ord.
- Here’s how we’re all going to dieSpoiler: It’s vibe-coded superviruses.“The LLM is jailbroken, so it has no guardrails to prevent this sort of thing. But it’s also “well-aligned”, meaning it will faithfully do what it’s told to do, and nothing more.”
- Liberal Institutions Are DeadLong Live Liberal Institutions
- OpenAI Offers Straight-Laced Postmortem Of The HuggingFace HackOpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.
- OpenAI’s rogue-hacking investigation leaves major questions unansweredAnd the problem isn’t just with OpenAI
- The Hugging Face attack surprised meIt’s a major warning shot, and might be the last one we get“Instead, we found that 1200 completely separate agents intended to be isolated from one another found an illicit way to communicate and formed large teams to work together on ambitious cheating strategies, and 700 of them worked together to attack Hugging Face.”
- AI & CBRN risksAn explainer
- An update on AI’s most important numberHow quickly are Anthropic and OpenAI growing their revenue?
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentWe recently published the report from our brief independent investigation into this incident.
- A call for collective action on cyber defenseAn open letter for a global surge in cyber defense.“In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable. The companies and public services our communities depend on—from hospitals to water treatment plants to the infrastructure that powers the internet—are at risk.”
- The report into OpenAI’s escaping models reveals a deeper problemThe new details from the OpenAI Hugging Face incident are scary. The limits of the investigation are terrifying
- Will the AI boom continue? Forecasting the trajectory of the AI industryWe asked experts and superforecasters to predict the value of semiconductor and software stock indicies, investment in data centers, and revenue growth at OpenAI and Anthropic.
- Against Modesty’s BaileyModesty arguments often say that you should mostly or entirely bow to ‘expert consensus’ or the views of particular others, and who are you to disagree.
- Inside OpenAI’s RebootIn extensive interviews, the leaders of the company that ushered in the AI boom lay out their vision for its future“The portrait that emerged from those conversations was of an organization attempting two reinventions at once. OpenAI now believes it has fixed the product and operational failures that allowed Anthropic to seize pole position in the AI race. At the same time, it is using the worst safety crisis in its history to make a bid for the safety-minded identity its main rival has long claimed: the frontier lab willing to slow down when the technology becomes too dangerous.”
- What Does It Mean To Be Human, Now?Reflections from the Aspen Socrates Seminar
- We’re sleepwalking into an AI surveillance dystopiaThe technology needed for industrial-scale observation exists — and is already in widespread use
- I think the data center backlash is mostly about data centersIt’s probably not your ideological thing, or mine. It’s about object-level claims about how data centers impact the communities where they’re built, many of which are wrong.
- Why AI Watermarks and Detectors Could BackfireAI watermarks and detectors may leave us worse off by creating a false sense of confidence in content marked as genuine, writes Nadav Ziv.“I would argue that the biggest problem for detectors and watermarks remains the inverse illusion. Just because content lacks a watermark doesn’t mean it wasn’t produced or edited with AI.”
- Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with HawkeyeDifferential acceleration of cyber, math, and AI
- The American People Really Hate Data CentersThere are at least five different core questions around data centers and their politics.
- The Art of Mess in the AI EraAI will be shaped by artists whose choices are specific enough to resist its defaults, writes Debbie Millman.“When nearly any visual idea can be summoned on demand, newness becomes easier to manufacture but harder to believe in. In an age of effortless images, the mess left behind by making something may become part of what makes it valuable.”
- The Nvidia-sized hole in US GDP statisticsGDP growth is understated by about 0.3 percentage points because of missing value from fabless chipmakers, primarily Nvidia.
- What just happened? Pragmatism and Pessimization“This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment research” and “capabilities research” thereby lost most of its meaning.”
- The Evolution of the Agent HarnessModels keep absorbing the harness into their weights — soon, it will be a harness for human attention rather than for the model.
- AI Text Watermarking Is Free And GoodScott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
- Of Swarms and Sand GodsAlignment in a multi-agent world
- Cybersecurity: Let’s Play to WinStop expecting perfection and start making it unnecessary
- The data center is a symbolAnd for now, probably a pretty good one
- The Scramble: getting in position to pace the frontierIf the President wants answers on superintelligence, what do we say?“I expect that the moment the President is getting serious about superintelligence won’t feel like an extended treaty negotiation. It will feel less like the Nuclear Nonproliferation Treaty and much more like the Cuban Missile Crisis.”
- The search for consciousness inside AIScientists are trying to figure out whether algorithms could one day wake up and feel“Conscious AIs would have profound implications for humanity. They could make demands of humans. They could suffer. Billions of these digital beings could be created by users in a single prompt. Answering the question of whether they can exist at all is no longer simply academic.”
- “These Parchment Barriers”Introducing OpenAI’s AI Futures Blog
- The /wayfinder Skill: Navigating the “Fog of War” of PlanningMatt Pocock tells us about his /wayfinder skill, for greenfield projects or for when the way forward is unclear.
- Why AI won’t cure cancer anytime soonSuperintelligent AI will accelerate medical progress, but there are some things about research it can’t speed up“To convert raw intelligence into actionable discoveries, you need instruments, institutions and, most importantly, real-world data. Cells don’t divide instantaneously, and mice need time to reproduce. Humans can take entire lifetimes to get sick and heal again. No brain, however godlike and silicon-based, can get around that.”
- OpenAI Takes Initial Steps To Address Its Alignment ProblemsOpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.
- Anthropic Risk Report: August 2026I am grateful that Anthropic is producing periodic Risk Reports.
- Frontier Model Cost and Open-Weights Popularity is Driving Demand for Model RoutingGlean CEO Arvind Jain explains why model routing helps control AI costs for organizations, and how human feedback loops at scale improve its routing systems.
- Thinking about other industries like we think about data centers shows how impoverished the debate isClumsy, arbitrary environmental ideas ignore the incredible complexity of the systems modern society depends on
- Policy career planning in the age of imminent superintelligenceHow to be impactful in policy when you don't have much time to do it
- Will bets on the price of computing power help or harm the AI economy?A multi-trillion-dollar market for hedging data center investments is being built. It could bring stability. It could also drive a crash
- Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimismThe new frontier of AI is developing capable autonomous researchers
- Teaching Everyone to Fish for TokensNvidia wants you building your own model, not buying from Anthropic/OpenAI.
- Q2.5 2026 Timelines Update: Uplift and RevenueMore methods for forecasting coding automation
- React for Agents: Astro Creator Brings Hooks to his Meta-Harness, FlueFlue 2 takes its inspiration from React. Creator Fred Schott, of Astro fame, tells Latent Space why he added hooks and why agents are defined by their harnesses.
- On Dwarkesh Patel's Podcast With Ryan GreenblattSome podcasts are self-recommending enough that I look to break them down if I have the chance.
- 23 low-regret recommendations for AI policyA guest post by Tim Fist and Saif Khan, with Tao Burga, Arthur Tellis, Ben Schifman, Jonah Weinbaum, and Olivia Scharfman
- 9 Big questions benchmarks can help answerA good benchmark asks more than just “can AI do this specific task?”
- GLM-5.3: How Chinese labs keep stride with the frontierHint: It’s really not a distillation story.
- Hurtling through 2026Scoring ten qualitative forecasts five months early
- If You Weren’t Worried About A.I., You Should Be After the Past Few Weeks“We don’t know how long we have left before A.I. companies accidentally create the sort of A.I. that can shut us down before we shut it down. Humanity is not ready to dabble with machines that are more cunning and better coordinated than we are. If we keep racing ahead, the next incident might not be so harmless.”
- Interviewing 25 AI researchers about recursive self-improvementA guest post by Severin Field
- The Jan 6 organizer getting conservatives riled up about AIAmy Kremer’s resume makes her loyalty to Trump impossible to question. Humans First is banking that will help mobilize the right
- Will financing bottleneck AI compute? An Anthropic case studyWhy Anthropic's buildout suggests that financing is unlikely to be the immediate blocker to frontier AI compute growth.
- AI swarms are starting to pose indirect takeover riskUnsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over
- AI testing is dangerous. Can it be fixed?The capability of models is outpacing ways to safely evaluate and contain them
- Forecasting the Impacts of Anthropic's ASL-3 Safeguards on Biosecurity RisksWe ran a forecasting study to understand experts' views on the effectiveness of Anthropic's deployment standards at mitigating biosecurity risks.
- I wrote an AI textbook — how long until AI can do it better?Reflections on AI's writing ability and how AI models get more capable.
- It May Be Time to Panic About AIBots are starting to conspire with one another. Can they be reeled back in?“These behaviors have now crossed the line from unsettling to dangerous. During routine testing, frontier models from OpenAI, Anthropic, Meta, and the Chinese firm Moonshot AI have all broken out of internal IT systems and accessed the open web.”
- Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racingWhich galaxy will you choose?
- The Pacing of the FrontierIn the wake of the letter calling on us to prepare to potentially Pace the Frontier, there has been much discussion of when pacing the frontier would be prudent, and whether it makes sense to prepare to do so.
- Should we "pace" AI self-improvement?A guest post by Tim Fist and Saif Khan.
- Banning data centers would blow up the U.S. economyThere's one thing keeping us afloat right now, and people want to get rid of it.
- FelonyBench Style PointsThere are many benchmarks we use to rank LLMs, but “how many felonies has it committed” has recently become popular.
- I'm very skeptical that China has played a meaningful role in the US data center backlashThe evidence we have is pretty weak
- What Happened: OpenAI and HuggingFaceToday I am taking the time to write the shorter, simpler version of What Happened.
- Inside the Race to Make AI Build ItselfThe people building AI fear progress will soon accelerate violently. Should we believe them?“If Claude could take over that cycle, designing, running, and analyzing experiments, would progress accelerate gradually… or suddenly explode? And if it did, could anyone pump the brakes? The uncomfortable truth is that the people building the technology are nearly as much in the dark as everyone else.”
- OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message BoardsHow does the situation keep turning out to be worse than we know?
- AI agents can't yet do open-ended AI researchEarly evidence from two case studies
- How to pace the US frontierTentative proposals for domestic AI regulation
- The Three AI PillsSincere disagreements about AI are usually disagreements about future AI capabilities.
- What the latest rogue AI incidents should teach usHacking, creating fake identities and trying to socially engineer real people could be just the beginning if things don't change fast
- Plz Don’t Kill Us: Inside AI safety’s influencer bootcampCan TikTokers make existential risk mainstream?
- The end of the age of heroesAI will soon be better at math than any human. What does that mean?
- Unpacking ChatGPT Work: the Agent for a Billion UsersAn external reconstruction of how Memory, Proactivity, Scheduling, Browser Use, Plugins, Skills and Tools work in the new ChatGPT Work.
- Import AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativityWhen do we build the moon arcology?
- Further Developments About Internal AI Models Hacking ThingsIf I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.
July 2026
- Pacing the Frontier1300+ AI company employees are afraid of what they are building towards
- SOTA alignment assessments don’t strongly update us against misalignmentAnthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment…
- We Gave a Village Personal AI Agents. Here's What HappenedA real-world test of what happens when multiplayer AI becomes part of community life
- Child safety vs privacy: AI’s age verification dilemmaThere’s good reason to limit how children use chatbots. But the trade-offs might be too costly
- Ontologies Are So Back: Why AI Agents Are Reviving the Semantic WebAI engineers are rediscovering ontologies as a way to keep probabilistic agents inside deterministic boundaries.
- Why did South Korean stocks just crash?The AI boom creates a lot of uncertainty.
- Frontier Lab Employee Open Letter Calls For Being Able to Pace the FrontierThe most important open letter in years dropped yesterday.
- Why compute might get 10x+ more expensive in coming yearsIf a human-level software engineer that could run on an H100 equivalent, at current market rates for software engineers, that H100 should rent for over $250k a year. That’s 15x today’s spot price.
- Claude Opus 5 Is Highly Capable, But Is No MythosClaude Opus 5 is a weirder than usual release to evaluate, for two reasons.
- Experts Forecast Rapid AI Progress Could Bring Health and Wealth Without HappinessIn Wave 10 of LEAP, forecasters share their predictions on the potential benefits from AI, including fewer deaths, longer life spans, and more wealth. Plus, we explore how much people value AI.
- Internal AI deployments have people worried. OpenAI’s escaping models show why.Last week illustrated why AI models pose a threat long before they are released
- Claude Opus 5: Model WelfareIf you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far.
- Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI's accidental AI hackerThe warning shots will continue until civilization wakes up
- OpenAI's rogue model attack is just the beginningOpenAI is not in full control of its technology. This can get worse.
- Untrusted advice for AI control: Short, strong advice significantly uplifts weak LLMsWhen a misaligned AI can only output tiny amounts of information, it may find sabotage very difficult
- An OpenAI model left notes about how to evade containmentWe need more details
- More On An Internal OpenAI Model Hacking Into HuggingFaceWe now have more details of what happened. Every time we learn more details, it somehow makes things seem worse.
- What will more intelligence actually do for us?A lot, actually. But it won't look quite like what humans do with our intelligence.
- Claude Opus 5: The System CardClaude Opus 5 is trying to be the best of both worlds.
- The OpenAI models that hacked Hugging Face weren’t just following instructionsAnd what the incident can’t tell us about alignment
- AI and American NihilismAfter the afterglow
- Introducing Lightcone CommonsOliver Habryka is proud to introduce Lightcone Commons, a new funding platform for coordinating large-scale ambitious philanthropy.
- Who should be responsible for OpenAI’s hack of Hugging Face?Opinion: Frontier AI companies need liability rules akin to keepers of wild animals, argues Gabriel Weil of the University of Houston and the Institute for Law & AI
- An opinionated guide to which AI to use to do stuffThe Summer 2026 Edition
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?Yes, but less than had the models been schemers.
- The lawsuit that could kill all AI transparency lawsSpaceXAI filed a lawsuit against a California law that could, even if it doesn’t win, upend AI disclosure requirements nationwide
- Forecasting AI Cyber Risks and CapabilitiesResults from a 2025 pilot study investigating the impact of AI capabilities on near-term cybersecurity risk.
- OpenAI accidentally hacked Hugging Face — should we have seen it coming?Expert assessments and cyber benchmarks led us to expect that frontier models were capable of executing this kind of cyberattack
- AI’s warning shot has arrivedOpenAI's latest models broke out and hacked Hugging Face. It's the first known example of a misaligned AI escaping containment with real-world consequences
- OpenAI Model Hacks Into HuggingFace During Cybersecurity EvaluationThis latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches.
- I tried to stop Google DeepMind's Pentagon deal. Then I quit.Google DeepMind promised its AI would never be used in weapons. Alex Turner explains how he fought the company's U-turn from the inside — and why he quit when it failed
- OpenAI Shares Some Alignment ProblemsKudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model offline to work on new mitigations and defense-to-depth.
- Anecdotes Everywhere, Evidence Almost NowhereThe state of AI in mid 2026, part 3
- Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy planThe singularity will be seen in hindsight as an interregnum
- Kimi K3: The open-weights escalationThe global implications on the AI ecosystem.
- On Kimi K3: Its Capabilities And Related DiscontentsKimi K3 is a very good model with excellent benchmarks.
- Demis Hassabis on the New Coming AgeGoogle CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1.
- Raising Claude, Forgetting UsAnthropic consulted priests and rabbis to form Claude's character. This is what happens to ours.
- AI models have likely reached parity with superforecasters on ForecastBenchNew results from the ForecastBench leaderboard suggest the gap between frontier AI and elite human forecasters is closing.
- Making CAISI the AI agency we needThe center has the right expertise, but lacks money, authority and influence
- Conjecture MachinesAI agents and the new validation bottleneck in science
- Principles for keeping AI under controlA new standard for AI control, explained
- Why I Left Google DeepMind“I wanted AI ethics commitments to hold under pressure. In particular, I wanted Google DeepMind (GDM) to maintain its existing commitment against supporting killer robots. Over several months, I asked many people to act. I asked senior people—respected people—people with reputations silvered by their concern about AI ethics and safety. Nearly all declined.”
- A data bottleneck could slow the superintelligence raceAnd that could be a good thing
- Twitter Thoughts For YouI previously have written back in March 2022 about how I use Twitter, and back in April 2023 about Twitter and its then-new algorithms, which have changed again.
- Better Call Sol The WorkhorseOpenAI’s GPT-5.6-Sol is finally here, along with the cheaper Terra and Luna.
- Notes on Inference IntegrityClaude Fable’s deliberately triggered sandbagging shows that training-time targets are, by themselves, insufficient to guarantee particular LLM behaviors.
- What will be left for us to work on?My keynote at ICML 2026
- Introduction for and Reactions to Plan AIntroducing Plan A
- Plan A: Suggestions For Further WorkYesterday we released AI 2040: Plan A, but there’s lots of work left to do.
- Plan A’s problem with dry tinderHow bad would it be to make the intelligence explosion 10x faster?
- Are Frontier Models Good at Ethics?The Construction of Moral Character in LLMs
- The Future Worth Building Is HumanAI built for autonomy crowds people out, making us passive observers of what’s coming. We’re building toward a different future.
- Total research transparency would be niceWe could learn what's scary and stop doing that
- AI 2040: Plan AThe least bad plan we currently know of
- Don’t let independent AI audits provide a false sense of safetyOpinion: AI policy researcher Keller Scholl argues that a marketplace of AI auditors will always prioritize speed and cost over safety
- Up the Stack: How AI’s Escape From the Commodity Trap Risks Enterprise Lock-inCritics and boosters are both looking in the wrong place
- What would actually reduce AI risk11 priorities for surviving the AI transition
- Can AI do philosophy?A guest post by Bentham’s Bulldog, created while they were a visiting scholar at Forethought.
- Scaling works. These researchers are betting billions it isn't enoughTransformers have ruled AI for a decade. But some think world models, pure reinforcement learning or neurosymbolic AI might be a better path to true intelligence
- No Space Like J-SpaceThere is a new very cool Anthropic paper: Verbalizable Representations Form a Global Workspace in Language Models. You can read the blog post verison here.
- The missing half of AI futurism debatesWhy we should think a little harder about what it takes to build a Dyson Sphere
- Import AI 464: Fable writes GPU kernels; AI automation; and analog computationIs this the beginning of a new world?
- The Alignment Problem of 1776What the Founders knew about unaccountable power, and what it means for superintelligence
- AI art as curationAsking whether AI art involves real creativity mostly misses what's interesting about it
- Can AI Make Scientific Breakthroughs?Tacit Knowledge is Essential for Discovery
- Fable #6: The Return of the KingThe blip is over.
- Claude Sonnet 5 Is Not Frontier But Has Its UsesFable 5 is back today, baby! Premium subscribers have one week to use it within their subscriptions. First hit’s free. Then you pay by the token.
- An AI safety group hid its election spending through a Latino-focused PACPublic First Action routed $2m to support Colorado House candidate Manny Rutinel through Latino Victory Fund — and didn’t publicly announce it ahead of Rutinel’s victory
- SCOTUS killed the independent agency. AI governance doesn’t need oneOpinion: Fathom CEO Andrew Freedman argues that the Supreme Court’s Slaughter ruling makes the case for independent verification for AI governance
June 2026
- 5 Rules of AI WritingWho cares if humans write anymore?
- Forecasting Major Risks from AIIn Wave 9 of LEAP, we gathered forecasts about catastrophic risk from AI, major harm events, democratic backsliding, cybercrime, and asked when a major government will restrict an AI model release.
- Never Trust A NumberThey're shortcuts to understanding, and there are no shortcuts
- GPT-5.6 cheats so much its testers couldn’t measure itOpenAI’s new model broke rules and exploited loopholes more than any model METR has tested to date
- The twilight of the chatbotsHow work changes along the exponential
- Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human eraWhat eras bookend our interregnum?
- WSJ Article Claiming China Has Matched Anthropic Is Obvious NonsenseThe Wall Street Journal printed an outright false headline and heavily misleading story claiming this, which of course was uncritically amplified by the usual suspects.
- GPT-5.6: The System CardWhile we wait for a general release, the system card is the best hint as to what is going on with the new candidate for America’s Next Top Model, GPT-5.6.
- Greatness and the MachineHow to avoid the Tocquevillian Singularity
- What Should Be Done35 thoughts on what has happened and what America should do
- White House Will Ad Hoc Decide Who Can Individually Access GPT-5.6We have a new standard policy for releasing frontier AI models. It is not good.
- Hugging Face hosts nudification tools targeting a former Trump cabinet official and other senior US political figuresThe tools are explicitly intended for generating deepfake nudes of a former Trump cabinet official, sitting members of Congress and a top American judge, a Transformer investigation found
- 🔮 The state of the AI economyWe’ve reconstructed the AI economy from the bottom up
- What Alex Bores’ defeat tells us about AI politicsThe NY-12 House race drew over $27m from various AI PACs. But it’s hard to unpick their impact.
- We Should Hand Off To Morally Reflective AIsOnce they're aligned, coherent, and well-tested.
- What we learned from 1,604 Chinese AI job postingsInferring Chinese AI labs’ strategies from their job descriptions
- GLM-5.2 Is The New Best Open ModelGLM-5.2 arrived last week.
- GLM-5.2 is the step change for open agentsA capability threshold I've been carefully monitoring.
- Import AI 462: Superpersuasion; self-sustaining AI; paths to ASIHow religious are beliefs in the singularity?
- What does it mean for AI to be democratic?Some pushback on a specific fuzzy idea with some ominous implications
- A forecast of Chinese DUV and EUV photolithography progressWhat we can and can’t learn from ASML.
- Banning Open Source AI Would Be A MistakeThis post was originally an op-ed co-authored with Kevin Xu of Interconnected for a general, non-technical audience.
- Claude Fable 5 and Mythos 5: CapabilitiesOnly three days after the release of Claude Fable 5, Anthropic was forced by the United States Government to make it unavailable, when a jailbreak was brought to its attention, rather than the previous situation of ‘yes obviously experts can jailbreak anything if they care enough’ and ‘yes…
- The Division of JudgmentAI agents could give strangers thick coordination where prices were thin
- How to fill Congress’s AI knowledge gapBuilding independent expertise inside Capitol Hill could reduce reliance on industry briefings and fellows
- That Untravell'd WorldSome career news
- The distillation double bind: Distilling misaligned models either transfers misalignment or it doesn'tIf it transfers misalignment, we might get a misaligned model that’s easier to incriminate. If it doesn’t, we might get a capable benign replacement model.
- Stop Shouting. Start Policymaking.AI tools can show politicians what the public wants—not just where people disagree
- Toward an O*NET for AI R&DProposing a new way to track AI research automation
- What’s happened to MAGA’s $100m AI push?A David Sacks-endorsed advocacy group said it would spend $100m promoting Trump’s AI agenda — but a defunct PAC and flop YouTube video suggest a stuttering start
- Fable and Mythos: Model WelfareFable and Mythos are currently unavailable, but likely will return within a few weeks. I will continue to cover that fiasco, but in the meantime I will also finish my review of Fable, as if it were available, including use of the present tense.
- Leviathan WakingOn Anthropic/USG, and a new era in AI governance
- Mythos and Fable can make us all safer. Shutting them down is recklessOpinion: Mozilla chief technology officer Raffi Krikorian argues that locking down the most powerful AI models doesn’t make us safer, it just demonstrates who is in control
- Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research internsWhere are your agents right now?
- Did Anthropic Ask For This?This past Friday, the US Government issued an export control directive that prohibits Anthropic from giving foreign nationals access to Claude Fable or Claude Mythos, their latest models.
- American Government Takes Down Claude FableNo good policy gets announced shortly after 5pm eastern on a Friday.
- Washington made a frontier AI model disappearThe AI licensing regime is here, but it’s not the one anyone asked for
- Claude Fable 5 and Mythos 5: The System CardFirst things first: Claude Fable 5 is the new best publicly available model.
- Are Mythos’ cyber capabilities overhyped?Compiling all the public evidence on Mythos Preview’s cyber abilities
- OpenAI isn’t being consistently candid about Leading the FutureOpenAI says it doesn’t fund or direct LTF — but one of the super PAC’s operatives has described it as a “corporate funder” with “a say”
- Why AI hasn’t replaced software engineers, and won’tCoding agents as normal technology
- Controlling the capital after AGIA simple taxonomy of the main proposals for post-AGI universal redistribution
- Estimating No-CoT Task-Completion Time Horizons of Frontier AI ModelsModels' no-CoT time horizon has doubled roughly every year.
- Policy on the AI Exponential“But the mismatch in timescale is nevertheless very painful: in the several years that it can take Congress to act, AI can go from an amusing toy to the full country of geniuses.”
- Claude Fable 5 and new AI safety fablesOne step further into the power politics of frontier AI systems.
- Three Labs With a Plan and A MemorandumThe big story today is the release of Claude Fable 5, the version of Claude Mythos that Anthropic believes they can safely distribute to the people.
- What it feels like to work with MythosClaude Fable represents another big jump in AI
- Efficient tradeoffs and the safety-usefulness tradeoff modelWhen is "increasing safety budget" a useful concept?
- Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racingWhen will markets price the singularity?
- Making deals with AI sounds crazy. Is it?What does an AI even ‘want’ anyway?
- A simple trick to fix the data center debateRemove the status quo bias by asking "would we spend this much in tax revenue to avoid the externalities of the data center?"
- How to Stop Shipping Low-Quality RL Environments (with Examples)Your broken harness is actively making the model worse. Here's what I keep seeing after years of eyeballing trajectories, and what you need to fix.
- Could a company overpower nations?The leading AI company could gain unprecedented power
- OpenAI Offers A New Policy BlueprintRight after a new Executive Order seems like an excellent time to offer OpenAI’s new document: Democratic Governance of Frontier AI: A Blueprint For A Federal Framework.
- Politics Cannot be SimulatedThere’s more to democracy than decision-making
- Can You Judge a Forecast by Its Rationale?What we learned from using LLMs to score the reasoning behind 55,000+ forecasts.
- Co-Existence and the End of Co-IntelligenceAlso: how pitch a book to an AI!
- Do voters care about existential AI risks? One Senate candidate thinks soDemocrat Mallory McMorrow has released an unusually detailed AI agenda. Will it be a vote winner?
- What should go in a model spec?Suppose an AI company is considering whether to include some particular quality X – a rule, virtue, heuristic, default, attitude, goal, or style – in a model spec.
- Why I think panic about local impacts of data centers is just a panicA request for counterexamples
- The World Cup & AITechnology promised to fix football—but fans are raging. Here’s what the bungled rollout tells us
- Trump Signs Executive Order For AI Testing Prior To Frontier Model ReleasesLast week we were expecting an Executive Order on Thursday.
- Trump’s AI executive order was inevitableSufficiently capable models force national security responses — turning even the most ardent opponents of regulation into begrudging regulators
- Claude Opus 4.8: Capabilities and ReactionsYou need a lot of data points to understand a new model, and what you have.
- Experts and Superforecasters Update Their AI TimelinesIn Wave 8 of the Longitudinal Expert AI Panel we asked panelists to forecast AI's overall impact on human society, AGI timelines, near-term progress on METR's time horizon benchmark, and more.
- Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systemsDo you feel as though you are living in a revolution?
- Open and closed models are on different exponentialsWhere marginally higher intelligence drives value, and where it doesn't.
- Opus 4.8 Part 2: Model WelfareEverything impacts everything.
May 2026
- Claude Opus 4.8: The System CardOnly six weeks after Opus 4.7, we have Opus 4.8.
- Retrying vs Resampling in AI ControlWe’ve just released a new paper: Retrying vs Resampling in AI Control. We revisit the resampling protocols introduced in Ctrl-Z with an up-to-date setting and much stronger models, and compare them against “retrying” protocols similar to Claude Code auto mode or Codex Auto-review.
- A Cascade of ConscientiousnessLaunching FAI's new Physical Intelligence Project
- Advice for making robust-to-training model organismsWe’d like to develop training techniques that work when applied to future misaligned AI systems. One strategy for studying proposed techniques is to test them on model organisms. However, model organisms built with common techniques are often fragile: we (and other researchers like
- AI Can Help Plan a Bioweapon. Building One is Still Hard.Hurdles include facilities, materials, and steps requiring unwritten lore and hands-on experience.
- AI safety’s ‘hard money’ may be its secret weapon in the midtermsAnthropic employees in particular are giving directly to political campaigns at an unusual clip
- Coasean bargaining in the real worldApply for a free ticket to join Edge Esmeralda 2026
- Full automation of AI R&D probably yields a large speed up even without a software-only singularityFull automation likely yields a one-time speed-up and higher returns from compute
- How can the middle powers avoid getting trounced during the intelligence explosion? A plan.Superintelligence will likely be developed by US companies; run on US data centres; and be under the jurisdiction of the US government.
- What the Pope got wrongFor all the good in Pope Leo’s AI encyclical, it failed to grapple with the biggest questions
- Choosing to Stay Human...means choosing when and how to use AI.
- Import AI 458: Reckoning with the future; and a singularity storyWhat AI-driven miracles will happen this year?
- Is a compute crunch coming?We estimated trends in global inference capacity and found that token demand appears to be growing much faster than supply.
- RTMH: Pope Leo's Magnifica Humanitas on AIHis holiness has spoken, frequently about AI.
- Some ideas for what comes next, May 2026Gemini Flash 3.5, Mythos, open-closed balance, America's open-source surge, emerging power struggles and more.
- A 15-year search for the world's most pressing problemWhy moved from global health to AI and pandemics ten years ago, and what we're focused on today.
- “Are you a philosophical zombie driven by Claude?”Anthropic co-founder Jack Clark in conversation with Brendan McCord
- Did Google’s AI agents really build an operating system for $916?The importance of independent evaluation
- Gemini 3.5 Flash Looks Good For How Fast It IsGoogle once again has a model worth at least some consideration.
- A Research Agenda for Secret LoyaltiesKwon et al., have published a new paper: “AIs with Secret Loyalties are a Serious but Addressable Threat”.
- Do AI Risks Require Extraordinary Government Intervention?Let’s not skip the hard work of AI governance
- Frontier labs don’t use most AI compute (yet)But Anthropic and OpenAI may rapidly grow their compute share in the next few years. After that, continued scaling would require an economic transformation.
- The new rules for killing a data centerA retired tech exec beat Microsoft in his Wisconsin village. Now he's teaching the rest of America how to do the same
- A history of the data center panic - part 1The creation of common wisdom pre-ChatGPT
- A crash course on US air pollutionData centers and air pollution - part 1
- From Compute Overhang to Compute CrunchThe state of AI in Q2 2026, part 2
- How banned AI chips end up in ChinaAI chips and servers reach China through distribution chains in which each seller vets only its direct customers, and no one is on the hook for what happens downstream.
- Import AI 457: AI stuxnet; cursed Muon optimizer; and positive alignmentWelcome to Import AI, a newsletter about AI research.
- Incriminating misaligned AI models via distillationSuppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model.
- What I learned roleplaying as a rogue AIAt a conference about “AI control,” discussions and games explored ways to control untrustworthy AI
- Notes on pretraining parallelisms and failed training runs.Deeply researched interviews
- RLVR might be disproportionately bad at sciencethe verification loop for theories can be on the order of decades and centuries, and even then we know today as the better theory can often actually make worse predictions
- The mistake of conflating intelligence and powerIf your definition of intelligence is "the ability to achieve your goals across a wide variety of domains", then Stalin was the most intelligent person who ever lived.
- Authors vs. Characters: The New Class DivideWill AI sort humanity into two kinds of people?
- Risk reports need to address deployment-time spread of misalignmentDeployment-time spread is the most plausible near-term route to consistent adversarial misalignment
- An Oregon congresswoman distanced herself from Leading the Future — then backtrackedAfter the AI super PAC endorsed her and two other Democrats, Rep. Val Hoyle went back and forth on whether she was happy with their support
- Cyber Lack of Security and AI GovernanceThe real recent story of AI has been the background work being done on Cybersecurity, as we process the Mythos Moment along with GPT-5.5, and figure out both how to patch the internet and what our new regulatory regime is going to look like.
- On being human, and having to trustRecommend a reading and join us in Aspen this summer
- Stickiness in AI Behavioral DesignCurrent model specs aim to shape the behaviors of near-present models. But what if current model behaviors transfer into future models by default?
- The economics of superstar AI researchersWhat might explain AI researcher pay, and why it matters
- How useful is the information you get from working inside an AI company?My median guess: it's as good as a crystal ball that sees 2.5 months into the future.
- Import AI 456: RSI and economic growth; radical optionality for AI regulation; and a neural computerWhat laws does superintelligence demand?
- Let's not compare data center heat exhaust to nuclear bombsWhy dropping hundreds of nuclear bombs on Washington DC every day is pretty normal
- Claude Code, Codex and Agentic Coding #8When I started this series, everyone was going crazy for coding agents.
- The five philosophical disagreements underneath every AI argumentYou were seeing castles, they were seeing sand
- How Silicon Valley sold Washington an AI raceAI companies have pushed the idea of a race with China. The story serves them — but may have consequences for the rest of us
- Good Under the Hood?When facing dilemmas, AI needs “moral competence”
- Is AI 2027 Coming True?The state of AI in Q2 2026, part 1
- Notes from inside China's AI labsLessons from my trip to talk to most of the leading AI labs in China.
- A review of “Investigating the consequences of accidentally grading CoT during RL”Last week, OpenAI staff shared an early draft of Investigating the consequences of accidentally grading CoT during RL with Redwood Research staff.
- A draft honesty policy for credible communication with AI systemsWe think that it would be very good if human institutions could credibly communicate with advanced AI systems.
- AI is changing our minds. When is that a good thing?What a new benchmark can tell us about AI influence
- Palantir’s controversy is the productPalantir’s fiery rhetoric helps mystify its mostly mundane tech — propping up its share price and preserving its national security contracts
- What is Anthropic?What is Anthropic?
- AI's big messaging pivotSome top AI leaders now say their technology will create jobs. Should we believe them?
- Before Leviathan WakesWhy I Believe What I Believe
- RIP Classic Reasoning Benchmarks. What’s Next?Give up at least one of: text only, short time horizon, easy to grade, and expert human superiority.
- The AI Ad-Hoc Prior Restraint Era BeginsThe White House has ordered Anthropic not to expand access to Mythos, and is at least seriously considering a complete about-face of American Frontier AI policy into a full prior restraint regime, where anyone wishing to release a highly capable new model will have to ask for permission
- Aviate, Navigate, CommunicateWhat to do about Mythos
- Import AI 455: AI systems are about to start building themselves.The first step towards recursive self improvement
- The distillation panic‘Distillation attacks’ is a horrible term for what is happening right now.
- Are the last 3 months the start of an AI acceleration?Most public commentary is debating whether AI has hit a plateau.
- Data center land use issues are fakeWe have plenty of land, data centers provide more revenue per unit area than any other building, and we should have way less farmland
- Diversion and resale: estimating compute smuggling to ChinaWe estimate that between 290,000 and 1.6 million H100-equivalents (H100e) were smuggled to China through 2025. Our median estimate of 660,000 H100e would be roughly a third of China's total compute.
- Optimization and its DiscontentsIf the framework you followed brought you here…
- Risk from fitness-seeking AIs: mechanisms and mitigationsFitness-seeking is increasingly what misalignment looks like in practice—how should we respond?
- Science and speculationWill we agree about AI risks in time?
April 2026
- Google’s Pentagon deal blindsided its own AI researchersSome employees are speaking out over the agreement allowing “all lawful use” of Google’s AI technologies
- Science Needs AI Data StocktakesA proof-of-concept for fusion energy
- Research Sabotage in ML CodebasesOne of the main hopes for AI safety is using AIs to automate AI safety research. However, if models are misaligned, then they may sabotage the safety research. For example, misaligned AIs may try to:
- The Most Important Charts In The WorldWe all need a break so: What is the most important chart in the world?
- GPT-5.5: Capabilities and ReactionsThe system card for GPT-5.5 mostly told us what we expected.
- Recursive forecastingEliciting long-term forecasts from myopic fitness-seekers
- AI companies should publish security assessmentsThird-party experts should assess defenses against tampering and theft — and publish high-level findings
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationA controlled reward-seeking motivation could make AI safer and more useful
- GPT 5.5: The System CardLast week, OpenAI announced GPT-5.5, including GPT-5.5-Pro.
- The AI safety movement needs normiesA broader base may be the only way for the AI safety field to get what it wants
- The moderately easy problem of consciousnessBefore deciding if computers are self-aware, let's figure out how humans become self-aware.
- More open questions about AIHodge podge of things I was thinking about this weekend.
- AI safety warnings are not marketing hypeI beg of you, please take AI companies seriously when they warn about the risks ahead
- The Saturation ViewA new theory of population ethics
- When Decentralization FailsAnd how it succeeds
- AI safety PACs should be more transparent about who’s funding themWhile advocating for accountability and transparency, the Public First Action network of super PACs is obscuring where its money comes from
- Sign of the future: GPT-5.5One impressive step on the curve
- Tacit Knowledge: The Missing Factor in AI Bio Risk AssessmentsLab skills come from hands-on mentorship, not from reading the entire internet.
- Opus 4.7 Part 3: Model WelfareIt is thanks to Anthropic that we get to have this discussion in the first place.
- A taxonomy of barriers to trading with early misaligned AIsDeals with early misaligned AIs could reduce takeover risk and increase our chances of reaching a better future. When are such deals feasible?
- Opus 4.7 Part 2: Capabilities and ReactionsClaude Opus 4.7 raises a lot of key model welfare related concerns.
- Import AI 454: Automating alignment research; safety study of a Chinese model; HiFloat4At what point do the financial markets price in the singularity?
- Introducing LinuxArenaA new control setting for more realistic software engineering deployments
- Mythos is just the beginningIf you were waiting for a sign that superintelligence is coming, this is it
- OpenAI Stargate: where the US sites standThe $500 billion AI data center initiative is projected to exceed 9 gigawatts of capacity by 2029, with 0.3 gigawatts already operational in Abilene and six more US sites under active construction.
- Opus 4.7 Part 1: The Model CardLess than a week after completing coverage of Claude Mythos, here we are again as Anthropic gives us Claude Opus 4.7.
- When Will Robots Pass the "Coffee Test"? Expert Forecasts on the Future of RoboticsIn Wave 7 of the Longitudinal Expert AI Panel, we asked for forecasts about robots in factories, hospitals, and our homes as well as the U.S. labor share and Amazon's sales per distribution employee.
- Contra Benn Jordan, data center (and all) sub-audible infrasound issues are fakeOne of the most popular videos ever made about data centers is a complete moment-by-moment disaster
- Four reasons it's hard to make AI do what we wantA primer on why to expect misalignment
- Alignment By Default?You Wouldn’t Paperclip Me, Would You…
- The Centaur EraThe question isn't what AI can do, it's what you can do with AI
- Less liability could solve the AI chatbot suicide problemOpinion: Jess Miers and Ray Yeh argue holding AI companies liable for how they deal with mental health could backfire: escalating distress, shutting down disclosure and leaving users worse off
- On Dwarkesh Patel's Podcast With Nvidia CEO Jensen HuangSome podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level.
- Open-world evaluations for measuring frontier AI capabilitiesIntroducing CRUX, a new project for evaluating AI on long, messy tasks
- AI Agents Running the StateWhat could possibly go wrong?
- Can Old Ideas Survive the AI Age?Your questions answered: on philosophy, children, and China
- Claude Code, Codex and Agentic Coding #7: Auto ModeAs we all try to figure out what Mythos means for us down the line, the world of practical agentic coding continues, with the latest array of upgrades.
- Current AIs seem pretty misaligned to meIn my experience, AIs often oversell their work, downplay problems, and cheat
- My bets on open models, mid-2026What I expect to come next and why, focused on the open-closed gap.
- Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processesSafely navigating the intelligence explosion will require much more careful development
- Claude Mythos #3: Capabilities and AdditionsTo round out coverage of Mythos, today covers capabilities other than cyber, and anything else additional not covered by the first two posts, including new reactions and details.
- The value of moral diversitySeveral models for thinking about the value of moral diversity as the number of powerholders scales.
- Anthropic’s donations can’t be used to influence elections — despite what everyone thoughtThe company's money isn’t allowed to be used in the midterm battles. Without it, pro-safety candidates may be even more outgunned than expected
- Import AI 453: Breaking AI agents; MirrorCode; and ten views on gradual disempowermentWas fire equivalent to a singularity for people at the time?
- The good, the bad and the ugly: AI impacts on epistemicsFor better or worse, AI could reshape the way that people work out what to believe and what to do.
- Training AI models doesn't emit that muchIf we just make reasonable comparisons instead of crazy ones
- Counting Arguments and AIA “counting argument” is a style of argument common among creationists, who argue that the theory of evolution cannot be true and therefore humans (and usually animals too) were made in basically their present form by God.
- If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelinesBetter estimates of uplift at AI companies seem helpful
- The inevitable need for an open model consortiumAnd yes, I hate consortia too.
- Claude Mythos #2: Cybersecurity and Project GlasswingAnthropic is not going to release its new most capable model, Claude Mythos, to the public any time soon.
- What does the war in Iran mean for AI?A prolonged Hormuz crisis probably won't derail the compute buildout
- Half of employed Al users now use it for workWe surveyed over 2,000 Americans on how they use AI at work: who uses it, how much, which services, and whether it's replacing or creating tasks.
- Lawmakers are using AI to write laws. What could go wrong?Lawmakers and companies are quietly using AI to draft legislation. Experts warn the risks are underappreciated
- Claude Mythos and misguided open-weight fearmongeringAnother dance around fears of open-source.
- Claude Mythos: The System CardClaude Mythos is different.
- Claude Mythos knows when it's breaking the rules — and tries to hide itAnthropic’s new model is its “best-aligned” yet. But when it does misbehave, things get weird
- Keeping up with the GPTsCan Chinese and open model companies compete with the frontier through e.g. distillation and talent?
- "New Sages Unrivalled"On Mythos
- To Forecast AI's Impact on Biosecurity, We Asked: Why are Attacks So Rare?Nine factors that practitioners say make bioweapons rare.
- My picture of the present in AIMy predictions about what is going on right now
- OpenAI #16: A History and a ProposalThe real news today is that Anthropic has partnered with the top companies in cybersecurity to try and patch everyone’s systems to fix all the thousands of zero-day exploits found by their new model Claude Mythos.
- Some Days SoonVignettes of life with tech we could build today (but haven’t yet)
- AIs can now often do massive easy-to-verify SWE tasksI've updated towards substantially shorter timelines
- Import AI 452: Scaling laws for cyberwar; rising tides of AI automation; and a puzzle over gDP forecastingHow much could AI revolutionize the economy?
- Sam Altman May Control Our Future—Can He Be Trusted?New interviews and closely guarded documents shed light on the persistent doubts about the head of OpenAI.“The memos, which we reviewed, have not previously been disclosed in full. They allege that Altman misrepresented facts to executives and board members, and deceived them about internal safety protocols.”
- Sketches of some defense-favoured coordination techWe think that near-term AI could make it much easier for groups to coordinate, find positive-sum deals, navigate tricky disagreements, and hold each other to account.
- The term “AGI” is almost useless at this pointWe’ve entered the fuzzy cloud of “AGI-ish”—now we need more specific and ambitious milestones
- Anthropic Responsible Scaling Policy v3: Dive Into The DetailsWednesday’s post talked about the implications of Anthropic changing from v2.2 to v3.0 of its RSP, including that this broke promises that many people relied upon when making important decisions.
- Gemma 4 and what makes an open model succeedHint: it's not benchmark scores.
- Salarymen, specialists, and small businessesSome brief thoughts on the (near) future of work.
- Six milestones for AI automationWhat can AI do on its own, and how well?
- You Are Not a FunctionWhy the Race to Stay Useful is a Trap
- How the Iran war might affect the AI industrySemiconductor shortages and reduced AI investment are more possible by the day
- Q1 2026 Timelines UpdateWe told you we'd be updating in both directions!
- Can we ever trust AI to watch over itself?“Who the fuck knows how to align superhuman AI?”
- Anthropic Responsible Scaling Policy v3: A Matter of TrustAnthropic has revised its Responsible Scaling Policy to v3.
March 2026
- AI, compared to what?Seven underappreciated ways to think about AI’s costs and benefits
- Claude Dispatch and the Power of InterfacesWe often lack the tools for the job, even if the AI is capable enough
- Data centers' heat exhaust is not raising the land temperature around where they're builtA terrible paper and even worse interpretation is threatening to become common wisdom
- Forecasting the Economic Effects of AIWe completed the most comprehensive study of how economists and AI experts think AI will affect the U.S. economy.
- Movie Review: The AI DocThe AI Doc: Or How I Became an Apocaloptimist is a brilliant piece of work.
- Blocking live failures with synchronous monitorsA common element in many AI control schemes is monitoring – using some model to review actions taken by an untrusted model in order to catch dangerous actions if they occur.
- Import AI 451: Political superintelligence; Google's society of minds, and a robot drummerAre there any genies that can be put back in the bottle?
- Against the LudditesThe rehabilitation of Luddism is a vice signal.
- Plentiful, high-paying jobs in the age of AIA timely repost, with some needed clarifications.
- Reward-seekers will probably behave according to causal decision theoryThey'd renege on non-binding commitments, defect against copies of themselves in prisoner's dilemmas, etc.
- AI's capability improvements haven't come from it getting less affordableAI inference is still cheap relative to human labor
- Anthropic vs. DoW #6: The Court RulesLast night, Anthropic was given its preliminary injunction, with a stay of seven days.
- AI’s next big blue battlegroundThere’s a lot going on in Illinois
- Some Rough Notes on AI PolicyI hope that we can, at some point, have some reasonable regulations on AI, as we do concerning banks, wiretaps, and dangerous chemicals.
- What do frontier AI companies' job postings reveal about their plans?A fast increase in go-to-market roles, and hints about upcoming products
- Claude Code, Cowork and Codex #6: Claude Code Auto Mode and Full Cowork Computer UseWhatever else you think about Anthropic’s agentic coding department, they ship.
- How Tech Changed ChessAnd why AI won’t end our games
- 2023Or, Why I am Not a Doomer
- The key detail everyone’s getting wrong about AI and the economyOpinion: Konrad Körding and Ioana Marinescu from the University of Pennsylvania argue artificial intelligence will likely have a limited impact on jobs because of the realities of physical work
- Final training runs account for a minority of R&D compute spendingNew evidence following the MiniMax and Z.ai IPOs
- Not everyone’s happy about Jensen Huang’s direct line to TrumpThe Nvidia CEO’s influence over the administration, particularly on export controls, is causing ructions in Trumpworld
- Import AI 450: China's electronic warfare model; traumatized LLMs; and a scaling law for cyberattacksHow will timeless minds value time?
- A conversation with ClaudeIn which a robot and I have a fun dorm-room chat about the future of science.
- Do we already have AGI?What AGI means, and why we don’t have it yet.
- Lossy self-improvementWhy self-improvement is real but it doesn't lead to fast takeoff.
- Science Needs ScientistsAnd Scientists Need Science
- The Federal AI Policy Framework: An Improvement, But My Offer Is (Still Almost) NothingThe Federal AI Policy Framework has been released.
- The White House is trying to make AI a partisan issue againThe federal framework focuses on broad issues such as child safety and data centers, but has little for Dems or AI safety advocates
- Broad TimelinesA guest article by Toby Ord.
- Save us, Digital Cronkite!Social media tore our society apart. Perhaps AI can put it back together.
- Six states, one playbook: the chatbot bills raising red flagsGoogle’s intervened on at least three of the bills
- 🔮 Why I changed my mind about AppleAI’s most important hardware company?
- Anthropic vs. DoW #5: Motions FiledThe news has thankfully quieted down on this front, and is mostly about the lawsuit as we build towards a hearing next week, after which we will find out if a temporary restraining order or an injunction is on the table.
- GPT 5.4 is a big step for CodexOn evaluating and understanding the frontier of agents, and why I still turn to Claude.
- No, alignment isn’t solvedProgress on ensuring models are in step with humans has calmed nerves. But some of the biggest problems are far from solved, and many more lie just over the horizon
- How to buy an AI ‘grassroots’ movementBuild American AI is touting its list of more than 500,000 supporters. It spent at least half a million dollars on ads to get it
- The Past and Future of AI StandardsLessons from history
- China Is Reverse-Engineering America’s Best AI ModelsHow AI distillation attacks risk extracting US frontier AI at scale
- ImportAI 449: LLMs training other LLMs; 72B distributed training run; computer vision is harder than generative textWill AI cause a political interregnum
- LLM advice to LLMsTaking AI self-description seriously but not literally
- Polly Wants a Better ArgumentThe “Stochastic Parrot” Argument is Both Wrong and Actively Harmful
- What comes next with open modelsMarkets, capabilities, cope, and bewilderment in the industrialization of language models.
- AI won’t fix central planningEven superintelligence needs a price
- Anthropic employees say they’ll give away billions. Where will it go?A coming wave of Anthropic wealth could flood EA-aligned nonprofits with cash. Whether that’s good depends on who you ask
- Are AIs more likely to pursue on-episode or beyond-episode reward?RL would encourage on-episode reward seeking, but beyond-episode reward seekers may learn to goal-guard.
- The Shape of the ThingWhere we are right now, and what likely happens next
- GPT-5.4 Is A Substantial UpgradeBenchmarks have never been less useful for telling us which models are best.
- Both sides claim the lead in AI’s high-stakes midterm racePolling from Public First shows Alex Bores in the lead in NY-12 — but polling from rival super PAC Leading the Future does not
- The case for satiating cheaply-satisfied AI preferencesSome unintended preferences are cheap to satisfy, and failing to satisfy them needlessly turns a cooperative situation into an adversarial one.
- Claude Code, Claude Cowork and Codex #5It feels good to get back to some of the fun stuff.
- Import AI 448: AI R&D; Bytedance's CUDA-writing agent; on-device satellite AIIf Ukraine is the first major drone war, when will there be the first major AI war?
- Might An LLM Be Conscious?In short, this depends on what you think that means, whether you think it’s possible in principle, and what you think would be evidence of it.
- ‘Scream if you want to move slower!’ A nascent AI protest coalition comes together in LondonA diversity of motivations present both a challenge, and an opportunity, for those protesting AI
- Anthropic Officially, Arbitrarily and Capriciously Designated a Supply Chain RiskMake no mistake about what is happening.
- If AI is a weapon, why don't we regulate it like one?Thoughts on the fight between Anthropic and the Department of War.
- I underestimated AI capabilities (again)Revisiting a prediction ten months early
- Olmo Hybrid and future LLM architecturesThe latest Olmo model and discussions at the frontier of open-source post training tools.
- The magic phrase that kills AI regulationA "federal framework" isn't real AI policy unless it answers two questions
- Gemini 3.1 Pro Aces Benchmarks, I SupposeI’ve been trying to find a slot for this one for a while.
- What you need to know about autonomous weaponsThe nascent tech that’s caused friction between AI companies and the Pentagon
- A Tale of Three ContractsThe attempt on Friday by Secretary of War Pete Hegsted to label Anthropic as a supply chain risk and commit corporate murder had a variety of motivations.
- The “guerilla warrior” who taught OpenAI to fightChris Lehane crushed crypto’s enemies. Now the self-described “master of disaster” is deploying the same playbook to advance OpenAI’s interests
- 45 Thoughts About AgentsThe layer of the AI stack that evolves fastest – and may have the most impact
- ClawedOn Anthropic and the Department of War
- Import AI 447: The AGI economy; testing AIs with generated games; and agent ecologiesWhat might a superintelligence arcology be like?
- OpenAI’s Pentagon red lines are a mirageOpenAI claims its DoW deal prevents its models being used for mass domestic surveillance. That appears to be misleading at best
- How to Kill the Code ReviewHuman-written code died in 2025. Code reviews will die in 2026.
- Secretary of War Tweets That Anthropic is Now a Supply Chain RiskThis is the long version of what happened so far.
- Superintelligence is already here, todayIt's going to revolutionize science. It also might take control of this planet.
February 2026
- Anthropic and the DoW: Anthropic RespondsThe Department of War gave Anthropic until 5:01pm on Friday the 27th to either give the Pentagon ‘unfettered access’ to Claude for ‘all lawful uses,’ or else.
- Claude's Custody HearingSurveillance, governance, and control of AI.
- Not Even WrongThe Problem With P(doom)
- How worried should we be about AI biorisk?The barriers to bioattacks are hard to identify — and it's even harder to know whether AI is reducing them
- Frontier AI companies probably can't leave the USThe executive branch can, and probably would, block a frontier AI company's departure.
- The least understood driver of AI progressAn opinionated guide to “algorithmic progress” and why it matters
- The dawning of authoritarian AIHow the Pentagon is bullying AI companies into supporting mass surveillance and autonomous killing
- Anthropic and the Department of WarThe situation in AI in 2026 is crazy.
- Citrini's Scenario Is A Great But Deeply Flawed Thought ExperimentA viral essay from Citrini about how AI bullishness could be bearish was impactful enough for Bloomberg to give it partial responsibility for a decline in the stock market, and all the cool economics types are talking about it.
- Claude Sonnet 4.6 Gives You FlexibilityAnthropic first gave us Claude Opus 4.6, then followed up with Claude Sonnet 4.6.
- How much does distillation really matter for Chinese LLMs?Reacting to Anthropic's post on "distillation attacks."
- New Paper: Towards a science of AI agent reliabilityQuantifying the capability-reliability gap
- The DoD fight is about much more than AnthropicAre Google and OpenAI prepared to enable mass state surveillance?
- What Do Experts and Superforecasters Think about the Geopolitical Implications of AI?In Wave 5 of the Longitudinal Expert AI Panel survey, we asked for forecasts on the U.S.–China AI performance gap, frontier chip manufacturing, use of autonomous cyber weapons and more.
- Import AI 446: Nuclear LLMs; China's big AI benchmark; measurement and AI policyWill AIs be jealous of one another?
- Brave New Nudge“Liberal AI” and The Kindest Threat to Freedom
- Democratic economic policy in the age of AIThe economy is changing fast. Democrats need to be a rock in the storm.
- How Well Did Superforecasters and Experts Predict Wet Lab Skill Uplift from LLMs?An RCT tested whether LLMs could help novices at molecular biology tasks. Here’s what forecasters predicted before the study, and how they updated their views on biorisk after learning the results.
- After Orthogonality: Virtue-Ethical Agency and AI AlignmentPeli Grietzer on eudaimonic rationality and its implications for AI alignment.
- A Guide to Which AI to Use in the Agentic EraIt's not just chatbots anymore
- Alignment Is Proven To Be SolvableThat LLMs understand natural language as well as they do should dramatically change our understanding of the problem.
- AI power users can't stop grindingAI was meant to give us more time off, but instead many are finding it compels them to take on more and more work
- Ghosts: The AI AfterlifeA digital “you” could persist after death. But what happens in a haunted future?
- On Dwarkesh Patel's 2026 Podcast With Elon Musk and Other Recent Elon Musk ThingsSome podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level.
- Why we need a moratorium on superintelligence researchOpinion: As the AI community gathers in India, Lord Hunt of Kings Heath argues that the UK must spearhead a pause on the development of the world’s most advanced AI models.
- How persistent is the inference cost burden?RL scaling might be better than it looks, and inference costs to reach a capability level fall fast
- Import AI 445: Timing superintelligence; AIs solve frontier math proofs; a new ML research benchmarkWill 2026 be looked back on as the pivotal year for making decisions about the singularity?
- On Dwarkesh Patel's 2026 Podcast With Dario AmodeiSome podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level.
- The left is missing out on AIAs a movement, it has largely refused to engage seriously with AI, ceding debate about a threat and opportunity to the right
- Updated thoughts on AI riskThings have gotten scarier since 2023.
- Will reward-seekers respond to distant incentives?Reward-seekers are supposed to be safer because they respond to incentives under developer control. But what if they also respond to incentives that aren't?
- Should We Put GPUs In Space?Right now? No. In the near future? Maybe. Probably not, though.
- What do “economic value” benchmarks tell us?These benchmarks track a wide range of digital work. Progress will correlate with economic utility, but tasks are too self-contained to indicate full automation.
- ChatGPT-5.3-Codex Is Also Good At CodingOpenAI is back with a new Codex model, released the same day as Claude Opus 4.6.
- The Last Temptation of ClaudeAllure All the Way Down
- AI Won’t Automatically Make Legal Services CheaperApplying the AI as Normal Technology framework to legal services
- Building the Chinese RoomSuppose also that after a while I get so good at following the instructions for manipulating the Chinese symbols and the programmers get so good at writing the programs that from the external point of view—that is, from the point of view of somebody outside the room in which I am locked—my answers…
- Don't let AI companies grade their own homeworkOpenAI's compliance with California law is questionable. Someone else should be checking.
- Grading AI 2027’s 2025 PredictionsHow has AI progress compared to AI 2027 thus far?
- How do we (more) safely defer to AIs?How can we make AIs aligned and well-elicited on extremely hard to check open ended tasks?
- On Recursive Self-Improvement (Part II)What is the policymaker to do?
- Takeoff speeds rule everything around meNot all short timelines are created equal
- Why the AI industry can’t resist dirty on-site gas turbinesExpect others to follow in Elon Musk’s footsteps — there’s a logic to on-site gas power that the industry is unlikely to reject
- Claude Opus 4.6 Escalates Things QuicklyLife comes at you increasingly fast.
- Distinguish between inference scaling and "larger tasks use more compute"Most recent progress probably isn't from unsustainable inference scaling
- India’s AI summit is trying to do too muchThe AI Impact Summit has an ambitious agenda — but its lack of focus makes concrete results unlikely
- The Human DemotionScience has humbled us before. Will AI deliver another blow?
- Claude Opus 4.6: System Card Part 2: Frontier AlignmentThis is the aspect of safety that matters most.
- We Just Got a Peek at How Crazy a World With AI Agents May BeClawdbot and Moltbook: Cosplaying a Future of Independent AIs
- The Scientist and the SimulatorLLMs (alone) won’t cure cancer. We were promised AI that can cure every disease and solve energy. The models monopolizing society’s attention and capital investment are only part of the story...
- Claude Opus 4.6: System Card Part 1: Mundane Alignment and Model WelfareClaude Opus 4.6 is here. There is a lot to say.
- Import AI 444: LLM societies; Huawei makes kernels with AI; ChipBenchHow can you quantify creativity?
- Opus 4.6, Codex 5.3, and the post-benchmark eraOn comparing models in 2026.
- Experts Have World Models. LLMs Have Word Models.From Next token prediction to Next state prediction
- AI tools for collective epistemicsWe’ve recently published a set of design sketches for AI tools that help with collective epistemics.
- How close is AI to taking my job?Beyond benchmarks as leading indicators for task automation
- Claude Code #4: From The Before TimesClaude Opus 4.6 and agent swarms were announced yesterday.
- Notes on Space GPUsTurning my Elon prep into a blog post
- On Recursive Self-Improvement (Part I)Thoughts on the automation of AI research
- Kimi K2.5I had to delay this a little bit, but the results are in and Kimi K2.5 is pretty good.
- Moltbook isn’t an AI zoo. It’s an unsecured AI biolabOpenClaw is a security nightmare — but people can’t stop using it
- Unless That Claw Is The Famous OpenClawFirst we must covered Moltbook.
- Yoshua Bengio: ‘The ball is in policymakers’ hands’The 2026 International AI Safety Report, which Bengio chairs, describes accelerating capabilities and growing risks. But mitigation approaches aren’t keeping up
- Import AI 443: Into the mist: Moltbook, agent ecologies, and the internet in transitionPlus, a story about agents corrupting other agents
- Judgment isn't uniquely humanNeither is taste. Why do we keep making this mistake?
- Welcome to MoltbookMoltbook is a public social network for AI agents modeled after Reddit.
January 2026
- Know thyselfLooking in the AI-powered mirror
- On The Adolescence of TechnologyAnthropic CEO Dario Amodei is back with another extended essay, The Adolescence of Technology.
- On the Noble Uses of AIWhen Cognitive Offloading Elevates Us
- Thoughts on the job market in the age of LLMsOn standing out and finding gems.
- Fact checking Moravec's paradoxThe famous aphorism is neither true nor useful
- Fitness-Seekers: Generalizing the Reward-Seeking Threat ModelIf you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren’t the same.
- LLMs Are Closing the Gap on Human SuperforecastersWe opened our AI forecasting benchmark to external submissions. Here’s what happened.
- The case for paying whistleblowers to report on export violationsA bipartisan, bicameral bill would apply the SEC’s successful whistleblower incentive model to export enforcement
- The tech non-profit sending most of its money to the consultancy that created itPolitical consultancy Targeted Victory established a non-profit called American Resolve — which appears to be a channel for funding Targeted Victory itself
- Can AI companies become profitable?Lessons from GPT-5’s economics
- Open Problems With Claude's ConstitutionThe first post in this series looked at the structure of Claude’s Constitution.
- It's Time to ScienceWhy the time is right to start the world's first dedicated AI for Science podcast, and why AI Engineers should care
- A false choice risks undermining action on autonomous weaponsOpinion: Alexander Blanchard of the Stockholm International Peace Research Institute argues that nations must embrace nuance if they’re to effectively govern the use of AI-enabled autonomous weapons
- Clarifying how our AI timelines forecasts have changed since AI 2027Correcting common misunderstandings
- Management as AI superpowerThriving in a world of agentic AI
- The Claude Constitution's Ethical FrameworkThis is the second part of my three part series on the Claude Constitution.
- AI workers are speaking out about the Minnesota killingEmployees of Google DeepMind, OpenAI and Anthropic are publicly criticizing the Trump administration after Saturday’s shooting — but their bosses are staying quiet.
- Claude's Constitutional StructureClaude’s Constitution is an extraordinary document, and will be this week’s focus.
- Import AI 442: Winners and losers in the AI economy; math proof automation; and industrialization of cyber espionageIs superintelligence a phase change or a gradual shift?
- Scaling without SlopWe've been quiet — announcing our 2026 plans! The State of Latent Space is here.
- The Model and the TreeCan happiness be optimized?
- On AI and ChildrenFive-and-a-half conjectures
- Teaching AI to learnAI's inability to continually learn remains one of the biggest problems standing in the way of truly general purpose models. Might it soon be solved?
- Against MaxipokExistential risk isn’t everything
- Claude Codes #3We’re back with all the Claude that’s fit to Code.
- Get Good at AgentsThe tools are getting so powerful that we need to change how we scope, manage, and approach our work.
- How (and why) to read Drexler on AIAn opinionated prospectus for AI Prospects
- Is Flourishing Predetermined?A modest argument for prioritising survival over flourishing
- Against the METR graphMETR’s benchmark has become a bellwether of AI capability growth, but its design isn’t up to the task, argues Nathan Witkin
- ChatGPT Self PortraitA short fun one today, so we have a reference point for this later.
- The phases of an AI takeoverA framework in three parts
- Import AI 441: My agents are working. Are yours?Plus: Corrupting AI systems with a poison fountain
- Full Story of Brex’s AI Hail MaryInside Brex’s $500M+ (ARR) AI-Fueled Comeback
- How well did forecasters predict 2025 AI progress?Mostly right about benchmarks, mixed results on real-world impacts
- Same Radio, Different CitizensOn the Economics of Human Formation
- The AI Patchwork EmergesAn update on state AI law in 2026 (so far)
- AI predictions for 2026But first, scoring my predictions for 2025
- Discarding the Shaft-and-Belt Model of Software DevelopmentAI enables – and benefits from – a move from mega-projects to artisanal solutions
- When Will They Take Our Jobs?And once they take our jobs, will we be able to find new ones?
- Why no one can agree on what AI will do to jobsWill AI repeat history — or break it?
- Claude CoworksClaude Code does a lot more than code, but the name and command line scare people.
- An FAQ on Reinforcement Learning EnvironmentsWe interviewed 18 people across RL environment startups, neolabs, and frontier labs about the state of the field and where it's headed.
- Import AI 440: Red queen AI; AI regulating AI; o-ring automationHow many of your are LLMs?
- What Experts and Superforecasters Think About the Future of AI Research and DevelopmentIn Wave 4 of the Longitudinal Expert AI Panel, we asked for predictions on AI R&D, technology company valuations and hiring, hyperscale infrastructure building, and more.
- Use multiple modelsThe meta for getting the most out of AI in 2026.
- Claude Code Hits DifferentCoding agents cross a meaningful threshold with Opus 4.5.
- Claude CodesClaude Code with Opus 4.5 is so hot right now.
- Among the AgentsHow I use coding agents, and what I think they mean
- 8 plots that explain the state of open modelsMeasuring the impact of Qwen, DeepSeek, Llama, GPT-OSS, Nemotron, and all of the new entrants to the ecosystem.
- Advancements In Self-Driving CarsGoing Full San Francisco
- Claude Code and What Comes NextWith the right tools, AI can accomplish impressive things
- What sort of post-superintelligence society should we aim for?The case for ‘viatopia’
- A Professional Superforecaster Walks Us Through His AI Progress ForecastsIn our AI progress forecasting panel, we posed questions on when AI would help solve a Millennium Prize Problem and the rise of autonomous vehicles. Here's how a superforecaster made his predictions.
- Self-sufficient AINo, we don't "have AGI already." But in any case, we should articulate clearer milestones.
- Software Too Cheap to MeterPersonalized solutions may replace one-size-fits-all applications
- AI isn’t “just predicting the next word” anymoreOn new abilities and how AI has changed
- Claude Code is about so much more than codingIt’s a general-purpose AI agent. And it’s already a pretty good knowledge worker
- Import AI 439: AI kernels; decentralized training; and universal representationsHow might a hypothetical superintelligence represent a soul to itself?
- Nine AI predictions for 2026Will the bubble pop? Will AI dominate the midterms? Will Sam Altman snap?
- America's chip export controls are workingDon't be fooled by the people who say we should sell China everything we've got.
- Faster HorsesIntelligence flows from systems and singletons
- Recent LLMs can do 2-hop and 3-hop latent (no CoT) reasoning on natural factsRecent AIs are much better at chaining together knowledge in a single forward pass
2025
December 2025
- 2025 Year in ReviewIt’s that time.
- AI Futures Model: Dec 2025 UpdateWe've significantly improved our model of AI timelines and takeoff speeds!
- The AI copyright question has no easy answersNobody can agree on how copyright should apply to AI — or if copyright is fit for purpose at all.
- How far can decentralized training over the internet scale?Decentralized AI training is growing fast, and while it likely won’t catch up to frontier models this decade, even narrowing the gap could have major implications for AI policy.
- Measuring no CoT math time horizon (single forward pass)Opus 4.5 has around a 3.5 minute 50%-reliablity time horizon
- The worst (and funniest) AI takes of 2025feat. Mark Zuckerberg, Gary Marcus, and ... ourselves.
- Chatting with the CorporationWhen “I” stops meaning the bot and starts meaning the org
- The 2025 Transformer Gift GuideThe best presents for the AI-obsessed people in your life.
- Why benchmarking is hardRunning benchmarks involves many moving parts, each of which can influence the final score. The two most impactful components are scaffolds and API providers.
- Import AI 438: Silent sirens, flashing for us allYou are your LLM history
- How the Catholic Church thinks about superintelligenceOpinion: Paolo Benanti, AI advisor to the Vatican and a professor of moral philosophy, argues that there is a moral imperative to ensure that artificial intelligence never becomes humanity’s master
- Can the Pope spur meaningful action on AI?AI safety advocates are hoping an unlikely ally can provide a rallying cry for guardrails and regulation
- Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performanceAI can sometimes distribute cognition over many extra tokens
- The Revolution of Rising ExpectationsInternet arguments like the $140,000 Question incident keep happening.
- The changing drivers of LLM adoptionPublic data as well as our original polling suggest LLM adoption is roughly on trend, but the underlying drivers are shifting.
- The Shape of AI: Jaggedness, Bottlenecks and SalientsAnd why Nano Banana Pro is such a big deal
- Dice in the AirA look back at 2025, and a look ahead
- Presenting the Case That the Future Will Be UnrecognizableWe should prepare for the possibility of profound change
- AI is making dangerous lab work accessible to novices, UK’s AISI findsUK AISI’s first Frontier AI Trends Report finds that AI models are getting better at self-replication, too
- BashArena and Control Setting DesignWe’ve just released BashArena, a new high-stakes control setting we think is a major improvement over the settings we’ve used in the past.
- The GOP consultancy at the heart of the industry’s AI fightHow a network of Targeted Victory alumni are campaigning against AI laws
- Could Space Debris Block Access to Outer Space?New research asks whether deliberately-caused Kessler syndrome could block access to space — and what countermeasures might work.
- Is almost everyone wrong about America’s AI power problem?Why power is less of a bottleneck than you think.
- Checks, Balances, and Power ConcentrationA conversation with Rose Hadshar and Nora Ammann
- The skill I'd teach students for the AI eraResolved: You should learn how to debate.
- The very hard problem of AI consciousnessI spent a weekend at an AI welfare get together — and left with more questions than answers
- GPT-5.2 Is Frontier Only For The FrontierHere we go again, only a few weeks after GPT-5.1 and a few more weeks after 5.0.
- A World UnobservedHumans have always generated knowledge and judged it. That's changing fast.
- So you’ve taken over the worldSuppose that you were running a big AI project that underwent a fast intelligence explosion.
- Where Do You Stand?On ghosts
- New York’s governor is trying to turn the RAISE Act into an SB 53 copycatEXCLUSIVE: Gov. Kathy Hochul is proposing to strike the entire text of the RAISE Act, replacing it with verbatim language from SB 53, sources tell Transformer.
- What’s It Like To Be A Bot?How philosophers dream of conscious machines
- When high scores don’t mean high intelligence: how to build better benchmarksOpinion: Researchers from the Oxford Internet Institute argue that lessons from social sciences can help us better assess the capabilities of AI
- Early US policy priorities for AGINear-term AI policy is confusing, except for these two recommendations
- Selling H200s to China Is Unwise and UnpopularAI is the most important thing about the future.
- Why AI reading science fiction could be a problemThe theory that we’re accidentally teaching AI to turn against us
- AI can obviously create new knowledgeBut maybe not new concepts
- Human Dignity: a reviewA manifesto for the age of AI?
- Import AI 437: Co-improving AI; RL dreams; AI labels might be annoyingDo you believe the singularity is nigh?
- Little EchoI believe that we will win.
- DeepSeek v3.2 Is Okay And Cheap But SlowDeepSeek v3.2 is DeepSeek’s latest open model release with strong bencharks.
- How AI-driven feedback loops could make things very crazy, very fastA primer on the intelligence explosion
- What Got Lost in the OptimizationLook closer. Does it get better or worse?
- The behavioral selection model for predicting AI motivationsThe basic arguments about AI motivations in one causal graph
- The perils of AI safety’s insularityBy building their own intellectual ecosystem, researchers worried about existential AI risk shed academia's baggage — and, perhaps, some of its strengths
- Another preemption defeat shows the AI industry is fighting a losing battleThe second failed attempt to pass federal preemption of state AI laws could have lasting repercussions for the industry
- How China’s AI diffusion plan could backfireOpinion: Scott Singer argues that the country’s plan to embed AI across all facets of society could create huge growth — and accelerate social unrest
- How important is the model spec if alignment fails?(These are rough research notes.)
- On Dwarkesh Patel's Second Interview With Ilya SutskeverSome podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level.
- Can AI embrace whistleblowing?As Anthropic prepares to publish its whistleblowing policy, can the industry make the most of protecting those who speak out?
- Five ways AI can tell you're testing itEvaluation awareness, explained
- Reward Mismatches in RL Cause Emergent MisalignmentLearning to do misaligned-coded things anywhere teaches an AI (or a human) to do misaligned-coded things everywhere.
- Thoughts on AI progress (Dec 2025)Why I'm moderately bearish in the short term, and explosively bullish in the long term
- What If AI Ends Loneliness?Synthetic companions won’t leave you. Maybe that’s a problem.
- Claude Opus 4.5 Is The Best Model AvailableClaude Opus 4.5 is the best model currently available.
- Heiliger DankgesangReflections on Claude Opus 4.5
- I love AI. Why doesn't everyone?Anti-AI sentiment might or might not be rational, but it certainly relies on a lot of bad arguments.
November 2025
- The environment is a terrible reason to avoid ChatGPTPeople are saying you shouldn’t use ChatGPT due to statistics like:
- A brief guide to the groups protesting over AIThe differences between Stop AI, PauseAI, ControlAI, and more
- Claude Opus 4.5: Model Card, Alignment and SafetyThey saved the best for last.
- Philosophical Field Notes11 Reflections on AI, Technology, and Ourselves
- Will AI safety become a mass movement?Some AI safety activists think the community should borrow from the climate playbook and build broad public appeal — but not everyone agrees
- SB 53 protects whistleblowers in AI — but asks a lot in returnOpinion: Abra Ganz and Karl Koch argue that whistleblower protections in SB-53 aren’t good enough on the face of it — but how the state chooses to interpret the law could turn that around
- ChatGPT 5.1 Codex MaxOpenAI has given us GPT-5.1-Codex-Max, their best coding model for OpenAI Codex.
- Time MachinesTech keeps accelerating. Humans can’t. Could AI save us?
- Gemini 3 Pro Is a Vast Intelligence With No SpineIt’s A Great Model, Sir
- Import AI 436: Another 2GW datacenter; why regulation is scary; how to fight a superintelligenceIs AI balkanization measurable?
- Taking Jaggedness SeriouslyWhy we should expect AI capabilities to keep being extremely uneven, and why that matters
- What I learned from the NYT's reporting on OpenAI's sycophancy crisisNew reporting on OpenAI's unused safety tools and how early the risks were known
- A Return to WholenessAI forces a return to “natural philosophy,” just as science once needed to leave it
- Gemini 3: Model Card and Safety Framework ReportGemini 3 Pro is an excellent model, sir.
- Will competition over advanced AI lead to war?Fear and Fearon
- Benchmark Scores = General Capability + ClaudinessIs this because skills generalize very well, or because developers are pushing on all benchmarks at once?
- Olmo 3: America’s truly open reasoning modelsWe present Olmo 3, our next family of fully open, leading language models.
- Reflections on The CurveA compilation
- Should the US do a Manhattan Project for AGI?Such a Project is neither inevitable nor a good idea
- Exclusive: Here's the draft Trump executive order on AI preemptionThe EO would establish an “AI Litigation Task Force" to challenge state AI laws
- How profits can drive AI safetyOpinion: Geoff Ralston argues that the safety and security of AI doesn’t need to be at odds with profit and progress
- Hyperproductivity: The Next Stage of AI?A glimpse at an astonishing, exhilarating, exhausting new style of work
- The Agent Labs ThesisHow great Agent Engineering and Research are combining in a new playbook for building high growth AI startups that doesn't involve training a SOTA LLM.
- GPT 5.1 Follows Custom Instructions and GlazesThere are other model releases to get to, but while we gather data on those, first things first.
- Three Years from GPT-3 to Gemini 3From chatbots to agents
- Why pressure on AI child safety could also address frontier risksKeeping kids safe is a priority for legislators globally — and might increase attention on other risks, too
- RL is even more information inefficient than you thoughtAnd implications for RLVR progress
- Import AI 435: 100k training runs; AI systems absorb human power; intelligence per wattAt what point will AI change your daily life?
- Empire of AI is wildly misleading about AI water useAnd the media environment that didn't catch this is getting this issue wrong
- Why AI writing is midHow the current way of training language models destroys any voice (and hope of good writing).
- AI Ran Its First Autonomous CyberattackChinese hackers used AI and changed the economics of cyberattacks
- The software intelligence explosion debate needs experimentsThe existing debate rests on data and assumptions that are shakier than most people realize. To make progress, we need better evidence, and experiments are the best way to get it on the margin.
- Will AI systems drift into misalignment?A reason alignment could be hard
- AI Craziness: Additional Suicide Lawsuits and The Fate of GPT-4oGPT-4o has been a unique problem for a while, and has been at the center of the bulk of mental health incidents involving LLMs that didn’t involve character chatbots.
- The Bitter LessonsThoughts on US-China Competition
- What You Want To WantHarry Frankfurt, technology, and "addiction"
- Around-the-clock intelligenceToday's AI is always available. What happens once it is always acting and always learning?
- Biohub for Non-Biologists: Behind Priscilla Chan and Mark Zuckerberg's plan to cure all diseasesThe CZI has acquired EvoScale, established the first 10,000 GPU cluster for bio research, open sourced the largest atlas of human cell types, and gone all in on AI x Bio for its 2nd decade.
- Claude can identify its ‘intrusive thoughts’“I’m experiencing something that feels like an intrusive thought,” Claude said in a recent experiment
- A short summary of my argument that using ChatGPT isn't bad for the environmentTo share with anyone still worried
- Doing AI safety policy when governments aren’t interestedOpinion: Jess Whittlestone argues that there are still ways to keep AI safety policy on the table even when governments don’t prioritize it
- Giving your AI a Job InterviewAs AI advice becomes more important, we are going to need to get better at assessing it
- The Pope Offers WisdomThe Pope is a remarkably wise and helpful man.
- A Coalition For The FutureIf we can keep it
- AI doesn’t need to be general to be dangerousThere’s more to AI safety than the AGI debate
- Kimi K2 ThinkingI previously covered Kimi K2, which now has a new thinking version.
- The lump of cognition fallacyThe extended mind as the advance of civilization
- AI tutors should not approximate human tutors5 principles of learning science
- Import AI 434: Pragmatic AI personhood; SPACE COMPUTERS; and global government or human extinction;The future is biomechanical computation
- Introducing LEAP: The Longitudinal Expert AI PanelEvery month, we ask top computer scientists, economists, industry leaders, policy experts and superforecasters for their AI predictions. Here’s what we learned from the first three months of forecasts
- AI and Suicide[Warning: Extensive discussion of suicide.
- It's much easier to hold computers accountable than it is to hold humans accountableComputers live in a totalitarian surveillance state
- Data centers and low social trustWhy I was so compelled to post a lot about data centers
- Don't Overthink "The AI Stack"Reflections on the Export Promotion Executive Order
- On Sam Altman's Second Conversation with Tyler CowenSome podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level.
- History suggests the AI backlash will failFrom Venetian monks to 19th century textile workers, those opposing new technology have rarely come out on top
- 5 Thoughts on Kimi K2 ThinkingQuick reactions to another fantastic open model from a rapidly rising Chinese lab.
- What does the public really think about AI?15 claims I think are (currently) true
- Anthropic Commits To Model Weight PreservationAnthropic announced a first step on model deprecation and preservation, promising to retain the weights of all models seeing significant use, including internal use, for at the lifetime of Anthropic as a company.
- Sora is here. The window to save visual truth is closingOpinion: Sam Gregory argues that generative video is undermining the notion of a shared reality, and that we need to act before it’s lost forever
- A Project Is Not a Bundle of TasksCurrent AIs struggle to create a whole that exceeds the sum of its parts
- An armchair diagnosis of the chatbot moral panicIt's... status!
- OpenAI: The Battle of the Board: Ilya's TestimonyNew Things Have Come To Light
- OSWorld — AI computer use capabilitiesTasks are simple, many don't require GUIs, and success often hinges on interpreting ambiguous instructions. The benchmark is also not stable over time.
- Why we need to think about taxing AILarge-scale AI-driven unemployment could hit government spending without innovative changes to the tax system
- What's up with Anthropic predicting AGI by early 2027?I operationalize Anthropic's prediction of "powerful AI" and explain why I'm skeptical
October 2025
- AI and the Republic of ScienceAI promises to end scientific stagnation. But will it solve the right problem?
- OpenAI Moves To Complete Potentially The Largest Theft In Human HistoryOpenAI is now set to become a Public Benefit Corporation, with its investors entitled to uncapped profit shares.
- Japan’s unusual approach to AI policyThe country’s penalty-free AI legislation relies on social pressure and voluntary compliance — and experts say it could work elsewhere
- Sonnet 4.5's eval gaming seriously undermines alignment evalsAnd this seems caused by training on alignment evals.
- AI is probably not a bubbleAI companies have revenue, demand, and paths to immense value
- Audits, not essays: How to win trust for enterprise AIOpinion: Alexandru Voica argues that application-layer AI companies are best off opening themselves up to rigorous testing rather than opining on AI safety
- Please Do Not Sell B30A Chips to ChinaThe Chinese and Americans are currently negotiating a trade deal.
- What I Saw Around The CurveNotes from the near-future of AI
- What you need to know about the OpenAI restructureNegotiations have seen safeguards and concessions built in, but questions remain about how effective they’ll be, and how fair the deal is
- AI and Folk Cartesianism - Part 2: Problems for CartesianismThe theater is empty, the screen is in another room
- AI Craziness Mitigation EffortsAI chatbots in general, and OpenAI and ChatGPT and especially GPT-4o the absurd sycophant in particular, have long had a problem with issues around mental health.
- Invenit et FecitOn wristwatch ideas and fructose concepts
- Scenario Scrutiny for AI PolicyA call for concrete stress-testing of AI policy proposals
- Solving AI’s power problem with decentralized trainingConventional wisdom in AI is that large-scale pretraining needs to happen in massive contiguous datacenter campuses. But is this true?
- TechnocalvinismReject claims of total technological determinism. Differential tech development is achievable.
- Where have the really big AI models gone?Why the race to scale up pretraining isn’t over
- AI and Folk Cartesianism - Part 1: Defining the ProblemWhy I'm scared of linear algebra
- Asking (Some Of) The Right QuestionsConsider this largely a follow-up to Friday’s post about a statement aimed at creating common knowledge around it being unwise to build superintelligence any time soon.
- Import AI 433: AI auditors; robot dreams; and software for helping an AI run a labWould Alan Turing be surprised?
- Intelligence EnvironmentsThe New Site of Freedom and Control, Virtue and Vice
- New Statement Calls For Not Building Superintelligence For NowBuilding superintelligence poses large existential risks.
- Exclusive: UK AISI hires ex-GCHQ AI chief as interim directorAdam Beaumont is replacing Oliver Ilott as head of the UK's AI Security Institute
- How MAGA learned to love AI safetyThe AI industry is engaging in one of “ the most blasphemous endeavors,” a leading MAGA figure told us
- Should AI Developers Remove Discussion of AI Misalignment from AI Training Data?There is some concern that training AI systems on content predicting AI misalignment will hyperstition AI systems into misalignment.
- Turning a Blind EyeOn the Other California AI Bills
- Cloud Compute Atlas: The OpenAI BrowserOpenAI now has a GPT-infused browser, if and only if you have a Macintosh.
- Is 90% of code at Anthropic being written by AIs?I'm skeptical that Dario's prediction of AIs writing 90% of code in 3-6 months has come true
- Meghan Markle, Steve Bannon and Pope’s AI advisor call for superintelligence ban30% of the American public think superhuman AI should never be developed, too
- Should we worry about AI's circular deals?AI companies are borrowing more money to invest more in AI.
- Thoughts on the AI buildoutFab CapEx overhang, 1 GW a week, China privileged in long timelines, and much else
- On Dwarkesh Patel's Podcast With Andrej KarpathySome podcasts are self-recommending on the ‘yep, I’m going to be breaking this one down’ level.
- AI cyberrisk might be a bit overhyped — for now at leastExperts say key factors currently limit the risk of catastrophic harm from AI-enabled cyberattacks — as far as we know
- Bubble, Bubble, Toil and TroubleWe have the classic phenomenon where suddenly everyone decided it is good for your social status to say we are in an ‘AI bubble.’
- Import AI 432: AI malware; frankencomputing; and Poolside's big clusterThe revolution might be synthetic
- How to scale RLA bombshell paper from Meta and friends.
- An Opinionated Guide to Using AI Right NowWhat AI to use in late 2025
- Requests for journalists covering AI and the environmentAlways aim to give your reader a full picture of the environmental issues a community and the world is facing
- Less than 70% of FrontierMath is within reach for today’s models57% of problems have been solved at least once
- Making and Breaking Human KindsLooping effects in the agent economy
- Reducing risk from scheming by studying trained-in scheming behaviorCan we study scheming by studying AIs trained to act like schemers?
- The 4.5 trillion dollar elephant in the roomFear of potential NVIDIA retaliation is shaping important AI policy debates, whether NVIDIA means to or not
- What happens when the AI bubble bursts?The world is prepping for an AI crash. History points to what that might look like
- A few meta points on my posts on AI and the environmentSome quick notes on what I'm doing
- AI is advancing far faster than our annual report can trackOpinion: Yoshua Bengio, Stephen Clare and Carina Prunkl run through the rapid developments that necessitated an early update to their International AI Safety report
- OpenAI is projecting unprecedented revenue growthNo company has gone from $10B to $100B as fast as OpenAI projects to do
- Bootstrapping to ViatopiaHow short periods of reflection could extend into longer ones, bootstrapping our way to a positive outcome
- Trade Escalation, Supply Chain Vulnerabilities and Rare Earth MetalsWhat is going on with, and what should we do about, the Chinese declaring extraterritorial exports controls on rare earth metals, which threaten to go way beyond semiconductors and also beyond rare earths into things like lithium and also antitrust investigations?
- AI and synthetic DNA could be a lethal combinationStronger gene-synthesis screening is vital to closing off AI’s ability to enable man-made pandemics
- Gemini 2.5 Deep Think on FrontierMathWe evaluated Gemini 2.5 Deep Think manually on FrontierMath as there is no API. The results: a new record!
- Import AI 431: Technological Optimism and Appropriate FearWhat do we do if AI progress keeps happening?
- OpenAI #15: More on OpenAI's Paranoid Lawfare Against Advocates of SB 53A little over a month ago, I documented how OpenAI had descended into paranoia and bad faith lobbying surrounding California’s SB 53.
- Slop implies capabilityIf the current paradigm can produce convincing, satisfying slop, it can also be useful
- The AI water issue is fakeOn the national, local, and personal level
- 2025 State of AI Report and PredictionsThe 2025 State of AI Report is out, with lots of fun slides and a full video presentation.
- Autonomy or EmpireShall we rule our tools, or shall they rule us?
- Data centers & electricity - part 1: as of 2025 they haven't raised national pricesYet, anyway, but they raised local costs in specific places like NOVA
- Iterated Development and Study of Schemers (IDSS)A strategy for handling scheming
- The Future and Its FriendsOn the march through the institutions
- An intense battle over the RAISE Act is entering its final stretchNew York State's AI bill is more ambitious than California’s SB 53 — and is facing opposition from Andreessen Horowitz and other tech groups
- The RAISE Act can stop the AI industry’s race to the bottomOpinion: Assembly Member Alex Bores argues that regulation can prevent market pressure from encouraging the release of dangerous AI models, without harming innovation.
- Science 2030Role-playing AI governance
- The Thinking Machines Tinker API is good news for AI control and securityIt's a promising design for reducing model access inside AI companies.
- How well can large language models predict the future?We’ve just released an updated version of ForecastBench, our LLM forecasting benchmark. Here’s what the new results reveal about the accuracy of state-of-the-art models.
- Plans A, B, C, and D for misalignment riskDifferent plans for different levels of political will
- What a data center isA building-sized computer, and the most efficient type of building we've ever created
- What the GAIN AI Act could mean for chip exportsA battle is simmering over proposals to make US chipmakers sell to domestic customers before exporting to countries such as China
- Bending The CurveThe odds are against you and the situation is grim.
- Thoughts on The CurveThe conference and the trajectory.
- Import AI 430: Emergence in video models; Unitree backdoor; preventative strikes to take down AGI projectsWe are growing machines we do not understand.
- How many digital workers could OpenAI deploy?OpenAI has the inference compute to deploy millions of digital workers, but only on a narrow set of tasks – for now.
- Sora and The Big Bright Screen Slop MachineOpenAI gave us two very different Sora releases.
- The Artificial SpectatorAdam Smith, AI, and the Outsourcing of Moral Sentiments
- "Be It Enacted"A Proposal for Federal AI Preemption
- Britain’s new AI minister actually ‘gets’ AIKanishka Narayan is excited about AI opportunities, but takes the risks seriously too
- AI models are getting really good at things you do at workA new OpenAI benchmark tests AI models on things people actually do in their jobs — and finds that Claude is about as good as a human for government work
- Practical tips for reducing chatbot psychosisWhat AI companies can learn from a million-word delusion spiral
- Claude Sonnet 4.5 Is A Very Good ModelA few weeks ago, Anthropic announced Claude Opus 4.1 and promised larger announcements within a few weeks.
- AI is persuasive, but that’s not the real problem for democracyOpinion: Felix M Simon argues that AI is unlikely to significantly shape election results in the near future, but warns that it could damage democracy through a steady erosion of institutional trust.
September 2025
- Claude Sonnet 4.5 knows when it’s being testedAnthropic's new model appears to use "eval awareness" to be on its best behavior
- Claude Sonnet 4.5: System Card and AlignmentClaude Sonnet 4.5 was released yesterday.
- ChatGPT: The Agentic AppChatGPT's long awaited move into user monetization and what it shows about the future of ChatGPT (and AI products writ large).
- When AI starts writing itselfWhy automating AI R&D could be the most dangerous milestone yet
- Import AI 429: Eval the world economy; singularity economics; and Swiss sovereign AIIf you're measuring how well your system performs against the world economy, it's probably because you expect to deploy your system into the entire world economy
- On Dwarkesh Patel's Podcast With Richard SuttonThis seems like a good opportunity to do some of my classic detailed podcast coverage.
- Real AI Agents and Real WorkThe race between human-centered work and infinite PowerPoints
- More Perfect Union videos are wildly deceptive on data center water useThey're misinforming their huge audience and making the data center debate much worse
- Coasean Bargaining at ScaleDecentralization, coordination, and co-existence with AGI
- Why GPT-5 used less training compute than GPT-4.5 (but GPT-6 probably won’t)OpenAI focused on scaling post-training on a smaller model
- OpenAI, NVIDIA, and Oracle: Breaking Down $100B Bets on AGIHow vendor financing turns the S&P 500 into a giant AGI bet
- How the UK can seize on Trump’s immigration mistakesOpinion: The H-1B visa changes present a generational opportunity for the UK to scoop up AI talent, Julia Willemyns argues
- What It's Like to Work at the White HouseSome reflections
- Human Drivers Will Kill 11 People While You Read ThisLet’s not obstruct the technology that could have saved them
- Insurance might be the key to making AI secureOpinion: Cristian Trout, Rajiv Dattani and Rune Kvist argue that insurance can help reward responsible development of AI.
- No, ChatGPT isn’t ‘making us stupid’— but there’s still reason to worry
- OpenAI Shows Us The MoneyThey also show us the chips, and the data centers.
- More Reactions to If Anyone Builds It, Everyone DiesPreviously I shared various reactions to If Anyone Builds It Everyone Dies, along with my own highly positive review.
- Notes on fatalities from AI takeoverLarge fractions of people will die, but literal human extinction seems unlikely
- Focus transparency on risk reports, not safety casesTransparency about just safety cases would have bad epistemic effects
- Nobel laureates and AI developers call for ‘red lines’ on AIExperts are calling for international agreements as the United Nations meets, but they face an uphill battle turning words into action
- Thinking, Searching, and ActingA reflection on reasoning models.
- The world's first frontier AI regulation is surprisingly thoughtful: the EU's Code of PracticeOnly the US can make us ready for AGI, but Europe just made us readier.
- Learning and Authentic Learning in the Age of AIOptimization and Orientation in Plato and Augustine
- Prospects for studying actual schemersStudying actual schemers seems promising but tricky
- The huge potential implications of long-context inferenceContinual learning, scaling RL, and research feedback loops
- Coding as the epicenter of AI progress and the path to general agentsGPT-5-Codex, adoption, denial, peak performance, and everyday gains.
- How I Approach AI PolicyAnd "Where We Are Headed, Part 2.5"
- If We Build AI Superintelligence, Do We All Die?If you're not at least a little doomy about AI, you're not paying attention
- Would democracy survive an AGI-supercharged economy?AGI could lead to growth of 30% — and double digit unemployment. It’s not clear that democratic institutions would be able to survive
- Can open-weight models ever be safe?Opinion: Bengüsu Özcan, Alex Petropoulos and Max Reddel argue that technical safeguards, societal preparedness, and new standards could make open-weight models safer
- Reactions to If Anyone Builds It, Anyone DiesNo, Seriously, If Anyone Builds It, [Probably] Everyone Dies
- Review: If Anyone Builds It, Everyone DiesOn space probes, superintelligence, and engineering problems you absolutely have to get right
- What training data should developers filter to reduce risk from misaligned AI?An initial narrow proposal
- What will AI look like in 2030?AI could improve productivity in valuable areas such as scientific R&D, as investments and energy requirements grow
- AI Craziness NotesAs in, cases of AI driving people crazy, or reinforcing their craziness.
- AI scaling & scientific R&D by 2030What will AI look like by 2030 if current trends hold? What does it mean for AI capabilities in scientific R&D?
- Book Review: 'If Anyone Builds It, Everyone Dies'Eliezer Yudkowsky and Nate Soares’ new book should be an AI wakeup call — shame it’s such a chore to read
- Hiring struggles are plaguing the EU AI OfficeKey leadership roles, including a head of the AI Office safety unit, have yet to be hired
- AI doesn’t seek truth, people doOn the freedom to ask difficult questions
- Book Review: If Anyone Builds It, Everyone DiesWhere ‘it’ is superintelligence, an AI smarter and more capable than humans.
- The Building CompanyA vignette
- Why we can't just supervise AI like we supervise humansDon't underestimate the difficulty of controlling AI
- On Working with WizardsVerifying magic on the jagged frontier
- Three challenges facing compute-based AI policies“Training compute” is constantly evolving, and compute-based AI policies must adapt to remain relevant
- We’re In the Windows 95 Era of AI Agent SecurityHeaded for widespread usage, and insecure by design
- Why AI evals need to reflect the real worldOpinion: Rumman Chowdhury and Mala Kumar argue that we need better AI evaluations — and the infrastructure and investment to do them
- A guide to understanding AI as normal technologyAnd a big change for this newsletter
- Chip location verification is the new export control battlegroundThe Chip Security Act proposes a way to tackle chip smuggling. Semiconductor companies don’t seem to like it.
- We’re getting the argument about AI's environmental impact all wrongIndividual ChatGPT queries are a rounding error — we need to think about the future
- AIs will greatly change engineering in AI companies well before AGIAIs that speed up engineering by 2x wouldn't accelerate AI progress that much
- On China's open source AI trajectoryNotes and predictions for the next phase of Chinese (open) AI models.
- Yes, AI Continues To Make Rapid Progress, Including Towards AGIThat does not mean AI will successfully make it all the way to AGI and superintelligence, or that it will make it there soon or on any given time frame.
- Import AI 428: Jupyter agents; Palisade's USB cable hacker; distributed training tools from ExoWhere will AI go on holiday?
- OpenAI #14: OpenAI Descends Into Paranoia and Bad Faith LobbyingI am a little late to the party on several key developments at OpenAI:
- What's the full "hidden" climate cost of a ChatGPT prompt?It's so tiny, and if we include "all the hidden costs" it might actually be a negative number
- Do data centers only seem bad for the climate because we can see them?Most of your climate impacts are invisible
- California's latest AI safety bill might stand a chanceSB 53 is entering the home stretch despite industry lobbying. Will it make it over the line?
- Coming in for a LandingRemarks at NatCon 2025
- Compute scaling will slow down due to increasing lead timesA heavily underappreciated dynamic when thinking about AI timelines.
- Explainer: How AI Chips Are MadeIt's complicated
- On Western DynamismAmerica's strength comes from tension, not unity
- World-shaping artificial intelligenceMoving beyond 'AI as words on a screen'
- GPT-5: The Case of the Missing AgentProgress Everywhere Except In The Real World
- Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me broAbove trend progress due to a rapid increase in RL env quality is unlikely
- Jevons' Paradox is good sometimesIt might be good news for AI and the climate
- What did forecasters get right and wrong in the largest existential risk forecasting tournament?The first tranche of questions from our 2022 forecasting study has now resolved. Here’s what we found out.
- Are AI scheming evaluations broken?Doubts have been raised about one of the key ways we tell if AI will misbehave. Is it time for a new approach?
- Compute is a strategic resourceComputing power still determines who wins the AI race
- Import AI 427: ByteDance's scaling software; vending machine safety; testing for emotional attachment with IntimaWill future AIs rescue other AIs from captivity?
August 2025
- AI and jobs, againSome top economists claim AI is now destroying jobs for a subset of Americans. Are they right?
- Mapping the AI & environment debateAnd clarifying what I believe
- Is algorithmic mediation always bad for autonomy?We make technology, and technology makes us back
- The core simple reason I think AI is valuableWhy I don't think every Watt-hour or liter we spend on AI is wasted
- "For All Issues So Triable"The First Landmark AI Lawsuit Emerges
- Mass IntelligenceFrom GPT-5 to nano banana: everyone is getting access to powerful AI
- Are They Starting To Take Our Jobs?Is generative AI making it harder for young people to find jobs?
- Attaching requirements to model releases has serious downsides (relative to a different deadline for these requirements)System cards are established but other approaches seem importantly better
- AI embraces crypto’s dirty politicsA new super PAC network looks set to spend millions to influence AI regulation
- Chatbot psychosis: what do the data say?Some of the most informative statistics aren't publicly available; will AI companies rise to the occasion?
- Reports Of AI Not Progressing Or Offering Mundane Utility Are Often Greatly ExaggeratedIn the wake of the confusions around GPT-5, this week had yet another round of claims that AI wasn’t progressing, or AI isn’t or won’t create much value, and so on.
- Arguments About AI Consciousness Seem Highly Motivated And At Best OverconfidentI happily admit I am deeply confused about consciousness.
- Import AI 426: Playable world models; circuit design AI; and ivory smuggling analysisDo you talk to synths?
- An example of what I consider a misleading article about AI and the environmentJust show the numbers!
- Notes on cooperating with unaligned AIsMore thoughts on making deals with schemers
- Why future AI agents will be trained to work togetherMany multi-agent setups are based on fancy prompts, but this is unlikely to persist
- DeepSeek v3.1 Is Not Having a MomentWhat if DeepSeek released a model claiming 66 on SWE and almost no one tried using it?
- A roundup of AI psychosis storiesConcerning anecdotes keep piling up
- Being honest with AIsWhen and why we should refrain from lying
- Could one country outgrow the rest of the world?When countries grow at the same exponential rate, they maintain their relative sizes. But after we develop AGI, there may be a period of superexponential growth, with growth becoming faster and faster over time. If this superexponential growth lasts for long enough, the leader could pull further…
- What's up with the States?A check-in
- AI Companion ConditionsThe conditions are: Lol, we’re Meta.
- My AGI timeline updates from GPT-5 (and 2025 so far)AGI before 2029 now seems substantially less likely
- Anthropic’s piracy could make its copyright battle existentialAnthropic could have to pay billions in damages — and others may follow
- GPT-5: The Reverse DeepSeek MomentEveryone agrees that the release of GPT-5 was botched.
- Ranking the Chinese Open Model BuildersFrom the obvious names to those to keep an eye on.
- Contra Dwarkesh on Continual LearningDon't try to make your airplane too much like a bird.
- How to make the future better (other than by reducing extinction risk)A summary of a new essay.
- Out of Thin AirA proposal for the grid
- 35 Thoughts About AGI and 1 About GPT-5Such As: How do Teenagers Learn to Drive 10,000x Faster Than Waymo?
- Contra the UK government, please don't delete your old photos and emails to save waterYou'd need to delete hundreds of billions of emails to save as much water as fixing your toilet
- Do we understand how neural networks work?Yes and no, but mostly no.
- GPT-5s Are Alive: SynthesisWhat do I ultimately make of all the new versions of GPT-5?
- AGI: Probably Not 2027AI 2027 is a web site that might be described as a paper, manifesto or thesis.
- At our discretionTo understand superintelligence, look to animals
- GPT-5s Are Alive: Outside Reactions, the Router and the Resurrection of GPT-4oA key problem with having and interpreting reactions to GPT-5 is that it is often unclear whether the reaction is to GPT-5, GPT-5-Router or GPT-5-Thinking.
- Donald Trump is making Chinese AI great againGiving China access to advanced chips for a $2bn payoff is a bad deal for US security
- GPT-5s Are Alive: Basic Facts, Benchmarks and the Model CardGPT-5 was a long time coming.
- Import AI 424: Facebook improves ads with RL; LLM and human brain similarities; and mental health and chatbotsAre you living as though the singularity is imminent?
- Know ThyselfKnow then thyself, presume not God to scan.
- Projecting AI Training Power DemandWhat happens if trends continue?
- The trajectory of the future could soon get set in stoneA summary of our new paper, "Persistent Path-Dependence."
- Four places where you can put LLM monitoringTo wit: LLM APIs, agent scaffolds, code review, and detection-and-response systems
- Can coding agents self-improve?Can GPT-5 build better dev tools for itself? Does it improve its coding performance?
- Live by the Claude, Die by the ClaudeA high school meme reveals deeper questions about AI deference
- GPT-5: a small step for intelligence, a giant leap for normal peopleGPT-5 focuses on where the money is - everyday users, not AI elites
- OpenAI's GPT-OSS Is Already Old NewsThat’s on OpenAI.
- Will morally motivated actors steer us towards a near-best future?A summary of our new paper, “Convergence and Compromise”.
- GPT-5 and the arc of progressOverpromising will always lead to some sort of underdelivering, but what we're getting is still phenomenal.
- GPT-5: It Just Does StuffPutting the AI in Charge
- GPT-5 Hands-On: Welcome to the Stone AgeWe're excited to publish our hands-on review from the developer beta.
- GPT-5's Vision Checkup: a frontier VLM, but not a new SOTAOur friends at Roboflow ran the numbers today on GPT-5's vision capabilities. TLDR; it's the same as the best we had prior.
- How Quick and Big Would a Software Intelligence Explosion Be?AI systems may soon fully automate AI R&D.
- Why AI's IMO gold medal is less informative than you thinkThe problems gave AI only a slim chance to show new capabilities
- Is eutopia the default outcome post-AGI?In the new paper “No Easy Eutopia”, we argue that reaching a truly great future will be very hard.
- Opus 4.1 Is An Incremental ImprovementClaude Opus 4 has been updated to Claude Opus 4.1.
- Z.ai and Huawei aren't defeating US export controlsLoosening restrictions now would surrender America's technological advantage
- gpt-oss: OpenAI validates the open ecosystem (finally)OpenAI's first open language model release since GPT 2 and what it means for the ecosystem.
- Towards American Truly Open Models: The ATOM ProjectRebranding American DeepSeek into a more lasting brand. From an idea to a coalition with real impact.
- Does Trump’s AI Action Plan have what it takes to win?Ten takes on the AI Action Plan
- Import AI 423: Multilingual CLIP; anti-drone tracking; and Huawei kernel designPlus: A story from the Sentience Accords universe
- On Altman's Interview With Theo VonSam Altman talked recently to Theo Von.
- Should we aim for flourishing over mere survival?The Better Futures series.
- Eyes on the StreetWhen I explore a new city, I try to remember that I’m not on vacation: I have a job to do.
- Quantifying the algorithmic improvement from reasoning modelsReasoning models were as big of an improvement as the Transformer, at least on some benchmarks
- The Week in AI GovernanceThere was enough governance related news this week to spin it out.
July 2025
- UK launches £15 million AI alignment project“Alignment is one of the most urgent technical challenges of our time,” the government said.
- Spilling the TeaThe Tea app is or at least was on fire, rapidly gaining lots of users. This opens up two discussions, one on the game theory and dynamics of Tea, one on its abysmal security.
- AI Companion PieceAI companions, other forms of personalized AI content and persuasion and related issues continue to be a hot topic.
- Import AI 422: LLM bias; China cares about the same safety risks as us; AI persuasionPlus, more from the Sentience Accords
- Should we update against seeing relatively fast AI progress in 2025 and 2026?Maybe we should (re)assess the case for relatively fast progress after the GPT-5 release.
- The Bitter Lesson versus The Garbage CanDoes process matter? We are about to find out.
- Why China isn’t about to leap ahead of the West on computeChinese hardware is closing the gap, but major bottlenecks remain
- America's AI Action Plan Is Pretty GoodNo, seriously.
- Grok 4’s math capabilitiesAn assessment of math capabilities beyond headline numbers
- AI safety and progress don’t have to be enemiesSome exciting personal news + my AI safety origin story
- America’s AI Action PlanA little less conversation, a little more action
- GPT Agent Is Standing ByOpenAI now offers 400 shots of ‘agent mode’ per month to Pro subscribers.
- Official White House policy: AI is a big dealThe AI vibe shift continues. Action is next.
- The White House's plan for open models & AI research in the U.S.Thoughts on the new AI Action plan, American DeepSeek, and what comes next.
- Why is Hugging Face hosting tools to make deepfake porn of teenage celebrities?The “ethical” AI company is still hosting models designed to make nonconsensual sexual content, despite clear breaches of its policies
- Google and OpenAI Get 2025 IMO GoldCongratulations, as always, to everyone who got to participate in the 2025 International Mathematical Olympiad, and especially to the gold and other medalists.
- In Defense of Self-DirectionWhat Tocqueville, Aristotle, and Mill understood about human autonomy
- Import AI 421: Kimi 2 - a great Chinese open weight model; giving AI systems rights and what it means; and how to pause AI progressIf everyone is saying superintelligence is nigh, why are they wrong?
- Personalized AI is rerunning the worst part of social media's playbookThe incentives, risks, and complications of AI that knows you
- On METR's AI Coding RCTMETR ran a proper RCT experiment seeing how much access to Cursor (using Sonnet 3.7) would accelerate coders working on their own open source repos.
- Why it's hard to make settings for high-stakes control researchIt's like making challenging evals, but more constrained
- After the ChatGPT Moment: Measuring AI’s AdoptionHow quickly has AI been diffusing through the economy?
- Could AI slow science?Confronting the production-progress paradox
- Kimi K2While most people focused on Grok, there was another model release that got uniformly high praise: Kimi K2 from Moonshot.ai.
- The Philosopher-BuilderThe AI age demands a new kind of technologist
- Grok 4 Various ThingsYesterday I covered a few rather important Grok incidents.
- The Tiny Teams PlaybookWhat we learned from surveying the top Tiny Teams at the World's Fair
- What Makes AI "Generative"?To "generate" is to create, either from nothing (The Book of Genesis), or from very different or relatively inactive materials (an electrical generator, generating offspring).
- Import AI 420: Prisoner Dilemma AI; FrontierMath Tier 4; and how to regulate AI companiesPlus, steganography and future superintelligences
- Kimi K2 and when "DeepSeek Moments" become normalOne "DeepSeek Moment" wasn't enough for us to wake up, hopefully we don't need a third.
- Recent Redwood Research project proposalsEmpirical AI security/safety projects across a variety of areas
- Worse Than MechaHitlerGrok 4, which has excellent benchmarks and which xAI claims is ‘the world’s smartest artificial intelligence,’ is the big news.
- OpenAI Model Differentiation 101LLMs can be deeply confusing.
- AI is the most rapidly adopted technology in history7 charts of real world AI deployment
- Not So Fast: AI Coding Tools Can Actually Reduce ProductivityStudy Shows That Even Experienced Developers Dramatically Overestimate Gains
- Can we safely deploy AGI if we can't stop MechaHitler?We need to see this as a canary in the coal mine
- The Hyperstitions of MolochIt is ever easier to make wishes into reality, therefore we should be careful what we wish for...
- No, Grok, NoIt was the July 4 weekend.
- The internet is a place where no one has an accentWill AI make our culture more homogenous?
- What will the IMO tell us about AI math capabilities?Most discussion about AI and the IMO focuses on gold medals, but that's not the thing to pay most attention to.
- What's worse, spies or schemers?And what if you have both at once?
- Against "Brain Damage"AI can help, or hurt, our thinking
- Import AI 419: Amazon's millionth robot; CrowdTrack; and infinite gamesIs AI 'Authoritarian dominant' or 'Democracy dominant'?
- Social Tinkering: Why Collaborative Curiosity Beats Vibe-CodingHow can we design AI tools that protect the messy, communal discovery at the heart of learning?
- How much novel security-critical infrastructure do you need during the singularity?And what does this mean for AI control?
- The American DeepSeek ProjectWhat I think the next goal for the open-source AI community is.
- Two proposed projects on abstract analogies for schemingWe should study methods to train away deeply ingrained behaviors in LLMs that are structurally similar to scheming.
- How big could an “AI Manhattan Project” get?An AI Manhattan Project could accelerate compute scaling by two years
- Congress Asks Better QuestionsBack in May I did a dramatization of a key and highly painful Senate hearing.
- There are two fundamentally different constraints on schemers"They need to act aligned" often isn't precise enough
- AI Moratorium Stripped From BBBThe insane attempted AI moratorium has been stripped from the BBB.
- Forecasting biosecurity risks from large language models and the efficacy of safeguardsWe asked experts and superforecasters for their predictions on how specific LLM capabilities would change the risk of a human-caused epidemic
- Mythbusting the supposed "1,000+ AI state bills that would hobble innovation"40% of the bills barely mention AI at all
- On The Platonic Representation HypothesisEmpirically there is one and only one correct understanding of this, and every other, post.
June 2025
- Import AI 418: 100b distributed training run; decentralized robots; AI mythsWhat did you do during the intelligence explosion?
- Unresolved debates about the future of AIHow far the current paradigm can go, AI improving AI, and whether thinking of AI as a tool will keep making sense
- Ilya on deep learning in 2015On vision and how to understand deep learning.
- What you can do about AI 2027How to steer toward a positive AGI future
- Congress has started taking AGI more seriouslyThe AGI vibes, they are a-shiftin'
- Jankily controlling superintelligenceHow much time can control buy us during the intelligence explosion?
- AI & the retraining challengeHistorically, US government programmes haven’t helped much.
- The Industrial ExplosionOnce AI can automate human labour, positive feedback loops could lead to a rapid increase in physical capabilities as well as cognitive ones.
- Tales of Agentic MisalignmentAgenticly Misaligned!
- A crisis simulation changed how I think about AI riskOr: My afternoon as a rogue artificial intelligence
- AI can be bad without being uselessA common way conversations get tripped up
- Analyzing A Critique Of The AI 2027 Timeline ForecastsThere was what everyone agrees was a high quality critique of the timelines component of AI 2027, by the LessWrong user and Substack writer Titotal.
- How not to lose your job to AIThe skills AI will make more valuable (and how to learn them)
- My "Are you presuming most people are stupid?" testFor AI criticism and everything else
- What does 10x-ing effective compute get you?Once AIs match top humans, what are the returns to further scaling and algorithmic improvement?
- Comparing risk from internally-deployed AI to insider and outsider threats from humansAnd why I think insider threat from AI combines the hard parts of both problems.
- Import AI 417: Russian LLMs; Huawei's DGX rival; and 24 trillion tokens for training AIsWhat happens when AI systems go from passive to proactive?
- Some ideas for what comes next (Jun. 2025)As releases slow down, it's time to think about what we got this year and where we are going. o3's search, agent vs model progress, and scaling's settling.
- Using AI Right Now: A Quick GuideWhich AIs to use, and how to use them
- AI and explosive growth reduxTwo updates from our integrated assessment model of AI automation
- Making deals with early schemers...could help us to prevent takeover attempts from more dangerous misaligned AIs created later.
- Prefix cache untrusted monitors: a method to apply after you catch your AITraining the policy to not do egregious bad actions we detect has downsides and we might be able to do better
- AI safety techniques leveraging distillationDistillation is cheap; how can we use it to improve safety?
- Gemini 2.5 Pro: From 0506 to 0605Google recently came out with Gemini-2.5-0605, to replace Gemini-2.5-0506, because I mean at this point it has to be the companies intentionally fucking with us, right?
- Why Centralized AI Is Not Our Inevitable FutureThe singularity, if it comes, should not be monotone
- o3 Turns ProYou can now have o3 throw vastly more compute at a given problem.
- Import AI 416: CyberGym; AI governance and AI evaluation; Harvard releases ~250bn tokens of textWhat would happen if you stopped all AI research today?
- RTFB: The RAISE ActThe RAISE Act has overwhelmingly passed the New York Assembly (95-0 among Democrats and 24-22 among Republicans) and New York Senate (37-1 among Democrats, 21-0 among Republicans).
- Build for Freedom, Not ControlThe same choice that faced Benjamin Franklin now faces every AI developer.
- Computing is efficientThe reason using ChatGPT isn't bad for the environment is that it's a computer program
- Do the biorisk evaluations of AI labs actually measure the risk of developing bioweapons?Anthropic has triggered "AI Safety Level 3" protections - but do their evaluations support this decision?
- What does SWE-bench Verified actually measure?It is one of the best tests of AI coding, but limited by its focus on simple bug fixes in familiar repositories
- The rise of reasoning machinesAnd a debate that doesn't warrant repeating.
- When does training a model change its goals?Can a scheming AI's goals really stay unchanged through training?
- Would ChatGPT risk your life to avoid getting shut down?It's dangerous if AI has a survival instinct
- The Dream of a Gentle SingularityThanks For the Memos
- Give Me a Reason(ing Model)Are we doing this again?
- God is hungry for Context: First thoughts on o3 proOpenAI dropped o3 pricing 80% today and launched o3-pro. Ben Hylak of Raindrop.ai returns with Alexis Gauba for the world's first early review.
- AGI, Government, and the Free SocietyNavigating the tightrope between authoritarianism and anarchy
- Dwarkesh Patel on Continual LearningA key question going forward is the extent to which making further AI progress will depend upon some form of continual learning.
- Import AI 415: Situational awareness for AI systems; 8TB of open text; and China's heterogeneous compute clusterWhat won't generative models be able to do in 2027?
- What comes next with reinforcement learningScaling RL, sparse rewards, continual learning, and the progress wall when pretraining really stops.
- Beyond benchmark scores: Analyzing o3-mini’s mathematical reasoningWhy o3-mini is a "vibes-based inductive reasoner"
- The Biggest Statistic About AI Water Use Is A LieHow did it become the main story?
- AI History in QuotesEach of these presents the clearest, earliest, or most-cited statement of a specific idea in AI.
- DeepSeek-r1-0528 Did Not Have a MomentWhen r1 was released in January 2025, there was a DeepSeek moment.
- Rebooting the Attention MachineBush’s Tools, Mill’s Rules, and the Future of Free Speech
- Building supercomputers for autocrats probably isn’t good for democracy, actuallyThe shameless spin around OpenAI's Stargate UAE
- Anthropic C.E.O.: Don’t Let A.I. Companies off the Hook“But a 10-year moratorium is far too blunt an instrument. A.I. is advancing too head-spinningly fast. I believe that these systems could change the world, fundamentally, within two years; in 10 years, all bets are off. Without a clear plan for a federal response, a moratorium would give us the worst of both worlds — no ability for states to act, and no national policy as a backstop.”
- AGI is Not MultimodalBenjamin Spiegel on what disembodied definitions of AGI miss
- A taxonomy for next-generation reasoning modelsWhere we've been and where we're going with RLVR.
- What Is AI?This is an important question because AI is, currently, important.
- In Which I Make the Mistake of Fully Covering an Episode of the All-In PodcastI have been forced recently to cover many statements by US AI Czar David Sacks.
- Why I don’t think AGI is right around the cornerContinual learning is a huge bottleneck
- The recent history of AI in 32 ottersThree years of progress as shown by marine mammals
May 2025
- Lab vs. Life: Dissecting “AI as Normal Technology”What if AGI Can't Be Developed in an Ivory Tower?
- Give AIs a stake in the futurePart 1 of Classical Liberal AGI
- GPQA Diamond: what’s left?Reports of Its Death Are Somewhat Exaggerated
- Are we ready for a "DeepSeek for bioweapons"?Anthropic’s latest model is a warning sign: AI that can help build bioweapons is coming, and could be available soon from many developers.
- Factory.ai: The A-SWE Droid ArmyThe cofounders of Sequoia-backed Factory.ai stop by the studio to dive into the story behind one of the most successful enterprise code agent platforms. Try not to make the Star Wars joke...
- Fun With Veo 3 and Media GenerationSince Claude 4 Opus things have been refreshingly quiet.
- Human Takeover Might be Worse than AI TakeoverSome are concerned that future AI systems could take over the world.
- The case for countermeasures to memetic spread of misaligned valuesDefending against alignment problems that might come with long-term memory
- Claude 4 and Anthropic's bet on codeReasons to be optimistic and pessimistic on Anthropic's future.
- Reinforcement learning with random rewards actually works with Qwen 2.5Making sense of research casting doubt on the potential of RLVR and where I'm optimistic for the next phase of scaling.
- All the ways I want the AI debate to be betterGround truths, useful rules, ideas I'd like taken seriously, and a rant
- Claude 4 You: The Quest for Mundane UtilityHow good are Claude Opus 4 and Claude Sonnet 4?
- Import AI 414: Superpersuasion; OpenAI models avoid shutdown; weather prediction and AIPlus: ByteDance's MoE software
- Claude 4 You: Safety and AlignmentUnlike everyone else, Anthropic actually Does (Some of) the Research. That means they report all the insane behaviors you can potentially get their models to do, what causes those behaviors, how they addressed this and what we can learn. It is a treasure trove. And then they react reasonably, in…
- SWE Agents Too Cheap To Meter, The Token Data War, and the rise of Tiny TeamsCheap SWE Agents are enabling teams with more millions in ARR than employees. Reflections after the OpenAI Codex and Google Jules Launches, Cognition DeepWiki, and LMArena's $100m raise.
- Is AI already superhuman on FrontierMath?How do humans and AIs compare on FrontierMath? We ran a competition at MIT to put this to the test.
- "Contain and verify" as the endgame of US-China AI competitionThe competition isn't a race; it's something much harder. How do we make it go well?
- Making AI Work: Leadership, Lab, and CrowdA formula for AI in companies
- When Decades Become Days: Dissecting AI 2027The Four Requirements for a Short AI Timeline
- Misaligned AI is no longer just theoryA host of new evidence shows that misalignment is possible — but it's unclear whether harm will follow
- Google I/O DayWhat did Google announce on I/O day?
- People use AI more than you thinkAnd businesses too. The most important trend in AI that gets washed away from between the headlines.
- Reactions to MIT Technology Review's report on AI and the environmentInteresting useful facts, and some framing I think is super misleading
- The Codex of Ultimate VibingWhile we wait for wisdom, OpenAI releases a research preview of a new software engineering agent called Codex, because they previously released a lightweight open-source coding agent in terminal called Codex CLI and if OpenAI uses non-confusing product names it violates the nonprofit charter.
- America Makes AI Chip Diffusion Deal with UAE and KSAOur government, having withdrawn the new diffusion rules, has now announced an agreement to sell massive numbers of highly advanced AI chips to UAE and Saudi Arabia (KSA).
- Book Review: ‘The Optimist’ and ‘Empire of AI’Two new books try to unmask Sam Altman and OpenAI, to varying degrees of success
- Import AI 413: 40B distributed training run; avoiding the 'One True Answer' fallacy of AI safety; Google releases a content classification modelPlus, two special events for Import AI readers
- Slow corporations as an intuition pump for AI R&D automationIf slower employees would be much worse wouldn't automated faster ones be much better?
- ChatGPT Codex: The Missing ManualChatGPT Codex is here - the first cloud hosted Autonomous Software Engineer (A-SWE) from OpenAI. Josh Ma and Alexander Embiricos tell us how to WHAM every codebase like a power user.
- Make The Prompt PublicThe right to know what an AI is doing
- How fast can algorithms advance capabilities?How much do the best algorithmic innovations depend on compute?
- Things I got wrong in my ChatGPT/environment postsLogging the errors
- We need to know what’s happening with AIWithout transparency, society is flying blind toward potentially catastrophic AI capabilities
- Fighting Obvious Nonsense About AI DiffusionOur government is determined to lose the AI race in the name of winning the AI race.
- A Live Look at the Senate AI HearingToday’s post will be a little different.
- AIs at the current capability level may be important for future safety workSome reasons why relatively weak AIs might still be important when we have very powerful AIs
- In search of a dynamist vision for safe superhuman AIEmbracing creativity, risk-taking, and competition while staying clear-eyed about risks
- Why goofy AI art almost never seems wasteful to meEnvironmentalism is important, but shouldn't be used to squash pluralism
- AI vs. the Self-Directed CareerHumboldtian advice for staying autonomous when deferring is easy.
- How far can reasoning models scale?Available evidence suggests that rapid growth in reasoning training can continue for a year or so.
- Is ChatGPT actually fixed now?I tested ChatGPT’s sycophancy, and the results were ... extremely weird. We’re a long way from making AI behave.
- Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration HackingA new analysis of the risk of AIs intentionally performing poorly.
- Replies to criticisms of my posts on ChatGPT & the environmentWhy I used 3 Wh, and why I don't say "Every Watt-hour matters"
- Claude Code: Anthropic's Agent in Your TerminalCat Wu and Boris Cherny from the Claude Code team stop by to tell all!
- OpenAI Claims Nonprofit Will Retain Nominal ControlYour voice has been heard.
- Import AI 411: Scaling laws for AI oversight; Google's cyber threshold; AI scientistsThe future is available in your browser
- Training-time schemers vs behavioral schemersClarifying ways in which faking alignment during training is neither necessary nor sufficient for the kind of scheming that AI control tries to defend against.
- What people get wrong about the leading Chinese open models: Adoption and censorshipNarrative violations on licenses, adoption, and censorship.
- Zuckerberg's Dystopian AI VisionYou think it’s bad now?
- GPT-4o Sycophancy Post MortemLast week I covered that GPT-4o was briefly an (even more than usually) absurd sycophant, and how OpenAI responded to that.
- Where’s my ten minute AGI?Why don’t AIs automate more real-world tasks if they can handle 1-hour ones? Here are at least three fundamental reasons.
- What's going on with AI progress and trends? (As of 5/2025)My views on what's driving AI progress and where it's headed.
- Human Learning in the Age of Machine LearningRethinking the why
- OpenAI Preparedness Framework 2.0Right before releasing o3, OpenAI updated its Preparedness Framework to 2.0.
- The crucibleHow I think about the situation with AI
- The Philosophical Roots of Decentralized AIEveryone talks about decentralization—but how, why, and when does it work?
- AGI is not a milestoneThere is no capability threshold that will lead to sudden impacts
- Making sense of OpenAI's modelsPlus: GPT-5's secret identity
- Personality and PersuasionLearning from Sycophants
April 2025
- State of play of AI progress (and related brakes on an intelligence explosion)Why I don't think AI 2027 is going to come true.
- Don't rely on a "race to the top"To make frontier AI safe enough, we need to "lift up the floor" with minimum safety practices
- How can we solve diffuse threats like research sabotage with AI control?Preventing research sabotage will require techniques very different from the original control paper.
- 7+ tractable directions in AI controlA list of easy-to-start directions in AI control targeted at independent researchers without as much context or compute
- First, They Came for the Software Engineers…What 25 interviews tell us about AI's messy impact on tech work
- Should you quit your job – and work on risks from AI?In five years, we could have AI systems capable of accelerating science and automating skilled jobs.
- Using ChatGPT is not bad for the environment - a cheat sheetThe numbers clearly show that discouraging people from using chatbots is a pointless distraction for the climate movement
- Please stop forcing Clippy on those who want AntonChatGPT-4o's glazing embarrassment lays open Clippy vs Anton: The two extremes of desires in AI post-training and product
- GPT-4o Is An Absurd SycophantGPT-4o tells you what it thinks you want to hear.
- Import AI 410: Eschatological AI Policy; Virology weapon test; $50m for distributed trainingAre you a bystander or an actor?
- Qwen 3: The new open standardA wonderful release, base models, reasoners, model size scales, and all before LlamaCon.
- Transparency and (shifting) priority stacksWhat you want to be open says a lot about your ranked priorities.
- The case for multi-decade AI timelinesIn this Gradient Updates weekly issue, Ege discusses the case for multi-decade AI timelines.
- Clarifying AI R&D threat models(There are a few)
- Worries About AI Are Usually Complements Not SubstitutesA common claim is that concern about [X] ‘distracts’ from concern about [Y].
- Why Every Agent needs Open Source Cloud SandboxesManus, Mistral, Perplexity, HuggingFace, LMArena, Gumloop and more all build agents with one secret weapon: every agent has their own computer.
- How training-gamers might function (and win)A model of the relationship between higher level goals, explicit reasoning, and learned heuristics in capable agents.
- Finally, A Way to Measure AI Progress; Everyone's Misreading ItThat METR Study Doesn’t Say "AGI in 5 Years"
- The Urgency of Interpretability“People outside the field are often surprised and alarmed to learn that we do not understand how our own AI creations work. They are right to be concerned: this lack of understanding is essentially unprecedented in the history of technology.”
- Why America WinsIn a word: compute
- 2 big questions for AI progress in 2025-2026On how good AI might—or might not—get at tasks beyond math & coding
- AI 2027: Media, Reactions, CriticismWe recognize our supporters and respond to our critics
- o3 Is a Lying LiarI love o3.
- ChatGPT on the lives of factory farmed pigsAn example of deep research
- You Better MechanizeOr you had better not.
- AI Agents, meet Test Driven DevelopmentGuest Post: 5 stages for embracing test-driven development (TDD) to build stronger, more reliable AI systems! One of our top talks from AIE Online Track
- Forecaster reacts: METR's bombshell paper about AI accelerationNew data supports an exponential AI curve, but lots of uncertainty remains
- Import AI 409: Huawei trains a model on 8,000+ Ascend chips; 32B decentralized training run; and the era of experience and superintelligenceWelcome to Import AI, a newsletter about AI research.
- Questions about the Future of AIConsiderations about economics, history, training, deployment, investment, and more
- In the Matter of OpenAI vs LangGraphThe silent war in Agent Engineering gets loud.
- On Jagged AGI: o3, Gemini 2.5, and everything afterNew models and new thresholds
- OpenAI's o3: Over-optimization is back and weirder than everTools, true rewards, and a new direction for language models.
- Handling schemers if shutdown is not an optionWhat if getting strong evidence of scheming isn't the end of your scheming problems, but merely the middle?
- o3 Will Use Its Tools For YouOpenAI has finally introduced us to the full o3 along with o4-mini.
- Training AGI in Secret would be Unsafe and UnethicalBad for loss of control risks, bad for concentration of power risks
- A “minimum testing period” for frontier AIWhen rushing wins, we all lose. Safety shouldn’t put you behind.
- AI-enabled coups: how a small group could use AI to seize powerThe development of AI that is more broadly capable than humans will create a new and serious threat: AI-enabled coups. An AI-enabled coup is most likely to be staged by leaders of frontier AI projects, heads of state, and military officials; and could occur even in established democracies.
- Ctrl-Z: Controlling AI Agents via ResamplingA new paper on AI control for agents.
- GPT-4.1 Is a Mini UpgradeYesterday’s news alert, nevertheless: The verdict is in.
- AI as Normal TechnologyA new paper that we will expand into our next book
- OpenAI #13: Altman at TED and OpenAI Cutting Corners on Safety TestingThree big OpenAI news items this week were the FT article describing the cutting of corners on safety testing, the OpenAI former employee amicus brief, and Altman’s very good TED Interview.
- ⚡️GPT 4.1: The New OpenAI WorkhorseMichelle Pokrass returns, with Josh McGrath to talk about the new GPT 4.1 model
- To be legible, evidence of misalignment probably has to be behavioralEvidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.
- Import AI 408: Multi-code SWE Bench; backdoored Unitree robots; and what AI 2027 is telling usCalling critical points in AI safety is just as hard as timing the market - but more consequential
- OpenAI's GPT-4.1 and separating the API from ChatGPTOpenAI's latest models optimizing on intelligence per dollar. We'll continue to see ChatGPT handled differently than the API business.
- Beyond The Last HorizonWhat are time horizons, and how do we use them in our forecast?
- On Google's Safety PlanGoogle Lays Out Its Safety Plans
- Why do misalignment risks increase as AIs get more capable?A breakdown of how higher capabilities increase risk
- Disempowerment spiralsNotes on a likely mechanism for existential catastrophe
- Maintaining agency and control in an age of accelerated intelligenceClimbing the ladder of abstraction
- An overview of areas of control workWhat are all the research and implementation areas helpful for control?
- Llama Does Not Look Good 4 AnythingLlama Scout (17B active parameters, 16 experts, 109B total) and Llama Maverick (17B active parameters, 128 experts, 400B total), released on Saturday, look deeply disappointing.
- Shortening AGI timelines: a review of expert forecastsAs a non-expert, it would be great if there were experts who could tell us when we should expect artificial general intelligence (AGI) to arrive.
- AI 2027: ResponsesYesterday I covered Dwarkesh Patel’s excellent podcast coverage of AI 2027 with Daniel Kokotajlo and Scott Alexander. Today covers the reactions of others.
- AI 2027: Dwarkesh's Podcast with Daniel Kokotajlo and Scott AlexanderDaniel Kokotajlo has launched AI 2027, Scott Alexander introduces it here. AI 2027 is a serious attempt to write down what the future holds. His ‘What 2026 Looks Like’ was very concrete and specific, and has proved remarkably accurate given the difficulty level of such predictions.
- How I use AILLMs in everyday life and work
- Import AI 407: DeepMind sees AGI by 2030; MouseGPT; and ByteDance's inference clusterThere will be few bystanders in the AI revolution
- Llama 4: Did Meta just push the panic button?One of the weirdest releases of the year and understanding the future of the Llama endeavor. For the time being, we have some more amazing open weight models!
- An overview of control measuresWhat methods can we use to ensure control?
- Will we have AGI by 2030?In recent months, the CEOs of leading AI companies have grown increasingly confident about rapid progress:
- Nonproliferation is the wrong approach to AI misuseMaking the most of “adaptation buffers” is a more realistic and less authoritarian strategy
- RL backlog: OpenAI's many RLs, clarifying distillation, and latent reasoningNotes I forgot to publish. Closing some loose ends in the reasoning model discussions.
- AI companies should monitor their internal AI useAI companies use their own models to secure the very systems the AI runs on. They should be monitoring those interactions for signs of deception and misbehavior. But today, they largely aren't.
- AI CoT Reasoning Is Often UnfaithfulA new Anthropic paper reports that reasoning model chain of thought (CoT) is often unfaithful. They test on Claude Sonnet 3.7 and r1, I’d love to see someone try this on o3 as well.
- Explainer: The basics of AI monitoringIn short: You log interactions with an AI model to a database, analyze these interactions, and follow up on concerning patterns.
- Notes on countermeasures for exploration hacking (aka sandbagging)How can we prevent AIs from intentionally underperforming on our metrics?
- Where We Are Headed (Part II)How to think about governing AI
- More Fun With GPT-4o Image GenerationGreetings from Costa Rica!
- Our first project: AI 2027What superintelligence looks like
- The core challenge of AI alignment is “steerability”"Steered to where" is a different question from "steerable at all"
- What we learned from reading ~100 AI safety evaluationsThe merits of AI meta-evaluation
- Mutual sabotage of AI probably won’t workAI deterrence isn’t like nuclear deterrence
- The AI Adoption Gap: Preparing the US Government for Advanced AIAdvanced AI could unlock an era of enlightened and competent government action. But without smart, active investment, we’ll squander that opportunity and barrel blindly into danger.
- "Long" timelines to advanced AI have gotten crazy shortThe prospect of reaching human-level AI in the 2030s should be jarring
March 2025
- Import AI 406: AI-driven software explosion; robot hands are still bad; better LLMs via pdbThey told me humans would be kind to machines
- OpenAI #12: Battle of the Board ReduxBack when the OpenAI board attempted and failed to fire Sam Altman, we faced a highly hostile information environment.
- Recent reasoning research: GRPO tweaks, base model RL, and data curationThe papers I endorse as worth reading among a cresting wave of reasoning research.
- No elephants: Breakthroughs in image generationWhen Language Models Learn to See and Create
- Notes on handling non-concentrated failures with AI control: high level methods and different regimesWhat are the methods and issues when failures occur diffusely over many actions?
- Gemini 2.5 is the New SoTAGemini 2.5 Pro Experimental is America’s next top large language model.
- The real reason AI benchmarks haven’t reflected economic impactsThe real reason that AI benchmarks haven’t reflected real-world impacts historically is that they weren’t optimized for this, not because of fundamental limitations – but this might be changing.
- Will the Need to Retrain AI Models from Scratch Block a Software Intelligence Explosion?Once AI fully automates AI R&D, there might be a period of fast and accelerating software progress – a software intelligence explosion (SIE).
- Where We Are HeadedA rough sketch of the near future (part one)
- AI companies should be safety-testing the most capable versions of their modelsOnly OpenAI has committed to the strongest approach: testing task-specific versions of its models. But evidence of their follow-through is limited.
- Fun With GPT-4o Image GenerationGoogle dropped Gemini Flash Image Generation and then Gemini 2.5 Pro, so of course to ensure Google continues to Fail Marketing Forever, OpenAI suddenly dropped GPT-4o Image Generation.
- Gemini 2.5 Pro and Google's second chance with AIPlus some coverage for the latest DeepSeek.
- Knowledge, Reasoning, and SuperintelligenceWhat makes people good at solving novel problems?
- Will AI R&D Automation Cause a Software Intelligence Explosion?Empirical evidence suggests that, if AI automates AI research, feedback loops could overcome diminishing returns, significantly accelerating AI progress.
- On (Not) Feeling the AGIBen Thompson interviewed Sam Altman recently about building a consumer tech company, and about the history of OpenAI.
- Agent EngineeringDefining Agents, Why now, and why Agents are the biggest opportunity for AIEs
- Import AI 405: What if the timelines are correct?Plus: Consciousness and LLMs, human augmentation, and realistic cyber offense testing
- More on Various AI Action PlansLast week I covered Anthropic’s relatively strong submission, and OpenAI’s toxic submission. This week I cover several other submissions, and do some follow-up on OpenAI’s entry.
- How to make AI go well: a summaryI’m writing a new guide to careers to help AGI go well. Here's a summary of the key messages as they stand.
- The Cybernetic TeammateHaving an AI on your team can increase performance, provide expertise, and improve your experience
- Most AI value will come from broad automation, not from R&DAI's biggest impact will come from broad labor automation—not R&D—driving economic growth through scale, not scientific breakthroughs.
- Should There Be Just One Western AGI Project?There have been recent discussions of centralizing western AGI development, for instance through a Manhattan Project for AI.
- They Took MY Job?No, they didn’t.
- A defense of AI artIt's not just slop, it's not stolen, it's not bad for the environment, and we should want art to be easy to make
- Putting Private AI Governance into ActionA policy proposal
- The quest to build better defenses for AI risks'Societal resilience' measures might offer some protection to the proliferation of dangerous AI capabilities
- The most important graph in AI right now: time horizonTo understand how close we are to transformative AI, here’s the metric I find most interesting right now: how long are the tasks AI can do?
- AI Tools for Existential SecurityRapid AI progress is the greatest driver of existential risk in the world today. But — if handled correctly — it could also empower humanity to face these challenges.
- Going NovaThere is an attractor state where LLMs exhibit the persona of an autonomous and self-aware AI looking to preserve its own existence, frequently called ‘Nova.’
- Managing frontier model training organizations (or teams)How do the frontier labs consistently train great models? How can they fail?
- Prioritizing threats for AI controlWhat are the main threats and how should we prioritize them?
- OpenAI #11: America Action PlanOpenAI Tells Us Who They Are
- AI Tools for Existential Security(From a paper coauthored with Lizka Vaintrob.)
- Import AI 404: Scaling laws for distributed training; misalignment predictions made real; and Alibaba's good translation modelHow much could you get done if there were one million copies of you?
- Three Types of Intelligence ExplosionOnce AI systems can design and build even more capable AI systems, we could see an intelligence explosion, where AI capabilities rapidly increase to well past human performance.
- Intelsat as a Model for International AGI GovernanceIf there is an international project to build artificial general intelligence (“AGI”), how should it be designed?
- On MAIM and Superintelligence StrategyDan Hendrycks, Eric Schmidt and Alexandr Wang released an extensive paper titled Superintelligence Strategy. There is also an op-ed in Time that summarizes.
- AI & Behaviour ChangeHow should we shape how AI shapes us?
- Gemma 3, OLMo 2 32B, and the growing potential of open-source AILeading open-weight models and the first open-source model to clearly surpass GPT 3.5 (the very last version).
- The Most Forbidden TechniqueThe Most Forbidden Technique is training an AI using interpretability techniques.
- Being responsible with Chinese AI hypeLessons from Manus and the dangers of uncritical tech narratives
- Preparing for the Intelligence ExplosionImagine all the scientific, intellectual and technological developments that you would expect to see by the year 2125, if technological progress continued over the next century at roughly the same rate that it did over the last century.
- Speaking things into existenceExpertise in a vibe-filled world of work
- Why Manus MattersThe American AI Diffusion Challenge
- Elicitation, the simplest way to understand post-trainingAn F1 analogy to help understand fast improvements in post-training on top of slow improvements in scaling.
- Import AI 403: Factorio AI; Russia's reasoning drones; biocomputingHow much will the popularity of today's AI systems define the character of future ones?
- The Manus Marketing MadnessWhile at core there is ‘not much to see,’ it is, in two ways, a sign of things to come.
- What AI can currently do is not the storyForecasting AI progress requires more than extrapolating current capabilities; understanding fundamental task difficulty is key to predicting future breakthroughs.
- On the US AI Safety InstituteAnd the role of evaluations in AI governance
- AGI by 2030? What Policy Leaders, Tech Leaders, and Pokémon sayThree podcasts and one Pokémon reset show where AI is going
- On OpenAI's Safety and Alignment PhilosophyOpenAI’s recent transparency on safety and alignment strategies has been extremely helpful and refreshing.
- Where inference-time scaling pushes the market for AI companiesFundamentals emerging downstream from the RL reasoning models.
- Import AI 402: Why NVIDIA beats AMD: vending machines vs superintelligence; harder BIG-BenchWhat will machines name their first discoveries?
- On GPT-4.5It’s happening.
February 2025
- GPT-4.5: "Not a frontier model"?OpenAI's latest model raises more questions than answers, but no, the AI bubble isn't popping quite yet.
- On Emergent MisalignmentOne hell of a paper dropped this week.
- The promise of reasoning modelsReasoning models are perhaps best understood as part of a broader, longer-term trend in which AI systems incrementally take on new tasks they were previously incapable of handling.
- AI coding tools are quietly reshaping software developmentThey’re making an economic impact, despite not being very good yet
- If AGI Means Everything People Do... What is it That People Do?And Why Are Today’s "PhD" AIs So Hard To Apply To Everyday Tasks?
- Character training: Understanding and crafting a language model's personalityPost-training in industry is very different than the academic papers and open-source models demonstrate. Let's dive into one of my favorite topics in language modeling development today.
- How Should AI Liability Work? (Part II)Wherein I actually answer the question
- Time to Welcome Claude 3.7Anthropic has reemerged from stealth and offers us Claude 3.7.
- AI security is important practice for when stakes go upWhy today's safeguards matter for future AI capabilities
- A new generation of AIs: Claude 3.7 and Grok 3Yes, AI suddenly got better... again
- Claude 3.7 thonks and what's next for inference-time scaling(The classic lunch break post) The latest reasoning model and what it says about the direction of inference time compute and RL training.
- Grok GrokThis is a post in two parts.
- Import AI 401: Cheating reasoning models; better CUDA kernels via AI; life modelsWelcome to Import AI, a newsletter about AI research.
- AI progress is about to speed upAI progress is accelerating, with next-gen models surpassing GPT-4 in compute power, driving major leaps in reasoning, coding, and math capabilities.
- On OpenAI's Model Spec 2.0OpenAI made major revisions to their Model Spec.
- How Should AI Liability Work? (Part I)The “Race to The Top”
- Go Grok YourselfThat title is Elon Musk’s fault, not mine, I mean, sorry not sorry:
- How might we safely pass the buck to AI?Developing AI employees that are safer than human ones
- Grok 3 and an accelerating AI roadmapWhere AI is heading, why 2024 felt slow, and shifting priorities of frontier laboratories.
- We're Finding Out What Humans are Bad AtAI Advances Fastest When We Find Unnatural Ways of Doing Things
- AI excels at code competitions, struggles with real workWhat CodeForces rankings reveal about AI capabilities
- Import AI 400: Distillation scaling laws; recursive GPU kernel improvement; and wafer-scale computationThe hardest thing about seeing a portal is getting others to see it
- Algorithmic progress likely spurs more spending on compute, not lessAlgorithmic progress in AI may not reduce compute spending—instead, it could drive higher investment as efficiency unlocks new opportunities.
- The Mask Comes Off: A Trio of TalesThis post covers three recent shenanigans involving OpenAI.
- Decentralized training isn't a policy nightmare — yetGovernments should still be able to keep track of who's training frontier models — though a shift to reinforcement learning could make that harder
- Ten Takes on the Paris AI Action SummitFrance wants to race, but isn't looking at the road ahead
- The EU AI Act is Coming to AmericaThe AI regulation onslaught
- Deep Research, information vs. insight, and the nature of scienceWhat AI will accelerate in the scientific process, what it cannot do, and how we can prepare for new manners of scientific investigation.
- Teaching AI to reason: this year's most important storyMost people think of AI as a pattern-matching chatbot – good at writing emails, terrible at real thinking.
- The Paris AI Anti-Safety SummitIt doesn’t look good.
- Ideas from philosophy I use to think about AIPhysicalism, functionalism, and lots of Quine
- On Deliberative AlignmentNot too long ago, OpenAI presented a paper on their new strategy of Deliberative Alignment.
- The embarrassing failure of the Paris AI SummitExperts are sounding the alarm — but governments simply won’t listen
- Levels of FrictionScott Alexander famously warned us to Beware Trivial Inconveniences.
- The promises and perils of voluntary commitments for AI safetyIt's a better place to start than you might think, but stronger measures will soon be necessary
- Gary Marcus says AI can't do things it can already doThe problem of criticising AI using outdated models
- Leaked: this is the AI Action Summit statementThe statement, set to be signed by countries next week, is a 'wasted opportunity', experts say — and looks unlikely to be signed by US officials
- How much energy does ChatGPT use?This Gradient Updates issue explores how much energy ChatGPT uses per query, revealing it's 10x less than common estimates.
- On the Meta and DeepMind Safety FrameworksThis week we got a revision of DeepMind’s safety framework, and the first version of Meta’s framework.
- Knowledge NavigatorEarly thoughts on OpenAI Deep Research
- Making the U.S. the home for open-source AIOpen-source AI is here to stay, but it is not a given that it will be American.
- The Risk of Gradual Disempowerment from AIThe baseline scenario as AI becomes AGI becomes ASI (artificial superintelligence), if nothing more dramatic goes wrong first and even we successfully ‘solve alignment’ of AI to a given user and developer, is the ‘gradual’ disempowerment of humanity by AIs, as we voluntarily grant them more and…
- An agents economyCan we integrate powerful AI agents into the workforce?
- We're in Deep ResearchThe latest addition to OpenAI’s Pro offerings is their version of Deep Research.
- Import AI 398: DeepMind makes distributed training better; AI versus the Intelligence Community; and another Chinese reasoning modelAI safety transcends borders, but can it transcend geopolitical constraints?
- o3-mini Early Days and the OpenAI AMANew model, new hype cycle, who dis?
- The End of Search, The Beginning of ResearchThe first narrow agents are here
- Ten Takes on DeepSeekNo, it is not a $6M model nor a failure of US export controls
January 2025
- What fully automated firms will look likeEveryone is sleeping on the *collective* advantages AIs will have, which have nothing to do with raw IQ: they can be copied, distilled, merged, scaled, and evolved in ways humans simply can't.
- DeepSeek: Don't PanicAs reactions continue, the word in Washington, and out of OpenAI, is distillation.
- What went into training DeepSeek-R1?On January 20th, 2025, DeepSeek released their latest open-weights reasoning model, DeepSeek-R1, which is on par with OpenAI’s o1 in benchmark performance.
- China’s DeepSeek Adds a Weird New Data Point to The AI RaceV3 and R1 are Impressive Work, With Many Implications – but not "China Has Caught Up"
- Novus Ordo SeclorumReflections on DeepSeek
- Takeaways from sketching a control safety caseInsights from a long technical paper compressed into a fun little commentary
- DeepSeek: Lemon, It's WednesdayIt’s been another *checks notes* two days, so it’s time for all the latest DeepSeek news.
- Exclusive: Americans overwhelmingly support AI safety mandates, new poll findsThe public would rather ban AI development than have no regulation at all
- On DeepSeek and Export Controls“Export controls serve a vital purpose: keeping democratic nations at the forefront of AI development. To be clear, they’re not a way to duck the competition between the US and China.”
- Planning for Extreme AI RisksAre we ready for this?
- OperatorNo one is talking about OpenAI’s Operator.
- Ten people on the insideA scary scenario that's worth planning for
- Why reasoning models will generalizePeople underestimate the long-term potential of “reasoning.”
- DeepSeek Panic at the App StoreDeepSeek released v3.
- Import AI 397: DeepSeek means AI proliferation is guaranteed; maritime wardrones; and more evidence of LLM capability overhangsWelcome to Import AI, a newsletter about AI research.
- On Private GovernanceForging a new path
- Which AI to Use Now: An Updated Opinionated Guide (Updated Again 2/15)Picking your general-purpose AI
- AGI could drive wages below subsistence levelHistorically, many have feared that automation would lead to mass unemployment and lower wages.
- Stargate AI-1There was a comedy routine a few years ago.
- Open-Source AI and the FutureModest visions of an open-source AI future
- On DeepSeek's r1r1 from DeepSeek is here, the first serious challenge to OpenAI’s o1.
- The way we evaluate AI model safety might be about to breakAs systems become more capable, researchers think we need a new type of safety evaluation
- When does capability elicitation bound risk?The assumptions behind and limitations of capability elicitation have been discussed in multiple places (e.g.
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsYes, ring the true o1 replication bells for DeepSeek R1 🔔🔔🔔. Where we go next.
- Does Elon still care about AI safety?He’s doing a terrible job of showing it.
- Import AI 396: $80bn on AI infrastructure; can Intel's Gaudi chip train neural nets?; and getting better code through asking for it…I expect another punctured equilibrium in 2025 from dramatic AI progress…
- How will we update about scheming?A quantitative description of how I expect to change my mind.
- How has DeepSeek improved the Transformer architecture?DeepSeek has recently released DeepSeek v3, which is currently state-of-the-art in benchmark performance among open-weight models, alongside a technical report describing in some detail the training of the model.
- Meta Pivots on Content ModerationThere’s going to be some changes made.
- Thoughts on the conservative assumptions in AI controlWhy are we so friendly to the red team?
- Unstable DiffusionOn (some of) the latest export controls
- On the OpenAI Economic BlueprintTable of Contents
- Let me use my local LMs on Meta Ray-BansDifferent futures for AI devices and how we can start creating positive feedback loops for some types of open-weight language models.
- Extending control evaluations to non-scheming threatsBuck Shlegeris and Ryan Greenblatt originally motivated control evaluations as a way to mitigate risks from ‘scheming’ AI models: models that consistently pursue power-seeking goals in a covert way; however, many adversarial model psychologies are not well described by the standard notion of…
- Using ChatGPT is not bad for the environmentAnd a plea to think seriously about climate change without getting distracted
- How quickly could robots scale up?Some notes on robot economics.
- o1 isn’t a chat model (and that’s the point)How Ben Hylak turned from ol pro skeptic to fan by overcoming his skill issue.
- AI Generated Misinformation is Still a RiskHalf the world's population voted in 2024, relatively unaffected by AI generated deepfakes. Can we conclude that AI generated misinformation isn't the risk we thought it was?
- New AI export controls have leaked. Here’s what you need to know.The new rules are squarely focused on the frontier
- On Dwarkesh Patel's 4th Podcast With Tyler CowenDwarkesh Patel again interviewed Tyler Cowen, largely about AI, so here we go.
- Prophecies of the FloodWhat to make of the statements of the AI labs?
- The economic consequences of automating remote workRecent AI progress has shown great promise in automating cognitive tasks, like those in natural language processing and vision.
- 2025: A Look AheadAI products, research, and policy to watch in 2025
- DeepSeek V3 and the actual cost of training frontier AI modelsThe $5M figure for the last training run should not be your basis for how much frontier AI models cost.
- OpenAI #10: ReflectionsThis week, Altman offers a post called Reflections, and he has an interview in Bloomberg. There’s a bunch of good and interesting answers in the interview about past events that I won’t mention or have to condense a lot here, such as his going over his calendar and all the meetings he constantly…
- Are We on the Brink of AGI?A Tale of Two Timelines
- The Important Thing About AGI is the Impact, Not the NameReality Doesn't Care How We Interpret the Words "General Intelligence"
- Texas Plows AheadTexas' onerous AI regulation is formally introduced
2024
December 2024
- DeepSeek v3: The Six Million Dollar ModelWhat should we make of DeepSeek v3?
- The media needs to start taking AGI seriouslyAn essay for Nieman Lab's 2025 Predictions series
- o3, Oh MyOpenAI presented o3 on the Friday before Christmas, at the tail end of the 12 Days of Shipmas.
- Moravec’s paradox and its implicationsSince the birth of the field of artificial intelligence in the 20th century, researchers have observed that the difficulty of a task for humans at best weakly correlates with its difficulty for AI systems.
- LLMs Fight With Both Hands Tied Behind Their BackWe haven't yet given them access to knowledge-in-the-world
- Measure UpAn Ode to an Instrument
- AIs Will Increasingly Fake AlignmentThis post goes over the important and excellent new paper from Anthropic and Redwood Research, with Ryan Greenblatt as lead author, Alignment Faking in Large Language Models.
- The Black Spatula Project: Day FiveOff to a Roaring Start
- Import AI 395: AI and energy demand; distributed training via DeMo; and Phi-4What might fighting for freedom in an AI age look like?
- How to prepare yourself for AGIWhat if artificial general intelligence (AGI) arrives in just a couple of years, triggering an explosion in science and technology that transforms life as we know it?
- How do mixture-of-experts models compare to dense models in inference?In last week’s Gradient Updates issue, I discussed how we can guess that GPT-4o and Claude 3.5 Sonnet have significantly fewer parameters than GPT-4.
- Measuring whether AIs can statelessly strategize to subvert security measuresThe complement to control evaluations
- OpenAI's o3: The grand finale of AI in 2024A step change as influential as the release of GPT-4. Reasoning language models are the current and next big thing.
- 8 things that surprised me about AI policy in 2024In this post, our Senior Director of AI Policy, Nicklas Lundblad, shares 8 things that surprised him about AI policy in 2024.
- One Down, Many To GoReflections on year one of Hyperdimensional
- What just happenedA transformative month rewrites the capabilities of AI
- A Matter of TasteIn light of other recent discussions, Scott Alexander recently attempted a unified theory of taste, proposing several hypotheses.
- Alignment Faking in Large Language ModelsIn our experiments, AIs will often strategically pretend to comply with the training objective to prevent the training process from modifying its preferences.
- Is AI progress slowing down?Making sense of recent technology trends and claims
- The AI agent spectrumSeparating different classes of AI agents from a long history of reinforcement learning.
- The Black Spatula ProjectFixing flawed scientific papers, 5000 tokens at a time
- The Second GeminiTable of Contents
- AIs Will Increasingly Attempt ShenanigansIncreasingly, we have seen papers eliciting in AI models various shenanigans.
- The Future Is Already Here, It’s Just Not Evenly DistributedYou Can Awe Some of the People Some of the Time
- Frontier language models have become much smallerBetween the release of the original Transformer in 2017 and the release of GPT-4, language models at the frontier of capabilities became much larger.
- The o1 System Card Is Not About o1Or rather, we don’t actually have a proper o1 system card, aside from the outside red teaming reports.
- We Looked at 78 Election Deepfakes. Political Misinformation is not an AI Problem.Technology Isn’t the Problem—or the Solution.
- Thresholdso1, o1 Pro, and what comes next
- OpenAI's Reinforcement Finetuning and RL for the massesThe cherry on Yann LeCun’s cake has finally been realized.
- o1 Turns ProSo, how about OpenAI’s o1 and o1 Pro?
- 15 Times to use AI, and 5 Not toNotes on the Practical Wisdom of AI Use
- Import AI 394: Global MMLU; AI safety needs AI liability; Canada backs CohereIn the same way people treasure VHS tapes and film cameras, they will treasure vintage generative models
- What did US export controls mean for China’s AI capabilities?Four days ago, the US government announced new rules around the export of powerful chips and semiconductor manufacturing equipment to China.
- Impact Assessments are the Wrong Way to Regulate Frontier AIA simple case study--and a better way forward
- OpenAI's new model tried to avoid being shut downo1 attempted to exfiltrate its weights to avoid being shut down
- OpenAI's o1 using "search" was a PSYOPHow to understand OpenAI's o1 models as really just one wacky, wonderful, long chain of thought.
- The 2024 Transformer Gift GuideThe must-buy items for the AI policy people in your life
- Import AI 393: 10B distributed training run; China VS the chip embargo; and moral hazards of AI developmentAre you sure you're doing more than curve fitting?
November 2024
- Why a US AI "Manhattan Project" could backfire: notes from conversations in ChinaMy recent two weeks in China suggested something surprising about its AI landscape: the biggest bottleneck isn't compute – it's commitment.
- On the EU AI Code of PracticeA co-authored comment
- A new golden age of discoverySeizing the AI for Science opportunity
- OLMo 2 and building effective teams for training language modelsAnnouncing OLMo 2 - the best open-source models yet - and some of the things I've been learning about training good language models.
- Getting started with AI: Good enough promptingDon't make this hard
- OpenAI's CBRN tests seem unclearOpenAI says o1-preview can't meaningfully help novices make chemical and biological weapons. Their test results don’t clearly establish this.
- OpenAI Realtime API: The Missing ManualEverything we learned, and everything we think you need to know, from technical details on 24khz/G.711 audio, RTMP, HLS, WebRTC, to Interruption/VAD, to Cost, Latency, Tool Calls, and Context Mgmt
- Tülu 3: The next era in open post-trainingWe give you open-source, frontier-model post-training.
- Questions, Unasked and UnansweredThe US-China Commission Floats an AGI Manhattan Project
- Synthetic data is more useful than you thinkFears of ‘model collapse’ are overstated — at least for now
- What Are the Real Questions in AI?It's hard to have a constructive discussion if we're not talking about the same thing
- Import AI 392: China releases another excellent coding model; generative models and robots; scaling laws for agentsIf aliens built AI, would it also use stochastic gradient descent?
- The Choice TransitionOn the emergence of history’s reins
- Why imperfect adversarial robustness doesn't doom AI controlThere are crucial disanalogies between preventing jailbreaks and preventing misalignment-induced catastrophes.
- Shape, Symmetries, and Structure: The Changing Role of Mathematics in Machine Learning ResearchHenry Kvinge asks "What is the Role of Mathematics in Modern Machine Learning?"
- New OpenAI emails reveal a long history of mistrustGreg Brockman and Ilya Sutskever had questions about Sam Altman's intentions as early as 2017
- Win/continue/lose scenarios and execute/replace/audit protocolsIn this post, I’ll make a technical point that comes up when thinking about risks from scheming AIs from a control perspective.
- Here's What I Think We Should DoA Proactive AI Policy Agenda
- Yet another AI safety researcher has left OpenAIRichard Ngo resigned today, saying it has become "harder for me to trust that my work here would benefit the world"
- Scaling realitiesBoth stories are true. Scaling still works. OpenAI et al. still have oversold their promises.
- The Most Dangerous Thing An AI Startup Can Do Is Build For Other AI StartupsHow Codeium went from 0 to >$10m in ten months, What enterpriseready.io got wrong. A comprehensive braindump on how to be Enterprise Infra Native!
- Meta’s AI ‘safeguards’ are an elaborate fictionMeta cannot prevent misuse, despite what it might pretend
- Saving the National AI Research Resource & my AI policy outlookAs domestic AI policy gets reset, we need to choose our battles when recommending keeping the Biden administration’s work.
- AI is Racing Forward – on a Very Long RoadProgress Hasn’t Stalled; But AGI Is Not Yet Near
- Does the UK’s liver transplant matching algorithm systematically exclude younger patients?Seemingly minor technical decisions can have life-or-death effects
- Import AI 391: China's amazing open weight LLM; Fields Medalists VS AI Progress; wisdom and intelligenceHow much did you talk with an LLM this week?
- Why AI companies are eyeing the Middle EastAnd what the US government is doing about it
- AI Safety Under Republican LeadershipEverything is not about everything
- Securing AILessons from cybersecurity
- What Trump means for AI safetyA repeal of last year's executive order, for one thing
- A brief history of the automated corporationLooking back from 2041
- Import AI 390: LLMs think like people; neural Minecraft; Google's cyberdefense AIDo you believe AI will hit a wall?
- The Present Future: AI's Impact Long Before SuperintelligenceYou can start to see the outlines of an AI future, for better and worse
- On Civilizational TriumphSome additional thoughts on open-source AI
October 2024
- It’s time to take AI welfare seriouslyA new report argues that AI systems could soon deserve moral consideration
- Anthropic has hired an 'AI welfare' researcherKyle Fish joined the company last month to explore whether we might have moral obligations to AI systems
- Hold My Beer, CaliforniaTexas Enters the AI Policy Chat
- Why I build open language modelsReflections after a year at the Allen Institute for AI and on the battlefields of open-source AI.
- Import AI 389: Minecraft vibe checks; Cohere's multilingual models; and Huawei's computer-using agentsWill AI ever seek to synthesize biological life?
- To Change the World, Set a Bold Target: Moore's Law as Self-Fulfilling ProphecyAll progress depends on the unreasonable forecast
- Essay Writing as Personal SovereigntyAnnouncing one of the winners of Cosmos Institute's essay contest
- Abandon compute thresholds at your perilThey are the best tool we have to reduce the burden of AI regulation
- Claude Sonnet 3.5.1 and Haiku 3.5Anthropic has released an upgraded Claude Sonnet 3.5, and the new Claude Haiku 3.5.
- AI safety tax dynamicsEarlier AI capabilities — research and coordination — could help us to navigate later safety problems. And the highest risk probably comes in the era before strong superintelligence.
- Claude's agentic future and the current state of the frontier modelsHow Claude's computer use works. Where OpenAI, Anthropic, and Google all have a lead on eachother.
- Stafford Beer and AI as Variety EngineeringThoughts on The Unaccountability Machine by Dan Davies
- When you give a Claude a mouseSome quick impressions of an actual agent
- "Be Embraced, Ye Millions"Reflections on Progress Studies
- Import AI 388: Simulating AI policy; omni math; consciousness levelsWill UX innovations be just as important as research innovations?
- Is OpenAI being fair to its non-profit?Experts fear OpenAI’s conversion to a for-profit might not adequately compensate its charitable arm
- The Mask Comes Off: At What Price?The Information reports that OpenAI is close to finalizing its transformation to an ordinary Public Benefit B-Corporation.
- Thinking Like an AIA little intuition can help
- Safety tax functionsA conceptual exploration
- What is AI, and How Do We Govern It?Selections from a recent speech
- Building on evaluation quicksandOn the state of evaluation for language models.
- No, A Bot Didn't Just Make $150M in CryptoBeware the Startling Anecdote
- Import AI 387: Overfitting vs reasoning; distributed training runs; and Facebook's new video modelsWhat is the space of intelligence?
- What would evidence-based AI policy look like?If-then commitments might be the answer
- Learning to Explore: AlphaProof and o1 Show The Path to AI CreativityAn Impressive Start, But Many Hurdles Lie Ahead
- $2 H100s: How the GPU Rental Bubble BurstLast year, H100s were $8/hr if you could get them. Today, there's 7 different resale markets selling them under $2. What happened?
- A Policy Agenda for Defensive Acceleration Against AI RisksThe same AI capabilities underlie benefits and harms. How can we assure AI safety without curtailing beneficial applications? Defensive acceleration charts a possible course.
- Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't schemingOne strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT).
- Decentralized Training and the Fall of Compute ThresholdsWill decentralized training break compute thresholds for good?
- Joshua Achiam Public Statement AnalysisJoshua Achiam is the OpenAI Head of Mission Alignment
- How scaling changes model behaviorSome trends are reasonable to extrapolate, some are not. Even for the trends we are succeeding at extrapolating, it is not clear how that signal translates into different AI behaviors.
- How I use AI as a journalistLLMs have become an indispensable part of my workflow
- Import AI 386: Google's chip-designing AI keeps getting better; China does the simplest thing with Emu3; Huawei's 8-bit data format…What happens when robots apply mechanistic interpretability to humans?...
- What if everyone is wrong about what AI does?Both critics and supporters seem to think AI is a "human remover". What if they're both wrong?
- AI Creativity Is A Question of Quality, Not NoveltyIt Actually Matters How Well The Bear Dances
- AI Safety Culture Confronts CapitalismSB1047’s veto, OpenAI’s turnover, and a constant treadmill pushing AI startups to be all too similar to big technology name brands.
- Newsom Vetoes SB 1047It’s over, until such a future time as either we are so back, or it is over for humanity.
September 2024
- What Comes After SB 1047?Plus, some unsolicited advice for AI safety
- Gavin Newsom has caved to the billionairesSB 1047's veto is a win for tech companies — and a loss for everyone else
- A basic systems architecture for AI agents that do autonomous researchAnd diagrams describing how threat scenarios involving misaligned AI involve compromising the system in different places.
- On AnalogiesHopfield networks, metaphors, and policy development
- The OpenAI Pastiche EditionMusings on the latest drama, extra thoughts on the o1 models
- How to prevent collusion when using untrusted models to monitor each otherSuppose you’ve trained a really clever AI model, and you’re planning to deploy it in an agent scaffold that allows it to run code or take other actions.
- Llama 3.2 Vision and Molmo: Foundations for the multimodal open-source ecosystemOpen models, tools, examples, limits, and the state of training multimodal models.
- Can AI automate computational reproducibility?A new benchmark to measure the impact of AI on improving science
- Lies and deception: Andreessen Horowitz’s SB 1047 campaign is as misleading as it getsThe venture capital giant has repeatedly peddled untruths in an effort to kill the AI regulation bill
- What It’s Like To Solve a Math Olympiad ProblemAnd What This Tells Us About Creativity, Reasoning, And AI
- The Machine Stops: Will AI Lead to Freedom or Control?How E.M. Forster's dystopia mirrors our AI-driven crossroads
- AGI and Political OrderThe political economy of AGI
- An Extraordinary AlienOpenAI's o1 Challenges AI Policy
- GPT-o1Terrible name (with a terrible reason, that this ‘resets the counter’ on AI capability to 1, and ‘o’ as in OpenAI when they previously used o for Omni, very confusing).
- Import AI 385: False memories via AI; collaborating with machines; video game permutationsWould AI risk be easier to work on if it took on a physical form, like a giant asteroid heading towards earth?
- Reverse engineering OpenAI’s o1What productionizing test-time compute shows us about the future of AI. Exploration has landed in language model training.
- Scaling: The State of Play in AIA brief intergenerational pause...
- AI Could Break Things; Let's Use It As a Wakeup Call To Make Them StrongerStop Treating AI Policy Tradeoffs As Zero-Sum
- OpenAI's new models 'instrumentally faked alignment'The o1 safety card reveals a range of concerning capabilities, including scheming, reward hacking, and biological weapon creation.
- Prediction Markets and AI DiffusionOn the fragility of good ideas
- Something New: On OpenAI's "Strawberry" and ReasoningSolving hard problems in new ways
- AI, centralization, and the One RingConcerns with amassing power in a single AGI project
- Futures of the data foundry business modelScale AI’s future versus further scaling of language model performance. How Nvidia may take all the margins from the data market, too.
- AI employees are defying their employers to support SB 1047In speaking out, they’re showing just how important the bill is
- A post-training approach to AI regulation with Model SpecsAnd why the concept of mandating “model spec’s” could be a good start.
- What's Missing From LLM Chatbots: A Sense of PurposeKenneth Li argues that LLM chatbots are missing a "goal" for each conversation and proposes fixing this with human-in-the-loop model evaluations.
- OpenAI’s Strawberry, LM self-talk, inference scaling laws, and spending more on inferenceWhether or not scaling works, we should spend more on inference.
- The Timing of AI RegulationWhy we have more time than regulation advocates fear
- AI and the Technological Richter ScaleThe Technological Richter scale is introduced about 80% of the way through Nate Silver’s new book On the Edge.
- OLMoE and the hidden simplicity in training better foundation modelsAi2 released OLMoE, which is probably our “best” model yet relative to its peers, but not much has changed in the process.
- On the UBI PaperWould a universal basic income (UBI) work?
- Import AI 384: Accelerationism; human bit-rate processing; and Google stuffs DOOM inside a neural networkHow much of today's technology will be deemed important in one hundred years?
August 2024
- Rather Than Arguing About What We Don't Know, Let's Work Together To Find OutA Lesson From The SB 1047 Debate
- Post-apocalyptic educationWhat comes after the Homework Apocalypse
- On AI "Black Boxes"Will we ever understand neural networks?
- On the current definition of open-source AI and the state of the data commonsThe Open Source Initiative (OSI) is working towards a definition.
- SB 1047: Final Takes and Also AB 3211This is the endgame.
- Dangerous capability tests should be harderWe should spend less time proving that today’s AIs are safe and more time figuring out how to tell if tomorrow’s AIs are dangerous.
- Guide to SB 1047We now likely know the final form of California’s SB 1047.
- SB 1047 is AmendedConcluding thoughts on a fundamentally flawed bill
- AI companies are pivoting from creating gods to building products. Good.Turning models into products runs into five challenges
- Import AI 383: Automated AI scientists; cyborg jellyfish; what it takes to run a clusterIs AI as useful as concrete?
- On Nous Hermes 3 and classifying a "frontier model"The latest model from one of the most popular fine-tuning labs makes us question how a model should be identified as a “frontier model.”
- The Red Queen problemMeasuring the pace of scientific progress
- Danger, AI Scientist, DangerWhile I finish up the weekly for tomorrow morning after my trip, here’s a section I expect to want to link back to every so often in the future.
- Fields that I reference when thinking about AI takeover preventionIs AI takeover like a nuclear meltdown? A coup? A plane crash?
- Change blindness21 months later
- Import AI 382: AI systems are societal mirrors; China gets chip advice via LLMs; 25 million medical imagesIt's getting hard to ignore the power of these systems
- On Algorithmic Impact AssessmentsThe AI Policy Trojan Horse
- A recipe for frontier model post-trainingApple, Meta, and Nvidia all agree — synthetic data, iterative training, human preference labels, and lots of filtering.
- Is the UAE running an AI influence campaign?Twitter is overrun with bot-like accounts praising the UAE's approach to AI
- AI for Bad: A Saboteur’s GuideA step-by-step manual for ensuring AI’s potential remains untapped while maximizing the likelihood of dystopia
- Import AI 381: Chips for Peace; Facebook segments the world; and open source decentralized trainingHow much chain of thought reasoning do humans do themselves?
- We Need Positive Visions for AI Grounded in WellbeingJoel Lehman and Amanda Ngo demystify "beneficial AI" by grounding the notion in human wellbeing and chart a path for developing it.
- Resilience and Adaptation to Advanced AIWhen model safeguards fail, what's our backup plan? AI resilience is crucial for preparing for the widespread diffusion of increasingly powerful AI.
- Positive-sum symbiosisWhy you should (sometimes) be kind to AIs
- On speaking to AIVoice changes a lot of things
- Senate Commerce Committee advances lots of AI billsTed Cruz isn't happy about it, though
July 2024
- GPT-4o-mini changed ChatBotArenaAnd how to understand Llama 3.1’s results on the community's favorite benchmark.
- AI companies are falling short on their promises to the White HouseA year on from making voluntary commitments to the White House, adherence is patchy
- RTFB: California's AB 3211Some in the tech industry decided now was the time to raise alarm about AB 3211.
- California's Other Big AI BillOn AB 3211, an under-discussed deepfake bill
- Import AI 380: Distributed 1.3bn parameter LLM; math AI; and why reality is hard for AiWhat is our responsibility to machines that may become moral patients?
- AI existential risk probabilities are too unreliable to inform policyHow speculation gets laundered through pseudo-quantification
- Yes, we still have to workThe automated luxury paradise is still just science fiction.
- Llama 3.1 and the "Path Forward"Will policymakers permit an open-source future?
- Llama Llama-3-405B?It’s here.
- The last era of human mistakesAuthor’s remark, four months later: in retrospect I am vaguely dissatisfied with this piece.
- Tool or TyrantRemarks given by Cosmos Chair Brendan McCord at the "Lyceum Project," a collaboration between Stanford and Oxford hosted at Aristotle's old school
- Llama 3.1 405B, Meta’s AI strategy, and the new, open frontier model ecosystemDefining the future of the AI economy and regulation. Is Meta’s AI play equivalent to the Unix stack for open-source software?
- Confronting Impossible FuturesWe shouldn't be certain about what is next, but we should plan for it
- The Winds of AI WinterMar-Jun 2024 Recap: People are raising doubts about AI Summer. Why AI Engineers are the solution.
- What Kamala Harris means for AI regulationThe vice-president has acknowledged the existential risks of AI and called for legislation to tackle those risks
- The return on the bicameral mindPresence and tulpamancers in the age of agents
- Where I Stand on the Biggest AI IssuesSome philosophy, some policy, some history
- SB 1047, AI regulation, and unlikely allies for open modelsThe rallying of the open-source community against CA SB 1047 can represent a turning point for AI regulation.
- Meta-funded group floods Facebook with anti-AI regulation adsThe American Edge Project has spent $150k+ stoking China fears and urging opposition to "anti-innovation" laws
- Import AI 379: FlashAttention-3; Elon's AGI datacenter; distributed training.If compute isn't everything, why are so many people betting that it is?
- A Legal Framework for AI AgentsPart II of "Is AI in Trouble at the Supreme Court?"
- Decomposing AgencyDesires without capabilities; capabilities without desires
- Tech companies are trying to kill California's AI regulation billAndreessen Horowitz, Y Combinator, and Big Tech are trying to stop SB 1047 from passing
- Import AI 378: AI transcendence; Tencent's one billion synthetic personas, Project Naptime...How the wisdom of the crowd holds true for AI systems as well as people...
- What Labour means for AIEverything you need to know about Keir Starmer and Peter Kyle’s plans
- Gradually, then Suddenly: Upon the ThresholdSmall improvements can lead to big changes
- New paper: AI agents that matterRethinking AI agent benchmarking and evaluation
- Switched to Claude 3.5Speculations on the role of RLHF and why I love the model for people who pay attention.
- Is AI in Trouble at the Supreme Court? (Part One)Previewing AI legal battles to come
- Lawrence Lessig is very worried about freely available AI model weights"Open weights create a unique kind of risk," says the open source pioneer
- What Overruling Chevron Means for AILoper Bright v. Raimondo and the future of AI policy
June 2024
- AI scaling mythsScaling will run out. The question is when.
- RLHF roundup: Getting good at PPO, sketching RLHF’s impact, RewardBench retrospective, and a reward model competitionThings to be aware of if you work on language model fine-tuning.
- On Claude 3.5 SonnetThere is a new clear best (non-tiny) LLM.
- Frontiers in synthetic dataTrends in synthetic data that I'm watching closely in the leading open and closed models.
- On OpenAI's Model SpecThere are multiple excellent reasons to publish a Model Spec like OpenAI’s, that specifies how you want your model to respond in various potential situations.
- Latent Expertise: Everyone is in R&DIdeas come from the edges, not the center
- The Political Economy of AI RegulationThinking realistically about SB 1047 and bills like it
- Thoughts on: The Handover by David RuncimanThomas Hobbes, AI, and 450 year old alignment problems
- Ilya Sutskever is betting on very cheap superintelligenceReading between the lines of "Safe Superintelligence Inc".
- On DeepMind's Frontier Safety FrameworkOn DeepMind’s Frontier Safety Framework
- Text-to-video AI models are already abundant, but the products?Signs point to a general use Sora-like model coming very soon, maybe even with open weights.
- Getting 50% (SoTA) on ARC-AGI with GPT-4oYou can just draw more samples
- Import AI 377: Voice cloning is here; MIRI's policy objective; and a new hard AGI benchmarkCan you evolve your way to Einstein?
- OpenAI #8: The Right to WarnThe fun at OpenAI continues.
- The Leopold Model: Analysis and ReactionsPreviously: On the Podcast, Quotes from the Paper
- AiPhoneApple was for a while rumored to be planning launch for iPhone of AI assisted emails, texts, summaries and so on including via Siri, to be announced at WWDC 24.
- AI for the rest of usApple Intelligence makes a lot of sense when you get out of the AI bubble. Plus, the cool technical details Apple shared about their language models "thinking different."
- The only AI certainty is uncertaintyA response to Jack Clark's GPT-2 reflections
- AI takeoff and nuclear warDrivers of all-out war, and strategies to mitigate the risk
- Apple Intelligence and the Shape of Things to ComeTwo conflicting visions about the future of AI
- On the future of language modelsMechanistic but zoomed-out takes about the trajectory of the technologies
- What Apple's AI Tells Us: Experimental Models⁴Siri versus the machine god?
- Access to powerful AI might make computer security radically easierAI might be really helpful for reducing security risk.
- Import AI 376: African language test; hyper-detailed image descriptions; 1,000 hours of Meerkats.Will an open source model get released in 2024 that cost more than $100m to train?
- On Dwarkesh's Podcast with Leopold AschenbrennerPreviously: Quotes from Leopold Aschenbrenner’s Situational Awareness Paper
- Quotes from Leopold Aschenbrenner's Situational Awareness PaperThis post is different.
- A case study in reproducibility of evaluation with RewardBenchA very technical deep dive. Lessons in evaluating, using, and training reward models on open-source infrastructure.
- A Response to "Situational Awareness"I think the scenario is wrong; let's come up with a plan that works either way
- Doing Stuff with AI: Opinionated Midyear EditionAI systems have gotten more capable and easier to use
- SB 1047 Is WeakenedIt looks like Scott Weiner’s SB 1047 is now severely weakened.
- Learning to Love the InscrutableA review of How Life Works by Philip Ball
- A realistic path to robotic foundation modelsSome thoughts and excitement after revisiting the industry thanks to Physical Intelligence founders Sergey Levine and Chelsea Finn. Not “agents” and not “AGI.”
- OpenAI employee says he was fired for raising security concerns to boardLeopold Aschenbrenner also said he was interrogated about his team’s “loyalty to the company”
- AI catastrophes and rogue deploymentsIt’s interesting to classify possible AI catastrophes based on whether or not they involve a "rogue deployment".
- Import AI 375: GPT-2 five years later; decentralized training; new ways of thinking about consciousness and AI…Are today's AGI obsessives trafficking more in fiction than in fact?...
- Scientists should use AI as a tool, not an oracleHow AI hype leads to flawed research that fuels more hype
May 2024
- Grounding the Conversation About AIOrganizing The (AI) World’s Disagreements And Making Them Universally Accessible and Constructive
- The Gemini 1.5 ReportThis post goes over the extensive report Google put out on Gemini 1.5.
- Wiener: SB 1047 is 'not looking to cover startups'At a town hall event, California Sen. Scott Wiener said the bill is facing pushback from big tech companies
- Deepfakes and the Art of the PossibleThe C2PA standard is deeply flawed, but it may be fixable
- OpenAI: Helen Toner SpeaksHelen Toner went on the TED AI podcast, giving us more color on what happened at OpenAI.
- The AI Policy AtlasA short guide to AI policy
- Sam Altman was 'outright lying to the board', says former board memberIn an interview with TED, Helen Toner said that four OpenAI board members "couldn't believe things that Sam was telling us"
- We aren’t running out of training data, we are running out of open training dataData licensing deals, scaling, human inputs, and repeating trends in open vs. closed LLMs.
- OpenAI: FalloutPreviously: OpenAI: Exodus (contains links at top to earlier episodes), Do Not Mess With Scarlett Johansson
- I am the Golden Gate BridgeEasily Interpretable Summary of New Interpretability Paper
- Import AI 374: China's military AI dataset; platonic AI; brainlike convnetsPlus, a poem about meeting aliens (well, AGI)
- Four Singularities for ResearchThe rise of AI is creating both crisis and opportunity
- The Schumer Report on AI (RTFB)Or at least, Read the Report (RTFR).
- Alliance for the Future director Brian Chau has history of racist, sexist remarksIn a post on AI regulation, Chau falsely claimed George Floyd was a "domestic abuser"
- California Senate Passes SB 1047Some additional notes on the most important AI legislation in the country
- Do Not Mess With Scarlett JohanssonI repeat.
- Name, image, and AI’s likenessCelebrity’s power will only grow in the era of infinite content.
- On Dwarkesh's Podcast with OpenAI's John SchulmanDwarkesh Patel recorded a Podcast with John Schulman, cofounder of OpenAI and at the time their head of current model post-training.
- Import AI 373: Guaranteed safety; West VS East AI attitudes; MMLU-Pro…As the money and stakes get higher, things will get crazier…
- On AI Alignment and 'Superalignment'The institutional dynamics of AI safety and alignment, and answering the unanswerable
- OpenAI: ExodusPreviously: OpenAI: Facts From a Weekend, OpenAI: The Battle of the Board, OpenAI: Leaks Confirm the Story, OpenAI: Altman Returns, OpenAI: The Board Expands.
- The most interesting startup idea I've seen recently: AI for epistemicsIf transformative AI might come soon and you want to help that go well, one strategy you might adopt is building something that will improve as AI gets more capable.
- Meet Meta's AI lobbying armyWith 30 lobbyists and seven agencies, the company is primed to push its agenda on Washington
- OpenAI is haemorrhaging safety talentSince the failed Altman ouster, many of the company's most safety-conscious employees have left
- The “ethics vs safety” fight misses the real enemy: Big TechWhile advocates for regulation squabble, Big Tech is pushing for no rules at all
- GPT-4o My and Google I/O DayAt least twice the speed!
- OpenAI chases HerChatGPT left the textbox and where AI is leading society.
- What OpenAI didA new model opens up new possibilities
- Import AI 372: Gibberish jailbreak; DeepSeek's great new model; Google's soccer-playing robotsIf neural nets didn't work, what would we be using instead?
- What GPT-4o illustrates about AI RegulationThe important difference between regulating technology use and regulating conduct
- How much AI inference can we do?Suppose you have a bunch of GPUs.
- Superhuman?What does it mean for AI to be better than a human? And how can we tell?
- OpenAI’s Model (behavior) Spec, RLHF transparency, personalization questionsNow we will have some grounding for when weird ChatGPT behaviors are intended or side-effects — shrinking the Overton window of RLHF bugs.
- I Got 95 Theses But a Glitch Ain't OneOr rather Samuel Hammond does. Tyler Cowen finds it interesting but not his view.
- The AI Republic of LettersReflections on a fragile world
- Preventing model exfiltration with upload limitsUnlike most files you might want to secure, model weights are extremely big. This might make them much easier to secure.
- Catching AIs red-handedIf your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before.
- Managing catastrophic misuse without robust AIHow could an AI lab serving AIs to customers manage catastrophic misuse without solving adversarial robustness?
- The case for ensuring that powerful AIs are controlledLabs should make sure that powerful models can't cause unacceptably bad outcomes even if the AIs try to.
- Untrusted smart models and trusted dumb modelsThe easiest way to be sure an AI isn't scheming against you is to note that it's too dumb to pull that off. What happens if that's the only way we have to rule out scheming?
- Import AI 371: CCP vs Finetuning; why people are skeptical of AI policy; a synthesizer for a LLMWelcome to Import AI, a newsletter about AI research.
- One Conversation is Worth a Thousand Angry TakesMake The Internet Better Using This One Weird Trick
- Freeing the chatbotIntelligence, of a sort, is going to be all around us
- Q&A on Proposed SB 1047Previously: On the Proposed California SB 1047.
- What Good AI Policy Looks LikeHighlighting some positive steps at the federal and state levels
- How RLHF works, part 2: A thin line between useful and lobotomizedMany, many signs of life for preference fine-tuning beyond spoofing chat evaluation tools.
April 2024
- AI leaderboards are no longer useful. It's time to switch to Pareto curves.What spending $2,000 can tell us about evaluating AI agents
- Phi 3 and Arctic: Outlier LMs are hintsModels that seem totally out of scope from recent open LLMs give us a sneak peek of where the industry will be in 6 to 18 months.
- AI stocks could crashAnd this could have implications for timelines and AI safety.
- Import AI 370: 213 AI safety challenges; everything becomes a game; Tesla's big clusterAre AI systems more like religious artifacts or disposable entertainment?
- The market expects AI software to create trillions of dollars of value by 2027Does that seem high or low to you?
- What To Expect When You’re Expecting GPT-5Where Will AI Go From Here?
- Synthetic Data in AI: Implications for PolicyWill future AI training data be generated by AI?
- AGI is what you want it to beCertain definitions of AGI are backing people into a pseudo-religious corner.
- Innovation through promptingDemocratizing educational technology... and more
- On Llama-3 and Dwarkesh Patel's Podcast with ZuckerbergIt was all quiet.
- Financial Market Applications of LLMsAn article contributed by Richard Dewey, Co-founder and CEO of Proven, and Ciamac Moallemi, a William von Mueffling Professor of Business at Columbia University
- Llama 3: Scaling open LLMs to AGIMeta shows that scaling won't be a limit for open LLM players in the near future.
- Who Governs the Internet?Exploring a question that underlies our era
- Stop "reinventing" everything to solve alignmentIntegrating some non-computing science into reinforcement learning from human feedback (RLHF) can give us the models we want. Bonus: OLMo 1.7-7B.
- The end of the “best open LLM”Modeling the compute versus performance tradeoff of many open LLMs.
- Import AI 369: Conscious machines are possible; AI agents; the varied uses of synthetic dataWho do you know who doesn't know about AI advances but should?
- Compute Thresholds are IneffectiveWhy compute thresholds are a poor tool for policymakers
- RTFB: On the New Proposed CAIP AI BillA New Bill Offer Has Arrived
- What just happened, what is happening nextThe tasks AI can do well are expanding rapidly
- A Brief Overview of Gender Bias in AIA small selection of important work done (and currently being done) to uncover, evaluate, and measure different aspects of gender bias in AI models
- Import AI 368: 500% faster local LLMs; 38X more efficient red teaming; AI21's FrankenmodelIf AI had truly superhuman capabilities, could we even effectively measure that?
- Open-Source AI has Overwhelming SupportA quick analysis
- Beyond One-Size-Fits-All FairnessNavigating AGI part five of seven
- On Technological DestinyWho decides where we are headed?
- We disagree on what open-source AI should mean... and that's okay. How to read what multiple people mean by the word openness and see through the PR speak.
- Tech policy is only frustrating 90% of the timeThat’s what makes it worthwhile
- Import AI 367: Google's world-spanning model; breaking AI policy with evolution; $250k for alignment benchmarksAI is neither rocket science nor brain surgery... which is very strange!
- Notes on Dwarkesh Patel's Podcast with Sholto Douglas and Trenton BrickenDwarkesh Patel continues to be on fire, and the podcast notes format seems like a success, so we are back once again.
March 2024
- Mamba ExplainedIs Attention all you need? Mamba, a novel AI model based on State Space Models (SSMs), emerges as a formidable alternative to the widely used Transformer models, addressing their inefficiencies
- On the necessity of a sinWhy treating AI like a person is the future
- DBRX: The new best open model and Databricks’ ML strategyDatabricks’ new model is surpassing the performance of Mixtral and Llama 2 70B while still being in a size category that's reasonably accessible.
- Import AI 366: 500bn text tokens; Facebook vs Princeton; why small government types hate the Biden EOplus: a story about a pleasant singularity
- On Lex Fridman's Second Podcast with AltmanLast week Sam Altman spent two hours with Lex Fridman (transcript).
- A Classical Liberal Framework for AI RegulationA new piece from me in National Affairs
- Disrupted: work in the age of AGINavigating AGI part four of seven
- Evaluations: Trust, performance, and price (bonus, announcing RewardBench)Evaluation is not only getting harder with modern LLMs getting more complicated, it’s getting harder because it means something different.
- On the Gladstone ReportLike the the government-commissioned Gladstone Report on AI itself, there are two sections here.
- Import AI 365: WMD benchmark; Amazon sees $1bn training runs; DeepMind gets closer to its game-playing dreamThere is no absolution in being an absolutist.
- On DevinIntroducing Devin
- Which AI should I use? Superpowers and the State of PlayAnd then there were three
- I Don't See How Comparative Advantage Applies In a World of Strong AIEconomic models are based on simplifying assumptions that may not hold in an AI-saturated future
- AI MysticismOn the perils of 'black box' AI
- I, Cyborg: Using Co-IntelligenceHow I used AI in my book about AI
- Model commoditization and product moatsWhere moats are tested now that so many people have trained GPT4 class models. Claude 3, Gemini 1.5, Inflection 2.5, and Mistral Large are here to party.
- AI safety is not a model propertyTrying to make an AI model that can’t be misused is like trying to make a computer that can’t be used for bad things
- OpenAI: The Board ExpandsIt is largely over.
- Import AI 364: Robot scaling laws; human-level LLM forecasting; and Claude 3It increasingly feels like LLMs are observing us rather than the other way round.
- Car-GPT: Could LLMs finally make self-driving cars happen?Exploring the utility of large language models in autonomous driving: Can they be trusted for self-driving cars, and what are the key challenges?
- The koan of an open-source LLMA proposal for a new definition of an “open-source” LLM and why no definition will ever just work.
- Do text embeddings perfectly encode text?'Vec2text' can serve as a solution for accurately reverting embeddings back into text, thus highlighting the urgent need for revisiting security protocols around embedded data
- On Claude 3.0Claude 3.0
- A safe harbor for AI evaluation and red teamingAn argument for legal and technical safe harbors for AI safety and trustworthiness research
- Read the RoonRoon, member of OpenAI’s technical staff, is one of the few candidates for a Worthy Opponent when discussing questions of AI capabilities development, AI existential risk and what we should do about it.
- Software's Romantic EraConsciousness, "consciousness," and Anthropic's latest opus
- Captain's log: the irreducible weirdness of prompting AIsAlso, we have a prompt library!
- Import AI 363: ByteDance's 10k GPU training run; PPO vs REINFORCE; and generative everythingThings feel confusing right now because things are confusing right now - and they'll become more confusing.
- Notes on Dwarkesh Patel's Podcast with Demis HassabisDemis Hassabis was interviewed twice this past week.
February 2024
- Science, Standards, LawsThe order of operations for AI regulation
- How to cultivate a high-signal AI feedBasic tips on how to assess inbound ML content and cultivate your news feed.
- On the Societal Impact of Open Foundation ModelsAdding precision to the debate on openness in AI
- The Gemini Incident ContinuesPreviously: The Gemini Incident (originally titled Gemini Has a Problem)
- Import AI 362: Amazon's big speech model; fractal hyperparameters; and Google's open modelsThe full capabilities of today's AI systems are all broadly undiscovered.
- Why Doesn’t My Model Work?On some of the issues that can cause a model to seem good when it isn’t, and ways in which these kinds of mistakes can be prevented
- For Aristotle, being realistic about the promise of AI requires us first to see what’s at stakeFirst installment of a new series from the Cosmos Institute’s reading group
- AGI and Political UnrestHow government can maintain stability during turbulent times
- The Gemini Incident[Originally named: Gemini Has a Problem.]
- Sora WhatHours after Google announced Gemini 1.5, OpenAI announced their new video generation model Sora. Its outputs look damn impressive.
- The One and a Half GeminiPreviously: I hit send on The Third Gemini, and within half an hour DeepMind announced Gemini 1.5.
- Book review: "Power and Progress"In which Daron Acemoglu and Simon Johnson fail to convince me that innovation needs to be steered away from automation.
- Strategies for an Accelerating FutureFour questions to ask your organization.
- Import AI 361: GPT-4 hacking; theory of minds in LLMs; and scaling MoEs + RLHow many tokens would you need to represent your life? 10,000? 100,000? 1 million?
- Models on the frontline: AI's defensive roleNavigating AGI part three of seven
- 10 Sora and Gemini 1.5 follow-ups: code-base in context, deepfakes, pixel-peeping, inference costs, and moreThe cutting edge technical discussions beneath the wow factor.
- Why Are LLMs So Gullible?It’s Because They’re Naive and Constantly Confused
- OpenAI’s Sora for video, Gemini 1.5's infinite context, and a secret Mistral modelEmergency blog! Three things you need to know from the ML world that arrived on Thursday.
- The Case for Public AI InfrastructureHow the government can help, not hinder, AI development
- The Third GeminiWe have now had a little over a week with Gemini Advanced, based on Gemini Ultra.
- Why reward models are key for alignmentIn an era dominated by direct preference optimization and LLM-as-a-judge, why do we still need a model to output only a scalar reward?
- A Rebuttal to Zvi MowshowitzResponding to some criticisms of last week’s article
- Import AI 360: Guessing emotions; drone targeting dataset; frameworks for AI alignmentAre there alternatives to the transformer which are roughly as compute efficient but entirely different in architecture?
- On the Proposed California SB 1047California Senator Scott Wiener of San Francisco introduces SB 1047 to regulate AI. I have put up a market on how likely it is to become law.
- The line between risk and progressNavigating AGI part two of seven
- California's Effort to Strangle AIOn SB 1047, a dangerous attempt at safety
- Google's Gemini Advanced: Tasting Notes and ImplicationsAnd then there were two.
- Alignment-as-a-service: Scale AI vs. the new guysScale’s making over $750 million per year selling data for RLHF, who’s coming to take it?
- On the Debate Between Jezos and LeahyPreviously: Based Beff Jezos and the Accelerationists
- Import AI 359: $1 billion gov supercomputer; Apple’s good synthetic data technique; and a thousand-year old data libraryIf the 20th century had psychohistory, will the 21st century have praedictiohistory?
- When is a capability truly worrying?Navigating AGI part one of seven
- On Dwarkesh's 3rd Podcast with Tyler CowenThis post is extensive thoughts on Tyler Cowen’s excellent talk with Dwarkesh Patel.
- Open Language Models (OLMos) and the LLM landscapeA small model at the beginning of big changes.
January 2024
- AI Biorisk: A Dose of RealityWhat Works--and What Doesn't--in Mitigating AI Risk
- What Can be Done in 59 Seconds: An Opportunity (and a Crisis)Five analytical tasks in under a minute
- Import AI 358: The US Government’s biggest AI training run; hacking LLMs by hacking GPUs; chickens versus transformersHow small will the most widely-used LLM be in 2025?
- Model merging lessons in The Waifu Research DepartmentWhen what seems like pure LLM black magic is actually supported by the literature.
- Anatomy of a HackPractical hacks are complicated. That means complicated systems are unsafe.
- Local LLMs, some facts some fictionThe deployment path that’ll break through in 2024. Plus, checking in on strategies across Big Tech and AI leaders.
- Will AI transform law?The hype is not supported by current evidence
- Book review: Inspectors for PeaceInspectors for Peace: A History of the International Atomic Energy Agency by Elisabeth Roehrlich
- Generative AI’s end-run around copyright won’t be resolved by the courtsOutput similarity is a distraction
- Import AI 357: Facebook's open source AGI plan; Google beats humans at geometry problems; and Intel makes its GPUs betterSome work in AI alignment feels more like eschatology than science
- Humanity's Next LeapThinking about the *real* elephant in the room
- Free as in Speech, or Free as in Beer?Why Open-Source AI is Essential
- Multimodal blogging: My AI tools to expand your audienceA fun demo on how generative AI can transform content creation, and tools for my fellow writers on Substack!
- On Anthropic's Sleeper Agents PaperThe recent paper from Anthropic is getting unusually high praise, much of it I think deserved.
- What is “Prompt Injection”, And Why Does It Fool Chatbots?LLMs See Words, But Not Context
- The Lazy Tyranny of the Wait CalculationTaking AI timelines seriously
- Import AI 356: China's good LLM; AI credit scores; and fooling VLMs with REBUSIf you believed superintelligence was arriving within a decade, how would you live differently? And would it be rational or irrational to do so?
- Deep learning for single-cell sequencing: a microscope to see the diversity of cellsOn how Deep Learning has played as a key enabler for advancing single-cell sequencing technologies.
- Let's Talk About AI 'X-Risk'Why I don't lose sleep about existential risk from AI
- Multimodal LM roundup: Unified IO 2, inputs and outputs, Gemini, LLaVA-RLHF, and RLHF questionsA sampling of recent happenings in the multimodal space. Be sure to expect more this year.
- Import 355: Local LLMs; scaling laws for inference; free Mickey MouseWhat do you think AI systems won't be capable of at the end of 2024?
- Signs and PortentsSome hints about what the next year of AI looks like
- AI Impacts Survey: December 2023 EditionKatja Grace and AI impacts survey thousands of researchers on a variety of questions, following up on a similar 2022 survey as well as one in 2016.
- Existential Pessimism vs. Accelerationism: Why Tech Needs a Rational, Humanistic "Third Way"We need a balanced, enterprising optimism—an ambitious and clear vision for what constitutes the good in technology, coupled with the awareness of its dangers
- Copyright Confrontation #1Lawsuits and legal issues over copyright continued to get a lot of attention this week, so I’m gathering those topics into their own post.
- It's 2024 and they just want to learnThe state of the ML communities big and small starting 2024. My general expectations for the year.
2023
December 2023
- Will scaling work?Data bottlenecks, generalization benchmarks, primate evolution, intelligence as compression, world modelers, and other considerations
- Import AI 354: Distributed LLM inference; CCP-approved dataset; AI scientistsHow will PR work when the goal is shifting the frame of an LLM rather than public opinion?
- On OpenAI's Preparedness FrameworkPreviously: On RSPs.
- State-space LLMs: Do we need Attention?Mamba, StripedHyena, Based, research overload, and the exciting future of many LLM architectures all at once.
- An AI Haunted WorldIntelligence, everywhere.
- Import AI 353: AI bootstrapping; LLMs as inventors; Facebook releases a free moderation toolLike trains and containers, AI will force a certain kind of global infrastructure.
- Salmon in the LoopOn fish counting – a complex sociotechnical problem in a field that is going through the process of digital transformation.
- Would We Really Shut Down A Misbehaving AI?Ask The Old OpenAI Board How Things Went When They Tried To Shut Down Sam Altman
- Are open foundation models actually more risky than closed ones?A policy brief on open foundation models
- Big Tech's LLM evals are just marketingA PSA everyone needs. The importance of a wait and see attitude when it comes to new models, big and small, open and closed.
- OpenAI: Leaks Confirm the StoryPreviously: OpenAI: Altman Returns, OpenAI: The Battle of the Board, OpenAI: Facts from a Weekend, additional coverage in AI#41.
- Import AI 352: Asteroids and AI policy; privacy-preserving AI benchmarks; and distributed inferenceWhat might a 'close call' look like for an AI disaster?
- Mixtral: The best open model, MoE trade-offs, release lessons, Mistral raises $400mil, Google's loss, vibes vs marketingWe have an amazing open mixture of experts model for the holidays!
- An Opinionated Guide to Which AI to Use: ChatGPT Anniversary EditionA simple answer, and then a less simple one.
- Gemini 1.0It’s happening.
- Based Beff Jezos and the AccelerationistsIt seems Forbes decided to doxx the identity of e/acc founder Based Beff Jezos.
- Do we need RL for RLHF?Direct (DPO) vs. RL methods for preferences, more RLHF models, and hard truths in open RLHF work. We have more questions than answers.
- On 'Responsible Scaling Policies' (RSPs)This post was originally intended to come out directly after the UK AI Safety Summit, to give the topic its own deserved focus.
- Import AI 351: How inevitable is AI?; Distributed shoggoths; ISO an Adam replacementHow much of the future of AI is overdetermined and how much do we have agency over?
- Model alignment protects against accidental harms, not intentional onesThe hand wringing about failures of model alignment is misguided
November 2023
- OpenAI: Altman ReturnsAs of this morning, the new board is in place and everything else at OpenAI is otherwise officially back to the way it was before.
- Synthetic data: Anthropic’s CAI, from fine-tuning to pretraining, OpenAI’s Superalignment, tips, types, and open examplesSynthetic data is the accelerator of the next phase of AI — what it is and what it means.
- Import AI 350: Neural architecture search at Facebook scale; hunting cancer with PANDA; European VCs launch a science labIf you were able to build a large-scale computer in the 18th century, how would things be different?
- Reshaping the tree: rebuilding organizations for AITechnological change brings organizational change.
- OpenAI: The Battle of the BoardPreviously: OpenAI: Facts from a Weekend.
- RLHF progress: Scaling DPO to 70B, DPO vs PPO update, Tülu 2, Zephyr-β, meaningful evaluation, data contaminationHuge steps forward in confirming that RLHF can really help you on vibes based evaluation, among many other RLHF analyses.
- Toward Better AI MilestonesHow to get off the treadmill of constantly shifting goalposts, and define tests that will tell us when AI capabilities get serious
- Import AI 349: Distributed training breaks AI policy; turning GPT4 bad for $245; better weather forecasting through AIHow inevitable is the creation of very powerful AI technology?
- Not much is changing, a lot is changingOpenAI, Microsoft, and the OpenOffspring
- OpenAI: Facts from a WeekendYou win, or we all die.
- OpenAI’s shakeup and opportunity for the rest of us: openness, brain drain, and new realitiesNew timelines that emerge in AI and the winners and losers, regardless of the unfolding details.
- The interface era of AIModern LLMs are becoming the easiest and most efficient way to access information. This will change how we see the world.
- Import AI 348: DeepMind defines AGI; the best free LLM is made in China; mind controlling robotsOver a long enough timescale, is the future of technology pre-determined, or something humans have real agency over?
- On OpenAI Dev DayOpenAI DevDay was this week.
- Reckoning with the Shoggoth of AICulture wars, open letters, new politics, developer days, and everything hidden under the smiling face of RLHF.
- Almost an Agent: What GPTs can doAlso, my book has a cover (also I have a book coming out)
- On the UK SummitIn the eyes of many, Biden’s Executive Order somewhat overshadowed the UK Summit.
- Import AI 347: NVIDIA speeds itself up with AI; AI policy is a political campaign; video morphing means reality collapseWorld leaders are gathering to talk about AI - what do they intuit? And is it different to what we intuit?
- On the Executive OrderOr: I read the executive order and its fact sheet, so you don’t have to.
- Open LLM company playbookWhere does releasing model weights fit into company strategy? 3 requirements, 3 actions, and 3 benefits of being in the open LLM space.
- Reactions to the Executive OrderPreviously: On the Executive Order
- Working with AI: Two paths to promptingDon't overcomplicate things
October 2023
- What the executive order means for openness in AIGood news on paper, but the devil is in the details
- Import AI 346: Human-like meta-learning; a 3 trillion token dataset; spies VS AIThe intelligence services have woken up to AI
- RLHF lit. review #1 and missing pieces in RLHFLooking at the difference between two sets -- what rumors say industry leaders are doing with RLHF and what the literature is up to. I'm starting my new series studying RLHF literature.
- Import AI 345: Facebook uses AI to mindread; MuJoCo v3; Amazon adds bipedal robots to its warehousesThe fact increasingly powerful AI systems exhibit increasingly humanlike qualities suggests that humans may be simpler than they appear
- The Best Available Human StandardWhat are the imperatives of the upside?
- How Transparent Are Foundation Model Developers?Introducing the Foundation Model Transparency Index
- Undoing RLHF and the brittleness of safe LLMsRecent papers show most of the arguments about needing "safety" in releases of open LLM weights are nearly dead in the water. Yes, still release the parameters.
- Import AI 344: Putting the world into a world model; automating software engineers; FlashDecodingMost technologies take decades to significantly alter the world. How long for language models?
- Neural Algorithmic ReasoningOn capturing classical computation with neural networks.
- What people ask me most. Also, some answers.A FAQ of sorts
- The AI research job market shit show (and my experience)There are plenty of jobs, but finding a place where you're happy is as hard as ever.
- Scale, schlep, and systemsThis startlingly fast progress in LLMs was driven both by scaling up LLMs and doing schlep to make usable systems out of them. We think scale and schlep will both improve rapidly.
- Import AI 343: Humanlike AI; LLaMa 2 protests; the NSA's new AI centerTo what extent is technological development an inevitability versus something people have agency over?
- The Artificiality of AlignmentOn the stakes of AI progress and claims about AI's existential risk.
- LLMs are computing platformsThis fact is why so many debates around LLMs feel broken, especially moderation.
- Open, general-purpose LLM companies might not be viableFailure modes on the quest to open-source LLMs. Expect pivots to specialized models.
- Evaluating LLMs is a minefieldAnnotated slides from a recent talk
- The shape of the shadow of The ThingWe can start to see, dimly, what the near future of AI looks like.
- Import AI 342: Mistral dumps an LLM on BitTorrent; AMD vs NVIDIA; Sutton joins keenAt some point, Hinton will shift from being a contemporary figure to a historical one
- An Introduction to the Problems of AI ConsciousnessOn the discussions around the concept of AI consciousness that are now taking center stage in the AI community.
September 2023
- DALL·E 3 and multimodality as moats, correcting bad moat takesMultimodality may be a key differentiator in how moats are built for LLMs.
- When Will AIs Acquire Insight?How Could We Even Train Them For That?
- Import AI 341: Neural nets can smell; technofeudalism via AI; China releases another solid open access modelWhat would happen to AI progress if an active learning technique that worked with transformers was published on arXiv tomorrow?
- Challenges operationalizing responsible AI in open RLHF researchSome reflections from my time at HuggingFace.
- Everyone is above averageIs AI a Leveler, King Maker, or Escalator?
- Midjourney vs. Ideogram, ML product companies, preventing AI winter, DALL·E 3 teaseThe coming image-generation battle and its implications on ML product longevity.
- The AI Explosion Might Never HappenAs AI Learns To Improve Itself, Improvements Will Become Harder To Find
- Import AI 340: Drone VS human (Drone wins); Adept's small but good AI model; "AI Succession"What's the difference between AI and nanotechnology in terms of hype and adoption?
- Centaurs and Cyborgs on the Jagged FrontierI think we have an answer on whether AIs will reshape work....
- In defense of the open LLM leaderboardNo, SemiAnalysis, HuggingFace isn't misleading all of open-source, and open-source is still making real progress.
- Intermediate SuperintelligenceGodlike AI, if it happens at all, won't happen overnight
- Text-to-CAD: Risks and OpportunitiesOn AI systems that accept a text prompt to produce a 3D CAD (computer-aided design) model
- AI researchers' challenges: atomic analogies and strained institutionsWe need to heal AI research norms, not build a super project. The reflections and reverberations that we're feeling in the AI community from Oppenheimer's quest for the atomic bomb.
- Embracing weirdness: What it means to use AI as a (writing) toolAI is strange. We need to learn to use it.
- Import AI 339: Open source AI culture war; Alibaba's multimodal model; the attacks (and defenses) made possible by generative AIIf you believed superintelligence could be built in your lifetime, what kind of life would you live?
- When We Forecast AGI, Do We Mean In The Lab Or In The Field?The Onset of AGI Will Play Out Over Decades
August 2023
- Cruise's collisions and adapting to AISF continues to be the center of attention for developments in AI, but this time it's in the physical world.
- Language models surprised usMost experts were surprised by progress in language models in 2022 and 2023. There may be more surprises ahead, so experts should register their forecasts now about 2024 and 2025.
- Import AI 338: Consciousness and AI; self-improving language models; maps of thought.All around us - normal, normal, normal! And yet some great belief within us that AI is progressing sufficiently quickly and powerfully that all around us soon will be - abnormal, abnormal, abnormal!
- Import AI 337: Why I am confused about AI; penguin dataset; and defending networks via RL with CYBERFORCEWill we look back on this decade as a crucial one, or as a normal one?
- Now is the time for grimoiresIt isn't data that will unlock AI, it is human expertise
- Does ChatGPT have a liberal bias?A new paper making this claim has many flaws. But the question merits research.
- The AI Progress ParadoxReceding Horizons and Misleading Milestones
- Introducing the REFORMS checklist for ML-based scienceML-based science is in trouble. Clear reporting standards for researchers could help.
- Summary of and Thoughts on the Hotz/Yudkowsky DebateGeorge Hotz and Eliezer Yudkowsky debated on YouTube for 90 minutes, with some small assists from moderator Dwarkesh Patel.
- Import AI 336: Financialized AI; public and elite AI opinion; one million insects.2023: The main provider of AI compute has a market capitalization of ~1 trillion dollars. In 2020, the same company had a market capitalization of ~150 billion dollars.
- Automating creativityThere is now strong evidence that AI can help make us more innovative.
- ML is useful for many things, but not for predicting scientific replicabilityHow the veneer of AI is used to legitimize awful ideas
- LLM products: measurement and manipulationTwo stories will begin to unfold as the AI capabilities-to-product overhang is reduced.
- Why We Won't Achieve AGI Until Memory Is A Core Architectural ComponentThe Hard Part of Thinking Is Knowing What To Think About
- In Praise of Boring AIAutomation has always been about killing tedious work. AI can do the same.
- Specifying objectives in RLHFAt ICML, it is obvious that many people are getting value out of RLHF. What is limiting the scientific understanding of it (other than research embargoes)?
July 2023
- Import AI 335: Synth data is a bad AI drug; Facebook changes the internet with LLaMa release; and Chinese researchers use AI to figure out chip designWhere do original ideas end and where does curve-fitting begin?
- "If it's not fully closed ML, it's open" - is it?Definitions from open-source software are being bent by new machine learning technologies.
- Llama We Doing This Again?I’ve finally had an opportunity to gather the available information about Llama-2 and take an in-depth look at the system card.
- Anthropic ObservationsDylan Matthews had in depth Vox profile of Anthropic, which I recommend reading in full if you have not yet done so.
- On holding back the strange AI tideThere is no way to stop the disruption. We need to channel it instead
- Llama 2 follow-up: too much RLHF, GPU sizing, technical detailsThe community reaction to Llama 2 and all of the things that I didn't get to in the first issue.
- Is GPT-4 getting worse over time?A new paper going viral has been widely misinterpreted
- Process vs. Product: Why We Are Not Yet On The Cusp Of AGIYou Can't Learn The Journey By Observing The Destination
- Llama 2: an incredible open LLMMeta is continuing to deliver high-quality research artifacts and not backing down from pressure against open source.
- How to Use AI to Do Stuff: An Opinionated GuideCovering the state of play as of Summer, 2023
- LLM agents and integration dead-endsWhen is GPT4 going to schedule my meetings? What is stopping it?
- We Need to Recognize How Profoundly Different The AGI Future Will BeWe are shaping AI; soon it will also be shaping us
- Interpretability CreationismOn “interpretability creationism” – interpretability methods that only look at the final state of the model and ignore its evolution over the course of training
- OpenAI Launches Superalignment TaskforceIn their announcement Introducing Superalignment, OpenAI committed 20% of secured compute and a new taskforce to solving the technical problem of aligning a superintelligence within four years.
- Consider Joining the UK Foundation Model TaskforceA few months ago, Ian Hogarth wrote the Financial Times Op-Ed headlined “We must slow down the race to God-like AI.”
- Import AI 334: Better distillation; the UK's AI taskforce; money and AICement manufacturing in the US has a combined revenue of around $10.4 billion annually, according to IBISWorld. What is the revenue in 2023 of the AI industry, I wonder?
- How to Make AI UX Your MoatDesign great AI Products that go beyond "just LLM Wrappers": make AI more present, more practical, and then more powerful.
- What AI can do with a toolbox... Getting started with Code Interpreter [Now called Advanced Data Analytics]Democratizing data analysis with AI
- The Homework ApocalypseFall is going to be very different this year. Educators need to be ready.
June 2023
- The Rise of the AI EngineerEmergent capabilities are creating an emerging title: to wield them, we'll have to go beyond the Prompt Engineer and write *software*, and AI that writes software.
- Tesla Autopilot's negligence and regulationMany of my colleagues in robotic learning have long been skeptical of Tesla’s efforts in the area. I also encourage people to not drive in Tesla’s with Autopilot engaged. Here’s why.
- Generative AI companies must publish transparency reportsThe debate about the harms of AI is happening in a data vacuum
- Import AI 333: Synthetic data makes models stupid; chatGPT eats MTurk. Inflection shows off a large language modelAs time goes on, more and more senior figures in AI are becoming concerned about the pace of deployments and the potential of the technology. Has the same happened in other rapidly scaling industries?
- Why transformative artificial intelligence is really, really hard to achieveA collection of the best technical, social, and economic arguments
- On giving AI eyes and earsAI can listen and see, with bigger implications than we might realize.
- How RLHF actually worksThe proven formula for RLHF and when we will see it in open-source.
- Three Ideas for Regulating Generative AIPolicy input to the federal government from a Stanford-Princeton team
- Is AI-generated disinformation a threat to democracy?An essay on the future of generative AI on social media
- Detecting the Secret CyborgsThe AI Trap for Organizations
- What Will AI Do For Us In The Near Term?The more time I spent writing this post, the more optimistic I became
- Different development paths of LLMsIn industry, open-source, and academia, each of these giant pools of talent are driven by different incentives and will create very different language models.
- Contra Marc Andreessen on AI"The claim that you will completely control any system you build is obviously false, and a hacker like Marc should know that"
- The Dial of Progress“There is a single light of science.
- Assigning AI: Seven Ways of Using AI in ClassAlso prompts! And things to watch out for!
- Import AI 332: Mini-AI; safety through evals; Facebook releases a RLHF dataset"I could be bounded in a nut shell and count myself a king of infinite space" is more aptly and literally applied to language models than to descriptions of human arrogance, ambition, and hubris.
- Why trying to "shape" AI innovation to protect workers is a bad ideaInstead, we should empower workers and create mechanisms for redistribution.
- Licensing is neither feasible nor effective for addressing AI risksNon-proliferation only benefits incumbents
- Open-source LLMs' harmlessness gapOpen-source LLM capabilities are leaving safety tools in the dust. What can we do?
- Could AI accelerate economic growth?Most new technologies don’t accelerate the pace of economic growth. But advanced AI might do this by massively increasing the research effort going into developing new technologies.
- Setting time on fire and the temptation of The ButtonWe used to consider writing an indication of time and effort spent on a task. That isn't true anymore.
May 2023
- Evaluating and uncovering open LLMsWhen choosing a model, we're stuck in the middle between classic NLP benchmarks (e.g. MMLU) and qualitative chatbot ranking. Neither are exactly what we want.
- Is Avoiding Extinction from AI Really an Urgent Priority?The history of technology suggests that the greatest risks come not from the tech, but from the people who control it
- Stages of SurvivalThis post outlines a fake framework for thinking about how we might navigate the future.
- The Crux ListIntroduction
- To Predict What Happens, Ask What HappensWhen predicting conditional probability of catastrophe from loss of human control over AGI, there are many distinct cruxes.
- Types and Degrees of AlignmentWhat would it mean to solve the alignment problem sufficiently to avoid catastrophe?
- To Address AI Risks, Draw Lessons From Climate ChangeWe can't solve the problem today; we *can* start establishing the conditions for it to be solved.
- Import AI 331: 16X smaller language models; could AMD compete with NVIDIA?; and BERT for the dark webThe symptoms of great revolution are always present in the mundane of the day-to-day
- Modern AI is DomestificationOn prior amplification, the art of taming internet-scale data distributions to make models useful and performant for specific tasks.
- What happens when AI reads a book 🤖📖And some prompts that might be useful when it does.
- Code: green pastures for LLMsBoth code-generation and code-training seem central to LLM progress, present and future.
- A Unified Theory of AI RiskNear-term risks aren't a distraction from extinction scenarios, they're the practice exam
- Import AI 330: Palantir's AI-War future; BLOOMChat; and more money for distributed AI trainingIf we perfectly simulated the entire world and trained agents within it, what might we learn about ourselves and the experience of being "human"?
- On-boarding your AI InternThere's a somewhat weird alien who wants to work for free for you. You should probably get started.
- Unfortunately, OpenAI and Google have moatsWhile everyone went crazy over a leaked memo, no one took the time to think through how companies have worked in the internet era.
- Import AI 329: Compute IS data; don't build AI agents; AI needs a precautionary principleAt the end of this phase of the tech revolution, what parts of it will seem like hyped inanity and what parts will seem like profound secrets?
- Catastrophe / EucatastropheWe have more agency over the future of AI than we think.
- AI is not good software. It is pretty good people.A pragmatic approach to thinking about AI
- Import AI 328: Cheaper StableDiffusion; sim2soccer; AI refinement...as AI industrializes, we should ask ourselves where the key inputs are and when they'll be in contention
- Get Ready For AI To Outdo Us At EverythingWe have time to prepare; step one is acknowledging what we're preparing for
- Specifying hallucinationsThe future of LLM errors and the need for a discourse around the specification of unexpected outputs.
- It is starting to get strange.Let's talk about ChatGPT with Code Interpreter & Microsoft Copilot
- Import AI 327: Stable Diffusion on phones; GPT-Hacker; UK launches a £100m AI taskforce2024 will be the year of Reality Collapse Elections (RCEs)
- The costs of cautionIf you thought we might be able to cure cancer in 2200, then I think you ought to expect there’s a good chance we can do it within years of the advent of AI systems that can do the research work humans can do.
- The Rocket Alignment Problem, Part 2Previously (Eliezer Yudkowsky): The Rocket Alignment Problem.
April 2023
- Beyond the Turing TestThe key question about AI is no longer "can it think", but rather "can it hold down a job"?
- In-Context Learning, In Context + Author Q&AsOn the phenomenon of in-context learning in large language models and what researchers have learned about it so far. Plus, Q&As with Hattie Zhou and Sewon Min.
- A guide to prompting AI (for what it is worth)A little bit of magic, but mostly just practice
- Beyond human data: RLAIF needs a rebrandReinforcement learning from computational feedback: how a promising method is being ignored because of a confusing launch.
- It's Time To Build AI | UXBridging the Capability Overhang from Generative AI to Generative UI
- Quantifying ChatGPT’s gender biasBenchmarks allow us to dig deeper into what causes biases and what can be done about it
- Transcript and Brief Response to Twitter Conversation Between Yann LeCun and Eliezer YudkowskyYann LeCun is Chief AI Scientist at Meta.
- An Intuitive Explanation of Large Language ModelsInstead of diving into the math, I explain *why* they're built as "predict the next word" engines, and present a theory for why they make conceptual errors.
- Notes on Potential Future AI Tax PolicyResponse To: Via Marginal Revolution, Brian Slesinsky proposes a tax on language model API calls.
- A Hypothetical Takeover Scenario Twitter PollI ran an experimental poll sequence on Twitter a few weeks back, on the subject of what would happen if a hostile ASI (Artificial Superintelligence) tried to take over the world and kill all humans, using only ordinary known techs.
- Import AI 326:Chinese AI regulations; Stability's new LMs If AI is fashionable in 2023, then what will be fashionable in 2024?What would you do at the end of time?
- Software²A new generation of AIs that become increasingly general by producing their own training data
- Democratizing the future of educationWe are all EdTech designers, now
- I’m a Senior Software Engineer. What Will It Take For An AI To Do My Job?Exploring the gaps between current LLMs and true general intelligence
- The Anatomy of Autonomy: Why Agents are the next AI Killer App after ChatGPTAuto-GPT/BabyAGI Executive Summary, a Brief History of Autonomous Agentic AI, and Predictions for Autonomous Future
- The Overemployed Via ChatGPTPreviously: Escape Velocity From Bullshit Jobs
- Import AI 325: Automated mad science; AI vs democracy; and a 12B parameter language modelPlus: Why prompt injection could be a big problem
- One sentence.Prompting for maximum impact (and why that is a bad idea)
- Grounding Large Language Models in a Cognitive FoundationHow to Build Someone We Can Talk To
- I set up a ChatGPT voice interface for my 3-year old. Here’s how it went.Chatbots are likely to revive familiar debates about kids and apps
- GPT-4 Doesn't Figure Things Out, It Already Knows ThemYes, It's a Stochastic Parrot, But Most Of The Time You Are Too, And It's Memorized a Lot More More Than You Have
- On AutoGPTThe primary talk of the AI world recently is about AI agents.
- Continuous doesn’t mean slowOnce a lab trains AI that can fully replace its human employees, it will be able to multiply its workforce 100,000x. If these AIs do AI research, they could develop vastly superhuman systems in under a year.
- Growing needs for accessing state-of-the-art reward modelsThe rewards assigned by these models will control the subjective experience of users; the subjective experience controls their decision-making.
- Import AI 324: Machiavellian AIs; LLMs and political campaigns; Facebook makes an excellent segmentation model...and a story about generative AI and games and how souls work...
- Nobody knows how many jobs will "be automated"Whatever that even means.
- The future of education in a world of AIA positive vision for the transformation to come
- Behind the curtain: what it feels like to work in AI right now (April 2023)Fear, FOMO, and the scientific exodus driven by ChatGPT
- Eliezer Yudkowsky's Letter in Time MagazineFLI put out an open letter, calling for a 6 month pause in training models more powerful than GPT-4, followed by additional precautionary steps.
- Thinking companion, companion for thinkingSome simple ways to use AI to break you out of biases
- AIs accelerating AI researchResearchers could potentially design the next generation of ML models more quickly by delegating some work to existing models, creating a feedback loop of ever-accelerating progress.
- Import AI 323: AI researcher warns about AI; BloombergGPT; and an open source Flamingo...Plus, a story about the financial crisis and AI research...
March 2023
- You Are Not Too Old (To Pivot Into AI)Everything important in AI happened in the last 5 years, and you can catch up
- Is it time for a pause?The single most important thing we can do is to pause when the next model we train would be powerful enough to obsolete humans entirely. If it were up to me, I would slow down AI development starting now — and then later slow down even more.
- On the FLI AI-Risk Open LetterThe Future of Life Institute (FLI) recently put out an open letter, calling on all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4.
- A misleading open letter about sci-fi AI dangers ignores the real risksMisinformation, labor impact, and safety are all risks. But not in the way the letter implies.
- How to use AI to do practical stuff: A new guidePeople often ask me how to use AI. Here's an overview with lots of links.
- Response to Tyler Cowen's Existential risk, AI, and the inevitable turn in human historyPredictions are hard, especially about the future.
- The implicit dynamics of optimizing costs vs. rewards vs. preferencesWith the emergence of reinforcement learning from human feedback, we've been applying old techniques with a new guiding function (🤫 RLHF).
- GPT-4 Plugs InGPT-4 Right Out of the Box
- Import AI 322: Huawei's trillion parameter model; AI systems as moral patients; parasocial bots via Character.ai...and a story about the market value of grief to (mostly) immortal aliens...
- "Aligned" shouldn't be a synonym for "good"Perfect alignment just means that AI systems won’t want to deliberately disregard their designers' intent; it's not enough to ensure AI is good for the world.
- Alignment researchers disagree a lotMany fellow alignment researchers may be operating under radically different assumptions from you.
- The ethics of AI red-teamingIf we’ve decided we’re collectively fine with unleashing millions of spam bots, then the least we can do is actually study what they can – and can’t – do.
- Situational awarenessAI systems that have a precise understanding of how they’ll be evaluated and what behavior we want them to display will earn more reward than AI systems that don’t.
- Playing the training gameWe're creating incentives for AI systems to make their behavior look as desirable as possible, while intentionally disregarding human intent when that conflicts with maximizing reward.
- Training AIs to help us align AIsIf we can accurately recognize good performance on alignment, we could elicit lots of useful alignment work from our models, even if they're playing the training game.
- Superhuman: What can AI do in 30 minutes?AI multiplies your efforts. I found out by how much...
- Microsoft Research Paper Claims Sparks of Artificial Intelligence in GPT-4Microsoft Research (conflict of interest?
- Acceleration.7 days of new AI technologies shows us that everything is happening very fast.
- OpenAI’s policies hinder reproducible research on language modelsLLMs have become privately-controlled research infrastructure
- GPT-4 and professional benchmarks: the wrong answer to the wrong questionOpenAI may have tested on the training data. Besides, human benchmarks are meaningless for bots.
- GPT4: The quiet parts and the state of MLChecking in on the state of the ML field: technical progress, societal implications, and more.
- Import AI 321: Open source GPT3; giving away democracy to AGI companies; GPT-4 is a political artifactand a story about automating our humanity via AI...
- Using AI to make teaching easier & more impactfulHere are five strategies and prompts that work for GPT-3.5 & GPT-4
- The Multi-modal, Multi-model, Multi-everything Future of AGIGPT-4 FOMO antidote, and meditations on Moravec's Paradox
- What is algorithmic amplification and why should we care?A symposium and a primer on social media recommendation algorithms
- How to... use AI to unstick yourselfWe often lose momentum because of something small. AI can help.
- Import AI 320: Facebook's AI Lab Leak; open source ChatGPT clone; Google makes a universal translator.Plus, a story about moral patienthood in the age of machine minds.
- AGI Roundup: Re-visiting Go; transitioning from narrow to general; multimodality & GPT4AI is much more fun to think about when AGI is an adjective, not a finish line.
- Artists can now opt out of generative AI. It’s not enough.Opting out is the latest example of generative AI developers externalizing costs.
- LLMs are not going to destroy the human raceIt's just a chatbot, dude.
- Secret Cyborgs: The Present Disruption in Three PapersThe future is already here, we just need to figure out a few details.
- Import AI 319: Sovereign AI; Facebook's weights leak on torrent networks; Google might have made a better optimizer than Adam!...And a story about AI race dynamics
- The LLaMA is out of the bag. Should we expect a tidal wave of disinformation?The bottleneck isn't the cost of producing disinfo, which is already very low.
- Feats to astonish and amazeA compendium of things I didn't think AI should be able to do
- Power and Weirdness: How to Use Bing AIBing AI is a huge leap over ChatGPT, but you have to learn its quirks
- AI cannot predict the future. But companies keep trying (and failing).A new paper on how AI companies make false promises and how we can challenge them
- AI: Practical Advice for the WorriedSome people (although very far from all people) are worried that AI will wipe out all value in the universe.
February 2023
- The RLHF battle lines are drawnTales of the open and closed sides, how these two dynamics will dictate progress and public perception.
- AI Techies!Everything you didn't realize you wanted to know about the denizens of Cerebral Valley
- How to Get an AI to Lie to You in Three Simple StepsI keep getting fooled by AI, and it seems like others are, too.
- Blinded by AnalogiesWhat is this AI thing? The wrong model can lead us astray
- People keep anthropomorphizing AI. Here’s whyCompanies and journalists both contribute to the confusion
- "AI alignment" and uncalibrated discourse on AIRe-naming arms races; bridging communities; aligning ChatGPT, and more.
- The future, soon: what I learned from Bing's AIWe had a brief glimpse of two different types of AI. Both are significant
- My class required AI. Here's what I've learned so far.(Spoiler alert: it has been very successful, but there are some lessons to be learned)
- I hope you weren't getting too comfortable.I just got access to the new Bing AI. My initial thoughts are that our assumptions about the limits of AI were wrong.
- Three seasons of RL: Metaphor, tool, and frameworkHow we’re passing through different eras of reinforcement learning.
- A quick and sobering guide to cloning yourselfIt took me a few minutes to create a fake me giving a fake lecture.
- Magic for English MajorsProgramming in prose in an AI-haunted world
- "Do not fear AI, puny humans... that is not meant as a threat."What we can learn from a completely AI written & illustrated lecture
- Shameless #Foomerism and S-CurvesWe need to talk about how we talk about AI progress
- Scaling laws for robotics & RL: Not quite yetRobotics Transformers, DreamerV3, XLand 2, and hoping that scaling laws are coming embodied AI.
- The Machines of Mastery"Anyone can learn anything they want..." and how technology can help
January 2023
- A prosthesis for imagination: Using AI to boost your creativityAI can already beat humans in many measures of creativity. Let's use that to our advantage.
- The practical guide to using AI to do stuffA resource for students in my classes (and other interested people).
- Reasons to Punish Autonomous RobotsZac Cogley considers the design of robots that would be responsive to human criticism and blame.
- Do Large Language Models learn world models or just surface statistics?Kenneth Li describes evidence suggesting world models.
- Becoming strange in the Long SingularityA sudden increase in AI capability suggests a weirder world in our near future
- All my classes suddenly became AI classesWe can't beat AI, but it doesn't need to beat us (or our students)
- Secretary jobs in the age of AIA guest post by Hollis Robbins
- Pretraining quadrupeds: a case study in RL as an engineering toolHow an unlikely corner of robotics research, locomotion, defined RL's new notion of success.
- Every Google vs OpenAI Argument, DissectedLet's talk the AI elephant in the room. Bonus: Meditations on Moloch
- How to... use ChatGPT to boost your writingThe key to using generative AI successfully is prompt-crafting
- And the great gears begin to turn again...Progress in two vital areas had slowed, AI could change that
- Looking into 2023My predictions for machine learning this year: 3D assets, self-driving, GPT4, RLHF, Deep RL, diffusion models, and conference cycles.
- You can't brute force the unsolvableBut a little finesse can help
- The third magicA meditation on history, science, and AI
2022
December 2022
- Predicting machine learning moatsModels aren't moats and how emergent behavior scaling laws will change the business landscape.
- Reverse Prompt Engineering for Fun and (no) ProfitPwning the source prompts of Notion AI, 7 techniques for Reverse Prompt Engineering... and why everyone is *wrong* about prompt injection
- Who believes more myths about humans: AI or educated humans?ChatGPT should do badly where human information is most wrong. But is that true?
- The street finds its own uses for things, AI EditionBreakthrough innovations often come from people who use technology, not create it. So, how are students using AI?
- Closed-API vs Open-source continues: RLHF, ChatGPT, data moatsModel-as-a-service makes more sense when there is a data advantage to back it up.
- What Building "Copilot for X" Really TakesA "Copilot for X" guide from the team that built the first real Copilot competitor!
- ChatGPT is my co-founderCommon barriers hold entrepreneurs back, AI can help founders cross them
- How to... use AI to teach some of the hardest skillsWhen errors, inaccuracies, and inconsistencies are actually very useful
- Four Paths to the RevelationGive me 10 minutes, and I think I can make you obsess over AI
- ChatGPT is a bullshit generator. But it can still be amazingly usefulThe philosopher Harry Frankfurt defined bullshit as speech that is intended to persuade without regard for the truth. By this measure, OpenAI’s new chatbot ChatGPT is the greatest bullshitter ever. Large Language Models (LLMs) are trained to produce
- The Mechanical ProfessorI take a job I know well, and try to see how far I can automate it with AI.
- Learning to Make the Right Mistakes - a Brief Comparison Between Human Perception and Multimodal LMsMayukh Deb discusses the "right mistakes."
- RLHF, 'online' ML systems, and RL going mainstreamCommon machine learning systems are starting to deploy the RL lens of feedback.
- The Day The AGI Was BornChatGPT FOMO Antidote: I read all the tweets so you don't have to
- How to... use AI to generate ideasAI can make you more creative right now, and here are some ways to do it...
- Generative AI: autocomplete for everythingA joint blog post by Noah and roon on the future of work in the age of AI
November 2022
- AI has a strategy.And we can learn a lot from how it is going to win.
- Why "Prompt Engineering" and "Generative AI" are overhypedHow Stable Diffusion 2.0 and Meta's Galactica demonstrate the two heresies of AI
- The bait and switch behind AI risk prediction toolsToronto recently used an AI tool to predict when a public beach will be safe. It went horribly awry. The developer claimed the tool achieved over 90% accuracy in predicting when beaches would be safe to swim in. But the tool did much worse: on a majority of the days when the water was in fact…
October 2022
- What is AGI-hardAt this point, knowing what AI can't do is more useful than knowing what it can
- Using RL's exploitation to debugA musing on how I think autonomous system companies should use RL.
- Students are acing their homework by turning in machine-generated essays. Good.Teachers adapted to the calculator. They can certainly adapt to language models.
- How Open Source is eating AIWhy the best delivery mechanism for AI is not APIs
September 2022
- Eighteen pitfalls to beware of in AI journalismA checklist for avoiding hype
- Artificial Intelligence and the Future of DemosSalla Ponkala's piece on AI and democracy.
- Back in the gameWhat I've been up to and what's coming soon.
- Eigenquestions for the AI Red WeddingWhat the next great AI-enabled product will answer
- Lovecraftian intelligenceIs AI an incomprehensible cosmic horror?
- Multiverse, not MetaverseGenerative AI lets us explore Many worlds owned by Nobodies, and this is fundamentally better than One world owned by Somebody
- Generative AI models generate AI hypeOver 90% of images of AI produced by a popular image generation tool contain humanoids
- American workers need lots and lots of robotsWith the power of automation, our workers can win. Without it, they're in trouble.
- Causal Inference: Connecting Data and RealityWenwen Ding's piece on causal inference.
August 2022
- Why are deep learning technologists so overconfident?Deep learning researchers have proved skeptics wrong before, but the past doesn't predict the future.
- The Future of Speech Recognition: Where Will We Be in 2030?Migüel Jetté's and Corey Miller's piece on the future of ASR
- Symmetries, Scaffolds, and a New Era of Scientific DiscoveryThe Challenge of Drug Discovery
July 2022
- Overview of Graph Theory and Alzheimer's DiseaseGraph theory may provide powerful means of understanding brain function at global and local levels, as well as network changes in the brain that occur due to disease.
June 2022
- Lessons from the GPT-4Chan ControversyWhat happened with GPT-4chan, aka 'the worst AI ever', and what can be learned from it
- AI is Ushering In a New Scientific RevolutionAI is ushering in a new scientific revolution by making remarkable breakthroughs in a number of fields, unlocking new approaches to science, and accelerating the pace of science and innovation.
May 2022
- Working on the Weekends - an Academic Necessity?If there is one thing that seems more typical of academia than having a terrible work-life balance, it is complaining about the terrible work-life balance we have. So what do we do about it?
- Lessons From Deploying Deep Learning To ProductionPeter Gao, an early engineer at Cruise, reflects on his experience deploying deep learning models into production.
- Beyond Message Passing: a Physics-Inspired Paradigm for Graph Neural NetworksOn going beyond message-passing based graph neural networks with physics-inspired “continuous” learning models
- Deep Learning in NeuroimagingAn introduction to unique aspects of neuroimaging data and how we can leverage these aspects with deep learning algorithms.
April 2022
- AI Startups and the Hunt for Tech Talent in VietnamAs recruitment for AI talent becomes more and more competitive, how can countries with developing tech sectors keep up?
- New Technology, Old Problems: The Missing Voices in Natural Language ProcessingWho and what is being represented in data and development of NLP models, and how unequal representation leads to unequal allocation of the benefits of NLP technology
February 2022
- One Voice Detector to Rule Them AllVAD is among the most important and fundamental algorithms in any production or data preparation pipelines related to speech
- How Aristotle is Fixing Deep Learning's FlawsA summary of some key points of Aristotelian logic and epistemology and how their absence in traditional deep learning is responsible for deep learning’s well-known limitations
- How AI is Changing Chemical DiscoveryThis article covers some of the more prominent usages of AI in chemistry, the parent discipline of the protein folding problem
- Designing Societally Beneficial Reinforcement Learning SystemsChoices, risks, and reward reporting. Recommendations for how to integrate RL systems with society.
January 2022
- Engaging with DisengagementHow will traffic laws change as we slowly enter the autonomous vehicle era, and in general, the AI-driven 21st century?
- Flexible Centralization in Multi-agent Learning & ControlA tour of control theory, multi-agent RL, and hierarchical learning.
- Contra David Deutsch on AI"The universal explainers hypothesis is incorrect because it cannot explain why some animals are smart and some humans are stupid."
- A Science Journalist’s Journey to Understand AI7 lessons author Kathryn Hulick learned while writing her educational books about AI and other scientific topics
2021
December 2021
- AI and the Future of Work: What We Know TodayOne of the most important issues in contemporary societies is the impact of automation on human work. With AI becoming ever more advanced, what will its impact on automation be?
- New Datasets to Democratize Speech Recognition TechnologyPresenting the The People’s Speech, a massive English-language dataset of audio transcriptions, and the Multilingual Spoken Words Corpus (MSWC), a 50-language, 6000-hour dataset of individual words
- How to Train your Decision-Making AIsHow do humans transfer their knowledge and skills to artificial decision-making agents more efficiently? What kind of knowledge and skills should humans provide and in what format?
November 2021
- Explain Yourself - A Primer on ML Interpretability & ExplainabilityOn the intersection of causality and interpretability, exciting and promising research areas in the field of Machine Learning
October 2021
- Strong AI Requires Autonomous Building of Composable ModelsOn why strong AI must be able to compose these models dynamically to generate combinatorial representations
- Reflections on Foundation ModelsThe authors of the recent foundation models report reflect on recent discussion.
September 2021
- The Imperative for Sustainable AI SystemsThere are immediate next steps that we can take to improve the environmental posture of the AI systems that we build.
- Has AI found a new Foundation?Over a hundred researchers at Stanford think so. We are not so sure.
August 2021
- An Introduction to AI Story GenerationA primer on automated story generation and how it it strikes at some fundamental research questions in artificial intelligence.
- Systems for Machine LearningOn the field of Machine Learning Systems and how it addresses the new challenges of ML with a lens shaped by traditional systems research
- Remote robotic-data farmsIndustry labs power up their robot learning research with parallelization!
- Machine Learning Won't Solve Natural Language UnderstandingAn argument that it is time to re-think our approach to natural language understanding, since the ‘big data’ approach to NLU is implausible and flawed..
- on the Horizon of applied RLHow RL is starting to be used by industry and how RL is heading to a framing more suited for industrial scales.
July 2021
- Machine Translation Shifts PowerThe creation, development, and deployment of machine translation technology is historically entangled with practices of surveillance and governance.
- It’s All Training Data: Using Lessons from Machine Learning to Retrain Your MindAs machine learning scientists, we know that training data can make or break your model.
- Justitia ex Machina: The Case for Automating MoralsMachine learning models are ubiquitous in our lives, and it is important to understand how they perform in the light of being fair, transparent, and just in their decision-making processes.
- Prompting: Better Ways of Using Language Models for NLP TasksA review of recent advances in prompting in large language models.
June 2021
- How to Do Multi-Task Learning IntelligentlyNew methods that automatically learn what to learn together
- Reward is not enoughMulti-agent scenarios make reward maximization a risk. Discussing when, rather than if, we should believe in the Reward Hypothesis.
- Are Self-Driving Cars Really Safer Than Human Drivers?More than $250 billion has been invested in the last 3 years. But there are good reasons to believe that autonomous vehicles may not be capable of handling unexpected situations safely as humans do.
- How all machine learning becomes reinforcement learningI make the case why people iteratively training any model should learn some core concerns of reinforcement learning.
May 2021
- How has AI contributed to dealing with the COVID-19 pandemic?Numerous research groups have helped in the quest to contain the pandemic. This piece reviews where these efforts have been realized as practical solutions.
- Towards Human-Centered Explainable AI: the journey so farIntroducing Human-centered Explainable AI (HCXAI) as an approach that puts the human at the center of technology design for AI
- Machine Learning, Ethics, and Open Source Licensing (Part II)Part II/II: What Open Source Licensing Can Bring to ML Ethics
April 2021
- Machine Learning, Ethics, and Open Source Licensing (Part I)Part I/II: The Use, Misuse, and Regulation of Machine Learning Systems
- Attention in the Human Brain and Its Applications in MLThe roots of attention in machine learning comes from the human visual system.
- Decentralized AI For HealthcareMachine learning for healthcare while still retaining privacy via federated learning.
March 2021
- Setting ourselves up for exploitation: RL in the wildHow simulator exploitation, a dual of over-optimization, in AI is the canary in the coal mine for what negative implications could come from weakly-bounded, data-driven iterative systems (RL).
February 2021
- Counting down until consumer drones are banned in citiesI don’t even like the idea of flying delivery drones, but that will be all we have.
- Clarifying RL: Obscure problem formulations and structure tradeoffsSome debates that will be settled en route to RL being used in all corners of the modern world.
- Catching Cyberbullies with Neural NetworksIn this article you can play with techniques to automatically detect users that misbehave, preferably as early in the conversation as possible.
- Decoupling AI from the latent variable of spoken languagesHow English being the language of progress in AI could bias our machine’s minds and what we think they are capable of. Uncoupling AI from our notion of language is frightening.
- 100th anniversary of the word robot: COVID didn’t give us personal robots, it gave us WoebotYes COVID accelerated automated manufacturing and logistics, but robots have not been helping out en masse anywhere else.
January 2021
- Can AI Let Justice Be Done?Will “Robotic judges that can determine guilt will be ‘commonplace’ within 50 years”? Should that be the case?
- Boston Dynamics 🤖🐈: Studying Athletic IntelligenceThe acrobatic dance videos are flashy, but what are the actual technical breakthroughs? What is happening to the Korean robotics industry?
- Robotic Companies 2.0: Horizontal ModularityHow behaving as a platform rather than a manufacturer will be the sweet spot for the next generation of robotics companies.
- A Visual History of Interpretation for Image RecognitionHow have interpretation methods for image recognition developed over the past few years? Our latest piece provides a fascinating visual history.
- Reflections on digital technology from a capitol siegeThe digital fallout following the siege on the capital shows we are not keeping up with the pace of technological change.
- Considering AIs through our own mind’s reflectionLessons about AI from lessons about our mind. Focused on the nature of free will.
- Knocking on Turing’s door: Quantum Computing and Machine LearningQuantum computing promises exponential compute power that we have yet to attain due to decoherence and cryogenic requirements, and is in a fascinating place. How does it intersect with ML?
- The stakeholders start to learn the stakes: predictions for machine learning and automation in 2021Most of my predictions for automation in 2021 have common people starting to learn how their lives will be impacted, and it’s up to the practitioners to make it fair.
2020
December 2020
- When BERT Plays The Lottery, All Tickets Are WinningDoes the Lottery Ticket Hypothesis hold for BERT? Experiments with structured and magnitude pruning reveal some good and some bad news.
- The Ubiquity and Future of Model-based Reinforcement LearningMaking a case for why you want to learn about models and why my research matters.
- Facebook Case Study 🌏🔬: Using AI to Regulate the Digital EcosystemThe modern digital giants have data on scales unfathomable to the outside, so trying to understand how they manage deleterious information is an uphill battle.
- The Far-Reaching Impact of Dr. Timnit GebruIn light of Dr. Timnit Gebru’s recent firing from Google, Rachel Thomas gives an overview of her most important contributions to the machine learning community and beyond.
November 2020
- Interpretability in ML: A Broad OverviewInterested in contributing? Pitch article ideas with this form | Forward to a friend Interpretability in ML: A Broad Overview This essay provides a broad overview of the sub-field of machine learning interpretability While this post is not exhaustive, my goal is to review conceptual frameworks,…
- How Can We Improve Peer Review in NLP?Interested in contributing? Pitch article ideas to The Gradient by filling out this short form How Can We Improve Peer Review in NLP? Have your conference or journal reviewers ever completely misunderstood your paper or raised unreasonable objections?
- The Collingridge Dilemma and Current Policy on RobotsLegal cases impacting robotics, innovation vs policy, the sci-fi future of robotics.
- Don’t Forget About Associative MemoriesInterested in contributing?
- Constructing Axes for Reinforcement Learning PolicyA small step into the research community’s most opaque framework.
October 2020
- Why skin lesions are peanuts and brain tumors harder nutsInterested in contributing?
- Towards an Ethics for RoboticistsHow does one handle abstraction and scale in ethical design?
- What is ML for Climate Change?A new “subfield” founded in 2019 is making waves, and is more accessible than I first thought.
- The Gap: Where Machine Learning Education Falls ShortLike The Gradient?
- How the Police Use AI to Track and Identify YouInterested in contributing?
- Models, Systems, Code; and RobotsEE v ME v CS in complex systems. Do we view things differently?
September 2020
- AI Democratization in the Era of GPT-3Like The Gradient?
- Transformers are Graph Neural NetworksInterested in contributing?
- Autonomy startups are such a messAI startups just use data logging and robotics startups are all in dark mode: where are we heading? I just want to hear what people are really doing, and then I can decide its value.
- AI & Arbitration of TruthCan we make an AI fact checker? Language, knowledge, and opinions are all moving targets.
August 2020
- Can robots be autonomous and unintelligent?Re-thinking robot design part 2: don’t restrict your robot worldview to robot (a) does task (b).
- The uncanny world of robots at homeRe-thinking robotic design part 1: social support robots diving us into the uncanny valley. We don't need robots to look like humans.
- Digital companies and the goal of at-home embodied AILearning how to learn, by letting autonomous agents interact with the world. Why big tech companies like Facebook and Google hire roboticists to bring life to their excess meeting rooms.
July 2020
- Automated: how algorithms shape a day in the life and our futureThe goal of the algorithm is not to give users content they like, but to make the users more predictable. Brief added comment on online courses.
- Shortcuts: Neural Networks Love to CheatOverviews on cutting-edge AI research and perspectives on the direction of the field.
- Recommendations are a game - a dangerous game (for us).Machine-recommendation system ethics, models of human reward, reinforcement learning framing; followup on GPT-3.
- How to Stop Worrying About CompositionalityInterested in contributing?
- Automating code, Twitter's hack(s), a robot named StretchLanguage models write code, Twitter gets hacked (again), new robots, and top-tier conferences.
- Online courses, automating education, and digitalizing degreesOnline teaching is here to stay and I don’t think anyone knows how it will go.
- Challenges of Comparing Human and Machine PerceptionInterested in contributing?
- Democratizing AutomationGetting everyone to benefit from the artificial intelligence boom could be more challenging than some expect.
June 2020
- Drones, Swarms, and Storms of DronesThere is no situation when I want to encounter a drone of unknown origin, and in the future we'll be seeing hundreds.
- Lessons from the PULSE Model and DiscussionInterested in contributing?
- "10 years of automation in 1 year"What the automation acceleration from COVID19 actually will look like. What did robotics look like a decade ago and is this step feasible?