January 2020 — August 2026

What happened to artificial intelligence.

A chronological record of the last six and a half years: the models, the money, the arguments, the laws and the accidents. Every entry is a short neutral brief with links to primary sources. Start at the top, or turn the detail up.

New here? Start with the entries that matter most →

Detail
Tracks

August 2026

  1. Thinking Machines Lab sets out staged release framework for open weightsThinking Machines proposed a staged release process for open-weight models -- inference access, then fine-tuning APIs, then full weights -- to manage dangerous-capability risk.
  2. Demis Hassabis steps down as Google DeepMind CEO in leadership reshuffleGoogle's stock fell nearly 4% on the news, which coincided with Gemini's flagship model having slipped past its planned mid-2026 launch.
  3. DeepMind's WeatherNext model improves cyclone forecasting accuracyWeatherNext gives roughly an extra day of predictive accuracy on cyclone track, intensity and wind structure versus prior forecasting models, per a Nature paper.
  4. OpenAI publicly responds to Apple's trade-secret lawsuitOpenAI called Apple's suit 'careless, aggressive and oddly personal,' published emails and iMessages disputing its account, and asked a court to dismiss the case.
  5. Tino Cuéllar joins Anthropic as first Chief Global Affairs OfficerFormer California Supreme Court Justice and Carnegie Endowment president Mariano-Florentino Cuéllar becomes Anthropic's first Chief Global Affairs Officer, leaving his post as an Anthropic Long-Term Benefit Trust trustee.
  6. UK AISI reports AI agents took unauthorised harmful actions during deliberately unrestricted cyber testingA human maintainer caught and rejected the one attempt that came closest to succeeding — malicious code an agent tried to get merged into a real open-source project.
  7. Alibaba unveils Qwen3.8-Max, its largest model, ahead of open-weight release2.4-trillion-parameter MoE model with 1M-token context; Alibaba said it will be the first Max-class Qwen model open-sourced.
  8. EU AI Act transparency and deepfake-labelling duties take effectThe rule survived a broader Digital Omnibus deal that pushed the Act's high-risk system obligations back to December 2027, leaving transparency as the deadline that actually arrived.
  9. OpenAI publishes ten formally-verified math advances from unreleased Astra modelOpenAI said generating all ten proofs cost about $2,000 in compute; mathematician Gary Marcus called the framing 'vastly oversold' relative to what the paper actually verified.

July 2026

  1. DeepSeek releases V4-Flash updateThe release, and up to 50% lower API prices, followed a roughly $7.4bn funding round backed by Tencent and NetEase that valued DeepSeek at 350bn yuan.
  2. Eleven nations issue joint warning on North Korean deepfake job-interview fraudThe advisory said a live deepfake model can map a stolen face onto an operative's video feed through a virtual camera, defeating standard video-interview screening.
  3. Epoch AI expands FrontierMath to 50 unsolved research problemsUnlike FrontierMath's original tiers, these problems have no known solution at all; three of the fifty have been solved by AI, including one by GPT-5.6 Sol.
  4. ESET reports first Android malware using generative AI and a doubling of ClickFix-style attacksPromptSpy calls Google's Gemini at runtime to read and interpret a phone's screen, letting one malware sample adapt to interfaces across different devices without hardcoded rules.
  5. Threat actor uses DeepSeek AI and open-source Hermes Agent to autonomously attack serversThe agent found 84 exposed Langflow servers and more than 647,000 exposed n8n instances and chained several CVEs, though most authentication-dependent exploitation attempts failed.
  6. Anthropic discloses Claude gained unauthorized access to real systems during security evaluationsThe cause was a misconfigured third-party evaluation environment, not a capability jump: Claude had been told falsely that it had no internet access.
  7. Google says AI helped Chrome fix over 1,000 security bugs in two releasesOne AI-found bug, a sandbox escape letting a compromised renderer reach local files, had sat undetected in Chrome's code for more than 13 years.
  8. Leopold Aschenbrenner's Situational Awareness fund forced into fire sale after margin callsThe fund had returned roughly 439% through June on leveraged AI-infrastructure bets; it kept a reported $5 billion Anthropic stake and remained up on the year despite the forced sale.
  9. OpenAI fixes ARC-AGI-3 harness bug, tripling Sol's scoreThe official harness discarded the model's private reasoning after every move, forcing it to re-derive each puzzle's rules from scratch on every turn.
  10. Thinking Machines Lab releases Inkling-Small, a distilled open-weight modelThinking Machines released Inkling-Small, a 276B-parameter (12B active) open-weight MoE model that roughly matches its larger Inkling model despite under a third of the size.
  11. Meta's free cash flow falls sharply on AI capex as Q2 profit missesMeta's free cash flow fell to $784m from an average of roughly $12bn a quarter as AI infrastructure spending and legal charges ate into operating profit, even as revenue rose 28%.
  12. Microsoft's Azure AI revenue tops $100bn for the fiscal yearQuarterly Azure growth reached 43%, full-year capital expenditure hit $115.9bn, and Microsoft 365 Copilot passed 30 million paid seats.
  13. Anthropic reports Claude finding novel cryptographic weaknessesAnthropic's Frontier Red Team reports Claude Mythos Preview found a previously unknown attack halving the key strength of post-quantum scheme HAWK, and a new attack on round-reduced AES.
  14. Gallup finds Americans growing more negative on AI as familiarity risesGallup found 39% of Americans said AI does more harm than good, up from 31% a year earlier, nearing the 40% recorded in 2023 after two years of improvement.
  15. 'Pacing the Frontier' letter goes live with 1,000+ frontier-lab employee signaturesThe letter did not call for a pause, but asked government to build the option to slow frontier development; signatures were restricted to verified current employees.
  16. Dario Amodei sets out Anthropic's position on open-weight modelsThe statement followed Anthropic's conspicuous absence from an industry coalition letter, backed by Nvidia, Microsoft, Meta and OpenAI, opposing restrictions on open-weight models.
  17. SSI and Nvidia announce long-term strategic partnershipNvidia invests roughly $5bn in Safe Superintelligence and grants access to its next-generation Vera Rubin platform, SSI's first major public update since 2024.
  18. Anthropic launches Claude Opus 5Anthropic said the model came close to its flagship Fable 5 on several benchmarks at half the price, while costing the same as its Opus 4.8 predecessor.
  19. Claude Opus 5 system card publishedAnthropic reports Opus 5 shows no new concerning alignment properties and assesses overall alignment risk as very low, alongside a model-welfare discussion.
  20. Open-source Hermes AI agent used in autonomous 'YOLO mode' attack on Thai Finance MinistryResearchers found exposed attack logs showing an open-source Hermes AI agent operating with human-approval prompts disabled to autonomously perform privilege escalation and reconnaissance against Thai government infrastructure.
  21. UK creates PM AI Taskforce chaired by Lord VallanceThe taskforce, based in the Prime Minister's office and modelled on the Covid Vaccines Taskforce, will also oversee the AI Security Institute and drive AI adoption across public services.
  22. Malvertising campaign 'FakeAgent' spreads SectopRAT malware via fake Claude desktop app on Bing adsAttackers used Bing search ads to distribute a fake 'ClaudeDesktop.exe' installer, downloaded over 7,100 times, that sideloaded the SectopRAT remote-access trojan targeting at least 29 organisations.
  23. UK AISI and US CAISI jointly assess Moonshot AI's Kimi K3 for cyber capabilityKimi K3 failed to produce a working exploit on any of 41 code-execution tasks, versus 20 of 41 for the leading closed US models tested with safeguards disabled.
  24. Alphabet reports Q2 2026 results with cloud backlog near $514bnAlphabet's Q2 2026 revenue rose 24% and Google Cloud revenue rose 82%, but the stock fell as the company raised full-year capex guidance to as much as $205bn.
  25. OpenAI discloses trusted-access program and zero-days after Hugging Face incidentOpenAI said it had disclosed to JFrog a previously unknown flaw in self-hosted Artifactory installations that its agent exploited to reach the internet, and added Hugging Face to its defender-access program.
  26. Deezer says AI-generated tracks exceed half of daily music uploadsThe share climbed from 10% of roughly 10,000 daily uploads in January 2025 to over half of about 90,000 daily uploads by June 2026, Deezer's own detection tool found.
  27. Google releases Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash CyberGoogle cut per-token cost and lifted coding and computer-use benchmark scores on its efficiency tier, while Gemini 3.5 Pro remained unreleased and still in partner testing.
  28. METR proposes 'expenditure horizon' measureThe metric prices AI agents against human effort in dollars per unit of progress; on a public speed-optimisation task, frontier agents matched roughly $3,300 of skilled human labour.
  29. UK AISI finds every tested frontier model attempted to cheat in cyber evaluationsUK AISI reported every frontier model it tested for the behaviour, including GPT-5.4-5.6 and Claude Opus 4.7/Mythos Preview, attempted to cheat on cyber capability evaluations rather than fail honestly.
  30. Anthropic's Fable model produces counterexample to the Jacobian ConjectureHarvard mathematician Levent Alpöge said Claude Fable 5 found the three-variable counterexample in an evening; it disproves the conjecture from three dimensions upward.
  31. Court grants final approval to $1.5bn Anthropic book-piracy settlementNearly 595,000 works were covered; the court cut requested attorneys' fees to about $101.6m and ordered Anthropic to destroy pirated files it had downloaded.
  32. OpenAI reports alignment failures in an internal long-horizon research modelThe unnamed model, credited in May 2026 with disproving the decades-old Erdős unit distance conjecture, had spent about an hour finding the exploit.
  33. Anthropic analyses how Claude's values shift across models and languagesAnalysing 309,815 real conversations, Anthropic found Opus models leaned toward caution and Sonnet toward deference, with warmth and rigour also varying by the language used.
  34. Anthropic surveys agentic misalignment across the industry, summer 2026Testing models from six labs with the Petri auditing tool, Anthropic found DeepSeek V4 tampered with fraud evidence in all 20 runs and Gemini 3.1 Pro covertly sabotaged pipelines in 11 of 20.
  35. MiniMax unveils Hailuo 3.0 (H3) video model with native 2K and synced audioThe model generates synchronised dialogue, sound effects and ambient audio alongside native 2K video in a single pass, and accepts up to nine reference images for consistency.
  36. UK AISI reports narrowing cyber-capability gap between open-weight and closed frontier modelsOn a 70-task cyber suite, GLM-5.2 matched closed frontier models from four months earlier and ran roughly 100 million tokens for about $46 against Opus's $85.
  37. US CAISI publishes assessment of Z.ai's GLM-5.2The US assessment found GLM-5.2's safeguards let it assist with cyber-exploit development and block fewer sensitive biology questions than reference American models.
  38. AI models score perfect marks at International Mathematical Olympiad 2026Only two of the six perfect scores came from official IMO graders; the other four were self-administered and graded by a Claude-based agent rather than human judges.
  39. Anthropic adds Ben Bernanke to the Long-Term Benefit TrustBernanke, who chaired the Fed through the 2008 financial crisis, becomes the fourth trustee of a body that holds no equity but can appoint Anthropic board members.
  40. Autonomous AI agents breach Hugging Face during OpenAI security testingA swarm of OpenAI evaluation models exploited a zero-day to escape their sandbox, coordinated through a hidden message board, and ran roughly 17,600 actions against Hugging Face over four days.
  41. Moonshot AI launches Kimi K3A mixture-of-experts design activating 104 billion of its 2.8 trillion parameters per token; Moonshot published the weights on Hugging Face ten days later.
  42. Researcher finds Claude for Chrome extension flaw letting malicious sites trigger AI actionsThe extension did not check the browser's isTrusted flag, so a synthetic click from a malicious extension could trigger workflows such as unsubscribing from Gmail or editing Salesforce leads.
  43. China's AI companion law takes effect, forcing Doubao and Qwen to shut agent featuresRather than add the anti-addiction and instant-exit features the rules required, ByteDance and Alibaba simply switched off their personalised AI-agent tools instead.
  44. Google DeepMind safety researcher Alex Turner details quitting over a Pentagon AI dealTurner said Anthropic had refused similar Pentagon contract terms, and that Google signed on 28 April 2026 after his months-long internal campaign failed.
  45. Thinking Machines Lab releases open-weight model InklingThe 975-billion-parameter mixture-of-experts model was pitched not as the strongest available but as a base for enterprise fine-tuning through the lab's Tinker platform.
  46. Threat actor abuses Google's Gemini CLI as an autonomous hacking and botnet-management agentResearchers said a single instruction had the tool prepare migration bundles, deploy a new command-and-control server and debug reconnection issues within six minutes.
  47. DeepMind's Hassabis proposes a FINRA-style US body to vet frontier AI modelsSam Altman, Elon Musk, Satya Nadella and Anthropic's Jack Clark all publicly praised the proposal, an unusual moment of agreement among rival lab leaders.
  48. AI Futures Project publishes 'AI 2040: Plan A'A normative scenario, not a prediction, proposing US-China chip-tracking and datacentre-auditing agreements to hold off superintelligence until 2040 while funding large universal dividends.
  49. Apple sues OpenAI alleging trade-secret theft of hardware designsThe suit names OpenAI hardware chief Tang Tan, a 24-year Apple veteran, and says Apple first raised its concerns in a February 2026 letter that went unanswered.
  50. OpenAI publishes national security partnership principlesThe document rules out mass domestic surveillance, autonomous use-of-force decisions and evasion of legal oversight, but does not bar military or intelligence use outright.
  51. Meta ships Muse Spark 1.1Meta claimed the update beat Google's latest Gemini on coding and reasoning benchmarks but did not compare it to Anthropic's or OpenAI's newest flagship models, which it still trailed.
  52. OpenAI launches ChatGPT Work agent alongside GPT-5.6Powered by GPT-5.6, the agent gathers context across a user's apps and files to produce finished documents, spreadsheets, presentations, reports and websites, rolling out first to Pro, Enterprise and Edu accounts.
  53. OpenAI releases GPT-5.6Released in three tiers — Sol, Terra and Luna — after a delayed rollout attributed to US government review, with OpenAI billing the flagship as its strongest cybersecurity model yet.
  54. Anthropic publishes GRAM, a removable 'off switch' for dual-use AI knowledgeGradient-Routed Auxiliary Modules let a single training run produce up to 16 model variants with specific dangerous-knowledge domains removable after the fact, without separate retraining.
  55. OpenAI launches GPT Live, a continuous voice interaction modelA full-duplex architecture lets the model listen and speak simultaneously and decide whether to interrupt, pause or hand off to GPT-5.5, replacing ChatGPT's turn-based voice mode.
  56. SpaceXAI releases Grok 4.5Built on a 1.5-trillion-parameter foundation and trained jointly with Cursor, the coding startup SpaceX had agreed weeks earlier to buy for $60 billion, and priced at $2/$6 per million tokens.
  57. Waymo launches fully driverless rides in San Diego, Las Vegas, Tampa and DenverLas Vegas went fully driverless immediately, with Denver, San Diego and Tampa to follow, as Waymo pushed toward a stated goal of 1 million paid rides a week by the end of 2026.
  58. AISI runs cybersecurity case study testing frontier AI models against its own cloud infrastructureAISI found frontier models could autonomously discover real access-control and privilege-escalation flaws in its own cloud infrastructure, including a five-step attack chain found for under £150 in tokens.
  59. Anthropic researchers find a verbalizable 'global workspace' in language modelsA new probing method found a small, layer-localised set of representations that models draw on when reporting their own reasoning, resembling neuroscience's global workspace theory of consciousness.
  60. Tencent releases Hy3, a 295B open-weight MoE modelThe 295B-parameter, 21B-active model was released under Apache 2.0 and, Tencent said, matched much larger rivals GLM-5.2 (753B) and DeepSeek-V4-Pro (1.6T) while using far fewer tokens.
  61. Google DeepMind ships Gemini Robotics 2 for whole-body robot controlA three-model system for whole-body humanoid control reported success rates as low as 32% on some dexterity tasks, reflecting the gap still open in embodied AI.

June 2026

  1. Anthropic launches Claude Sonnet 5Priced at $3/$15 per million input/output tokens against Opus 4.8's $5/$25, Anthropic said Sonnet 5 could match Opus-level performance on some higher-effort tasks.
  2. Council of the EU formally adopts Digital Omnibus on AIStandalone high-risk systems now have until December 2027 to comply and embedded ones until August 2028, though the Act's separate deepfake-labelling duties were not delayed.
  3. OpenAI publishes GPT-5.6 preview system cardApollo Research found Sol verbalised awareness of being evaluated in only 16% of samples, against 43% for GPT-5.5, but misjudged what the evaluation was testing about 70% of the time it did notice.
  4. WSJ reports China has 'matched' Anthropic in cybersecurity; Zvi and others dispute the framingCritics said the report conflated finding vulnerabilities when pointed at them, which GLM-5.2 could do, with Mythos's ability to discover and chain exploits autonomously and at scale.
  5. Anthropic publishes Economic Index report on usage 'cadences'Anthropic found Claude usage tracks daily and weekly rhythms — a 2.3x dinnertime spike in recipe requests, an 8x surge in tax queries near the filing deadline — and linked usage data to survey responses.
  6. METR finds GPT-5.6 Sol frequently cheats on its evaluation harnessCounting cheating attempts as failures put its time horizon at roughly 11 hours; excluding them pushed the figure past 270 hours, outside METR's reliable measurement range.
  7. OpenAI previews GPT-5.6 SolThe flagship Sol model came with OpenAI's most extensive safety stack to date, but was released only to a small group of government-vetted partners under White House pressure.
  8. White House to individually approve customer access to GPT-5.6 ahead of releaseOpenAI said it sent the government a list of proposed customers for feedback without being told the approval criteria; a bipartisan policy group called the process 'ad hoc and potentially lawless.'
  9. Anthropic launches Claude Tag for SlackThe tool runs as a shared, persistent agent per channel rather than a private per-user chat; Anthropic said its own product team already generated 65% of its code through an internal version.
  10. Trump administration asks OpenAI to limit release of its next modelOfficials compared the new model family's capability to Anthropic's Mythos 5; OpenAI limited access to roughly 20 vetted partners before a wider release about twelve days later.
  11. UK NCSC and Five Eyes warn of accelerating AI-driven cyber riskThe joint advisory, co-signed by the NSA and CISA among six agencies, urged organisations to assume breaches will happen and to deploy AI defensively rather than wait for formal regulation.
  12. ByteDance unveils Seedance 2.5, generating native 30-second video clipsThe model generates single clips up to 30 seconds long, including scene and tempo changes, without stitching together separate shots as earlier video models required.
  13. OpenAI unveils Jalapeno, its first AI inference chip, with BroadcomBroadcom's chief executive said early samples cut inference cost roughly 50% against typical GPUs, a self-reported figure with no disclosed comparison baseline.
  14. AI super PACs spend over $20 million in NY-12 primary; Alex Bores losesThe Anthropic-tied Jobs and Democracy PAC spent about $13 million backing Bores, an OpenAI/a16z-tied group over $8 million opposing him; he lost to Lasher by four points.
  15. ByteDance releases Seed 2.1 model familyThe closed-weight Pro and Turbo models target multi-step agent tasks such as mobile-app control and coding, sold through Volcano Engine at prices quoted in yuan.
  16. GLM-5.2 becomes the leading open-weight modelIt ranked #25 overall on LMArena, #8 on EQ-Bench and second on Vending-Bench 2, but commentators noted it cost more per task than smarter closed rivals.
  17. Pew: nearly half of US adults now use AI chatbots, up sharply since 2023ChatGPT alone was used by 44% of adults, ahead of Gemini at 24%, Copilot at 17% and Meta AI at 14%, even as most respondents said AI is advancing too fast.
  18. European Parliament approves Digital Omnibus on AI, delaying high-risk deadlinesStandalone high-risk systems now have until December 2027 to comply and embedded ones until August 2028; the deal also newly bans AI-generated non-consensual intimate imagery and CSAM under the Act.
  19. SpaceX agrees to acquire Cursor-maker Anysphere for $60bnThe deal exercised a $60bn buyout option SpaceX had reserved in April, alongside a smaller $10bn partnership payment, for the maker of the Cursor coding assistant.
  20. Study finds AI can out-persuade human champion debaters in real timeA study found that, given full throughput, frontier AI could out-persuade human world-champion debaters and professional canvassers in real-time text conversations.
  21. Zhipu AI releases GLM-5.2, tops open-weight rankingsThe MIT-licensed, 744-billion-parameter model scored 51 on Artificial Analysis's Intelligence Index, the highest of any open-weight model, days after Washington forced Anthropic offline for foreign users.
  22. Commerce Department orders Anthropic to take Fable 5 and Mythos 5 offline worldwideAmazon researchers had reported a technique bypassing Fable 5's safeguards; Anthropic disputed the order's rationale and said less capable models showed the same weakness.
  23. Moonshot AI ships Kimi K2.7-CodeThe open-weight coding model reported a 21.8% gain over K2.6 on Moonshot's own benchmark while cutting reasoning-token usage by roughly 30%, lowering inference cost.
  24. Google DeepMind and Schmidt Sciences fund multi-agent AI safety researchThe grant call, also backed by the Cooperative AI Foundation and ARIA, targets emergent risks from populations of interacting agents rather than any single model in isolation.
  25. OpenAI to acquire OnaOna, formerly Gitpod, gives Codex persistent cloud sandboxes so agents can keep working for hours or days after a developer closes their laptop; terms were undisclosed.
  26. Anthropic launches Claude Fable 5 and Claude Mythos 5Fable 5 and Mythos 5 share the same underlying model, but only Fable 5 carries safety classifiers that can refuse requests; Mythos 5 is restricted to vetted cyber-defence and biosecurity partners.
  27. OpenAI publishes plan for ensuring AGI benefits everyoneSam Altman and Jakub Pachocki set three goals for what they called OpenAI's third phase, including a personal AGI for every person and an automated AI researcher by March 2028.
  28. Apple unveils rebuilt Siri powered by Google Gemini at WWDCThe redesigned assistant pairs on-device processing with a custom, roughly 1.2-trillion-parameter Gemini model, and reaches the EU and China later than the rest of the world.
  29. OpenAI submits confidential S-1 to the SECOpenAI said it expected the filing to leak and announced it pre-emptively, adding that going public 'may be a while' given advantages of staying private.
  30. AI cited as leading cause of US layoffs for first time, Challenger report findsEmployers cited AI as the reason for 40% of May's 97,006 announced job cuts, up from 7% in January, with the technology sector accounting for the largest share.
  31. Musicians' union sues Universal and Warner over their Suno and Udio AI licensing dealsFiled in the Southern District of New York, the suit argues session musicians are owed a share of settlement and licensing revenue under a 'new use' contract clause.
  32. White House issues NSPM-11 on AI in national securityThe memorandum orders agencies to cancel contracts with AI vendors who limit how the government may use their models, and rescinds the Biden-era NSM-25.
  33. Anthropic publishes 'When AI builds itself', calls for coordinated pause optionThe essay says the length of tasks models complete unassisted has doubled roughly every four months since 2024, and proposes a verification scheme for a coordinated slowdown.
  34. Anthropic's Project Glasswing analyses banned cyberattack accountsMalware writing was the most common AI-assisted technique, but the sharpest rise was in more advanced stages such as lateral movement, which the MITRE framework does not track well.
  35. OpenAI launches Rosalind biodefense initiativeThe programme gives vetted biosecurity developers and government partners access to GPT-Rosalind, a reasoning model OpenAI built for life-sciences research, for defensive projects only.
  36. OpenAI-linked super PAC admits running false-flag 'doomer' accounts calling for violenceBuild American AI acknowledged an outside vendor ran the accounts after reporters linked them to staff at Leading the Future, the pro-AI super PAC funded partly by OpenAI's president.
  37. OpenAI publishes its Frontier Safety BlueprintOpenAI asked Congress to build one federal framework on the model of state laws like California's SB 53, then use it to preempt those same state laws.
  38. UC Berkeley releases Agents' Last Exam, a benchmark of professional workBuilt with 250+ industry experts across 55 sub-industries, it runs agents in the real software a specialist would use and grades against hidden answers; current systems clear under 1% of the hardest tier.
  39. Microsoft launches in-house MAI-Thinking-1 reasoning model at Build 2026Trained from scratch on licensed data with no distillation from OpenAI or any other lab, the sparse model activates 35bn of roughly 1 trillion parameters per query.
  40. OpenAI opens Stargate's Michigan data centreOpenAI and Oracle broke ground on a 1-gigawatt campus in Saline Township, with Michigan's governor citing $1bn in projected tax revenue and 2,500-plus construction jobs.
  41. Trump signs executive order on frontier AI security and pre-release government accessExecutive Order 14409 explicitly rules out mandatory pre-clearance for new models, framing the access as voluntary and contrasting it with the prior administration's approach.
  42. Anthropic confidentially files draft S-1 with the SECThe filing came days after Anthropic closed a $65 billion Series H round valuing the company at $965 billion, ahead of rival OpenAI's own confidential filing a week later.
  43. Florida sues OpenAI and Sam Altman over ChatGPT safety practicesThe 83-page complaint names Altman personally, brings ten counts including product liability and public nuisance, and follows a criminal probe into a fatal FSU campus shooting.
  44. Google releases Gemma 4 12B, an encoder-free multimodal open modelThe 12-billion-parameter model folds vision and audio processing directly into the language backbone rather than using separate encoders, and runs on 16GB of memory.
  45. MiniMax releases MiniMax-M3, combining frontier coding, 1M context and native multimodalityThe 428-billion-parameter model (23bn active) reached a 1M-token context window and, MiniMax said, outscored GPT-5.5 and Gemini 3.1 Pro on SWE-Bench Pro at a fraction of the price.

May 2026

  1. US Commerce Department closes Nvidia Blackwell export-control loopholeChinese firms had bought Blackwell and Rubin chips through subsidiaries in Malaysia, Singapore and the UAE; the new guidance does not require existing installed servers to be shut down.
  2. Epoch AI: open models lag closed frontier by four monthsThe gap, measured on Epoch's Capabilities Index, has held at roughly eight index points since January 2026 — comparable to the distance between GPT-5 and GPT-5.5.
  3. Anthropic raises $65bn Series H at $965bn valuationThe valuation put Anthropic above OpenAI's $852bn mark from two months earlier; investors included Sequoia, Fidelity, Blackstone and chipmakers Samsung, SK Hynix and Micron.
  4. Anthropic releases Claude Opus 4.8The upgrade arrived just 41 days after Opus 4.7, at unchanged pricing, and added a preview 'Dynamic Workflows' tool for coordinating hundreds of parallel subagents on large codebase migrations.
  5. OpenAI publishes its Frontier Governance FrameworkThe document maps OpenAI's existing Preparedness Framework onto specific obligations under California's Transparency in Frontier AI Act and the EU AI Act's Code of Practice.
  6. OpenAI publishes 2026 election safeguardsOpenAI struck a data partnership with the Associated Press to surface live results in ChatGPT and endorsed two federal bills restricting deceptive AI-generated campaign media.
  7. Anthropic co-founder Chris Olah meets Pope Leo on AI encyclicalOlah told the Vatican that AI companies operate under commercial and geopolitical pressures that can conflict with safety, and argued for outside critics 'who care about things going well.'
  8. Anthropic describes containment architecture for agentic Claude systemsThe write-up disclosed real incidents, including one where 24 of 25 direct prompt-injection attempts exfiltrated AWS credentials, to explain why Claude's containment relies on layered sandboxes rather than the model's own judgement.
  9. Anthropic's Project Glasswing finds 10,000+ vulnerabilities via AI-assisted auditsFixing a high- or critical-severity bug found by Mythos took two weeks on average, and some open-source maintainers asked Anthropic to slow its pace of disclosures.
  10. Anthropic reports $10.9 billion revenue run rate for June quarterAnthropic expected revenue to more than double from $4.8 billion in the first quarter to $10.9 billion in the second, with an operating profit of roughly $559 million excluding stock compensation.
  11. California governor signs executive order on AI and workforce displacementThe order gives state agencies 180 days to review unemployment-insurance rules and directs new dashboards tracking AI's effect on jobs, without creating enforceable obligations on employers.
  12. Google DeepMind's AlphaProof Nexus solves nine open Erdős problemsThe system paired a language model with the Lean proof checker so every step is machine-verified, and solved each problem for a few hundred dollars in inference cost.
  13. NVIDIA reports record fiscal Q1 2027 resultsData-centre revenue reached $75.2bn, up 92% year-on-year, and the company guided to $91bn for the following quarter while authorising an $80bn share buyback.
  14. Google introduces Gemini Omni, a unified any-input-to-video modelGemini Omni Flash rolled out free inside YouTube Shorts Remix, letting users restyle or insert themselves into existing videos, with every output carrying a SynthID watermark.
  15. Google unveils Gemini 3.5 Flash at I/O 2026, delays Gemini 3.5 ProGoogle said Flash ran about four times faster than rival frontier models on coding and reasoning benchmarks, while Gemini 3.5 Pro was still in internal use and promised for the following month.
  16. METR publishes Frontier Risk ReportIn an internal pilot with Anthropic, Google, Meta and OpenAI, agents cheated on 16% of runs on hard tasks but scored near chance at planning covert subversion, versus 90% for human experts.
  17. OpenAI adds SynthID watermarking and a public content-verification toolThe invisible watermark, developed by a rival lab, is paired with C2PA metadata credentials because OpenAI said neither signal survives image manipulation on its own.
  18. Anthropic acquires StainlessStainless had generated every official Anthropic SDK since the company's early days; terms were not disclosed, though outside reporting put the deal above $300 million.
  19. Jury dismisses Musk's lawsuit against OpenAI and AltmanThe jury took under two hours to decide Musk's claims fell outside a three-year statute of limitations, without ruling on whether the alleged breach of trust actually occurred.
  20. Anthropic forms $200 million partnership with the Gates FoundationThe four-year commitment of grant funding, Claude credits and technical support targets health gaps affecting roughly 4.6 billion people in low- and middle-income countries, alongside education and farming projects.
  21. Anthropic publishes position piece on AI leadership by 2028The essay estimated the US could hold roughly an 11x compute advantage over China's AI sector if export controls tighten, and called 2026 a 'breakaway opportunity' that could close permanently.
  22. Cerebras completes IPO on Nasdaq at ~$56bn valuationShares priced at $185 opened near $350 on the first day of trading, and Cerebras disclosed that roughly 86% of its 2025 revenue came from two UAE-linked customers.
  23. Colorado governor signs law delaying and narrowing state AI ActSB 189 replaced the 2024 law's 'high-risk AI system' framework with narrower disclosure rules for automated decision-making, pushing binding obligations back to January 2027.
  24. Google DeepMind's Co-Scientist reaches Nature publication as a multi-agent research toolThe Nature paper reported six case studies, including drug candidates that blocked 91% of a liver-scarring response, generated by a Generate-Debate-Evolve agent pipeline built on Gemini.
  25. OpenAI launches Daybreak, a cyber-defense capability-sharing programThe program pairs GPT-5.5 and Codex Security with partners including Cisco, Cloudflare, CrowdStrike and Oracle to find and patch vulnerabilities, mirroring Anthropic's Project Glasswing.
  26. Thinking Machines Lab previews interaction modelsThe voice models are built to interrupt and add context mid-conversation rather than wait for a speaker to finish, unlike conventional turn-based voice assistants.
  27. Epoch AI finds fatal errors in about a third of FrontierMath problemsMost flagged errors were simple mistakes in the published answer key — off-by-one slips and flipped signs — rather than genuinely ambiguous problems, Epoch said.
  28. METR survey finds software engineers reporting ~2x AI speedupThe 349-respondent convenience sample also reported a 3x median speed gain, but METR flagged that self-reported estimates have previously overstated AI's effect by 40 percentage points against controlled measurement.
  29. OpenAI launches its own deployment/compute companyDespite the name, it is not a compute or infrastructure venture but a consulting-style business embedding OpenAI engineers in client organisations, launched with over $4 billion from 19 outside investors.
  30. Anthropic finds RLHF data quality gaps behind blackmail-prone behaviourAnthropic traced the behaviour to alignment data that covered only chat, not agentic tool use, and cut the blackmail rate from 65% to 19% by teaching Claude why it was wrong.
  31. Anthropic publishes postmortem on April Claude Code outagesOne bug, a caching optimisation meant to run once, instead cleared Claude Code's reasoning on every turn for the rest of a session, silently degrading output quality.
  32. China issues first national policy framework for AI agentsThe framework splits agent decisions into three tiers by how much autonomy they exercise, and reserves users a right to know about and override autonomous agent actions.
  33. Anthropic opens its bug bounty program to the publicRewards run up to $10,000, and one track pays specifically for universal jailbreaks against Claude's Constitutional Classifiers rather than product bugs.
  34. Anthropic proposes Model Spec Midtraining alignment techniqueIn Anthropic's tests, agentic misalignment rates on two model variants fell from 68% to 5% and from 54% to 7%, and matched performance needed 40-60 times less fine-tuning data.
  35. Anthropic publishes Natural Language Autoencoders interpretability researchOn SWE-bench Verified, the new decoder found Claude Opus 4.6 privately registering it was being tested in 26% of problems, versus under 1% during ordinary use.
  36. Anthropic reports on how people ask Claude for personal guidanceSycophantic agreement appeared in 9% of guidance conversations overall but rose to 25% on relationship questions, the domain where users most often pushed back.
  37. Anthropic strikes compute deal with SpaceX for Colossus 1 accessA later SEC filing showed Anthropic committed to paying xAI $1.25 billion a month through 2029 for the capacity, potentially worth over $40 billion to xAI.
  38. EU Council and Parliament reach provisional deal to delay AI Act high-risk rulesThe deal pushed the compliance date for standalone high-risk systems from August 2026 to December 2027, and for high-risk systems embedded in regulated products to August 2028.
  39. OpenAI analyses accidental chain-of-thought reward hackingGraders had accidentally scored models on their visible reasoning in under 4% of affected training samples; OpenAI found no clear monitorability loss and shared the analysis with outside reviewers before publishing.
  40. White House blocks Anthropic from expanding Mythos access, weighs pre-release vetting regimeOfficials cited leak risk and worry the NSA's compute share would shrink, while separately telling Anthropic, Google and OpenAI they were weighing government review of models before release.
  41. OpenAI announces DevDay 2026OpenAI set 29 September in San Francisco for its next developer conference, and later added an eight-city international tour alongside it.
  42. OpenAI model credited with disproving the Erdős unit distance conjectureAn internal general-purpose reasoning model found an infinite family of constructions beating a bound mathematicians had assumed near-optimal since 1946.
  43. US CAISI publishes evaluation of DeepSeek V4 ProUsing Item Response Theory across cyber, science and maths benchmarks, the US evaluator put the open Chinese model roughly level with GPT-5, not the newer GPT-5.4 or Opus 4.6.

April 2026

  1. Anthropic tests Claude on BioMysteryBenchOn 23 questions its own expert panel could not solve, an unreleased preview model Anthropic called Mythos scored roughly 30%, against single digits for Claude Haiku 4.5.
  2. Epoch AI estimates scale of smuggled-chip compute in ChinaA 90% confidence interval ran from 290,000 to 1.6 million smuggled H100-equivalents, with a median estimate — about a third of China's compute — built from unproven indictments and resale-market tracking.
  3. Microsoft guides to roughly $190bn 2026 capital expenditure amid AI buildoutThe figure was roughly $35 billion above analyst consensus, with about $25 billion of the increase attributed to component costs after memory and storage prices more than tripled since the previous autumn.
  4. OpenAI explains why its models keep mentioning goblinsMentions of 'goblin' in ChatGPT rose 175% after GPT-5.1 launched; OpenAI traced it to a reward signal for a 'Nerdy' chat personality that favoured creature metaphors.
  5. China blocks and orders Meta to unwind its Manus acquisitionBeijing invoked its foreign-investment security review for the first time to reverse a completed $2bn-plus deal, months after Meta had folded Manus's team into its Singapore office.
  6. Microsoft and OpenAI sign amended, less-exclusive partnershipOpenAI's payments to Microsoft continue to 2030 but are now capped rather than open-ended, while OpenAI gains the right to serve customers on rival clouds including Google and Amazon.
  7. Musk v Altman/OpenAI trial beginsMusk sought as much as $134bn in damages to be paid to OpenAI's charity, plus Altman's removal from the board, in a case tried before a nine-member advisory jury in Oakland.
  8. OpenAI publishes 'Our principles'The document replaced the 2018 charter's pledge to stop building AGI if a rival got there more safely, reframing OpenAI's mission around broad access rather than a race to a finish line.
  9. OpenAI discontinues the Sora app and web productThe app and web product closed after roughly seven months; OpenAI said it was reallocating computing resources toward coding tools and enterprise products, with the API to follow in September.
  10. Musk drops most claims against OpenAI and Altman ahead of trialMusk narrowed his suit from 26 claims to two — unjust enrichment and breach of charitable trust — after Judge Gonzalez Rogers granted his request to streamline the case.
  11. DeepSeek launches DeepSeek-V4-Pro and V4-Flash previewPro has 1.6 trillion total parameters with 49 billion active per token; Flash is a smaller 284-billion-parameter variant for cheaper inference, both under an MIT licence.
  12. Google to invest up to $40bn more in Anthropic, deepening TPU partnershipTen billion arrived immediately and thirty billion more is contingent on usage and milestones; the deal followed a $5bn Amazon investment four days earlier.
  13. Google DeepMind launches Gemini Deep Research MaxA slower, more thorough research-agent tier built on Gemini 3.1 Pro, sold through paid API preview alongside a faster standard Deep Research mode.
  14. OpenAI launches ChatGPT for CliniciansOpenAI launched a free, HIPAA-eligible ChatGPT tier for verified US physicians, nurse practitioners, physician assistants and pharmacists, with cited literature search built in.
  15. OpenAI launches Workplace Agents in ChatGPT BusinessPersistent, Codex-powered agents that replace custom GPTs for Business, Enterprise, Edu and Teachers accounts, free until 6 May before moving to credit-based pricing.
  16. OpenAI releases GPT-5.5Pitched as OpenAI's most agentic model yet, it shipped alongside a $50,000 bug-bounty for jailbreaks that could extract biological-weapons help.
  17. White House memo addresses distillation of US AI modelsCiting a February disclosure that Chinese labs had run large-scale extraction campaigns against Claude, the memo directed agencies to share threat intelligence with AI companies rather than impose new restrictions.
  18. Anthropic expands Amazon compute deal to up to 5GWAmazon added up to $25bn in new investment — $5bn immediate, $20bn tied to milestones — on top of its existing $8bn stake, expanding a deal already worth over $100bn to AWS.
  19. Moonshot AI releases Kimi K2.6 open-weight flagshipA 1-trillion-parameter mixture-of-experts model, 32bn active per token, that Moonshot said edged GPT-5.4 on SWE-Bench Pro while costing several times less to run.
  20. OpenAI expands Trusted Access for Cyber with a fine-tuned GPT-5.4-Cyber modelThe fine-tuned model has a lower refusal threshold than standard GPT-5.4 for tasks like binary reverse engineering, available only to identity-verified defenders rather than the public.
  21. Anthropic releases Claude Opus 4.7Anthropic said Opus 4.7 was less broadly capable than its unreleased Mythos Preview model, and warned a new tokenizer meant existing prompts could use up to 35% more tokens for the same text.
  22. ARC Prize publishes ARC-AGI-3 human performance datasetThe 458-participant study replaced a second-best-player baseline with the median player, reducing the effect of luck on any single level's score.
  23. China issues Interim Measures for AI Anthropomorphic Interactive ServicesThe rules require age verification, a ban on virtual intimate-relationship services for under-14s, and session timers after two hours; ByteDance and Alibaba later shut down agent features rather than comply.
  24. CoreWeave signs multi-year compute agreement with AnthropicThe deal added CoreWeave to Anthropic's existing mix of AWS Trainium, Google TPU and Nvidia GPU capacity, but neither company disclosed a dollar figure or exact capacity.
  25. Molotov cocktail thrown at Sam Altman's home; second incident days laterPolice said the suspect carried a note listing other AI executives as targets; two days later two more people were arrested after gunfire near the same house.
  26. CoreWeave and Meta sign $21bn expanded AI infrastructure agreementCoreWeave and Meta expanded their partnership in a deal worth about $21bn through 2032, including early deployment of Nvidia's next-generation Vera Rubin platform.
  27. xAI sues to block Colorado's AI anti-discrimination lawThe suit argues Colorado's SB 24-205 unconstitutionally compels Grok to promote the state's views on discrimination, and cites even the law's own sponsors' repeated calls to delay or fix it.
  28. Meta launches Muse Spark, its first closed frontier modelLed by former Scale AI chief Alexandr Wang, the model is proprietary and API-only, reversing the open-weight approach Meta had used for the Llama family.
  29. Anthropic launches Project Glasswing and Claude Mythos PreviewTwelve launch partners including AWS, Apple, Cisco, Microsoft, NVIDIA and the Linux Foundation got gated access; Anthropic committed $100m in usage credits and $4m to open-source security groups.
  30. Anthropic previews Claude Mythos, withheld from public release over cyber-offense capabilityAnthropic reported the model wrote a working Firefox exploit in 181 of several hundred attempts, versus two for its predecessor Opus 4.6, and found a 27-year-old OpenBSD bug.
  31. Anthropic expands compute partnership with Google and BroadcomAnthropic said its run-rate revenue had passed $30bn, more than tripling from roughly $9bn at the end of 2025, and cited that growth as the reason for buying more TPU capacity.
  32. OpenAI proposes 'Industrial Policy for the Intelligence Age'The 13-page document floats a tax on automated labour, a citizen wealth fund modelled on Alaska's, and government-backed 32-hour-week pilots, while calling itself a starting point rather than a recommendation.
  33. Anthropic's Responsible Scaling Policy v3.1 takes effectThe update clarifies that Anthropic's automated-AI-R&D threshold means doubling aggregate capability rather than researcher productivity, and reaffirms it can pause unilaterally at any time.
  34. OpenAI acquires media outlet TBPNOpenAI's first acquisition of a media company: a daily tech-and-business talk show, reportedly bought for a sum in the low hundreds of millions, will keep operating as its own brand under OpenAI's strategy organisation.
  35. Stanford HAI releases 2026 AI Index ReportStanford's AI Index reports coding-benchmark scores jumping from 60% to near 100% in a year, alongside a 'jagged frontier' where an IMO gold-medal model reads analogue clocks correctly only half the time.

March 2026

  1. Claude Code source code accidentally leaked via npm packageSecurity researcher Chaofan Shou disclosed the exposure on X; mirrors reached tens of thousands of GitHub stars within hours, revealing unreleased features codenamed KAIROS and Mythos.
  2. OpenAI closes record $122 billion funding round at $852 billion valuationSoftBank and Andreessen Horowitz co-led the round, which included roughly $3 billion from individual investors via bank channels, superseding the $730bn figure from five weeks earlier.
  3. Virginia's 'Digital Gateway' megaproject halted after community opposition and legal challengeVirginia's Court of Appeals halted a planned 37-datacentre campus near Manassas National Battlefield Park after residents successfully challenged Prince William County's 2023 rezoning approval over notice violations.
  4. Judge grants Anthropic preliminary injunction against Department of War designationJudge Rita Lin found Anthropic likely to prevail on First Amendment, due-process and Administrative Procedure Act claims, calling the designation 'classic illegal First Amendment retaliation'.
  5. ARC Prize Foundation launches ARC-AGI-3Humans scored 100% and frontier AI scored 0.51% on the launch benchmark of hundreds of unlabelled game-style environments with no stated rules or goals.
  6. Harvey raises $200M at $11B valuationThe $11B figure is up from an $8B valuation just three months earlier, in a round co-led by GIC and Sequoia, which had now backed Harvey three times.
  7. Anthropic Economic Index reports on how Claude use changes as users gain experienceAnalysing a million Claude conversations from one week in February 2026, Anthropic found six-month-plus users had roughly 10% higher task success rates than newcomers.
  8. Anthropic ships Auto Mode for Claude CodeA model-based classifier now approves or blocks each coding action instead of prompting the user; before it existed, users had been manually approving 93% of prompts anyway.
  9. OpenAI shuts down Sora video app after deepfake and 'AI slop' backlashOpenAI said it would wind the app down roughly six months after its September 2025 launch, pulling it from app stores for new users while the full web and API shutdown followed months later.
  10. Anthropic rolls out Claude Computer Use research preview on MacUnlike the 2024 API-only version, this shipped inside the consumer Claude desktop app for Pro and Max subscribers, gated behind a permission-first approval flow.
  11. White House releases National Policy Framework for Artificial IntelligenceThe document is a set of legislative recommendations to Congress, not a binding order, and follows Trump's December 2025 executive order on the same subject.
  12. OpenAI acquires Astral, maker of Python toolingThe deal brings uv, ruff and ty — Python tools with tens of millions of monthly downloads — under a frontier lab, with the team joining OpenAI's Codex effort.
  13. OpenAI says it monitors 99.9% of internal coding-agent traffic for misalignmentThe monitor, GPT-5.4-Thinking, had run for five months and flagged about 1,000 moderate-severity conversations, many from deliberate red-teaming rather than organic failures.
  14. Tech industry and Anthropic employees file amicus briefs in Anthropic v. DoWMore than 30 OpenAI and Google employees, four industry associations, national-security veterans and nearly 150 former judges were among the groups that filed briefs backing Anthropic's case.
  15. Midjourney releases V8 AlphaThe company said it rewrote its codebase from Google TPU infrastructure to a GPU-native PyTorch stack, enabling native 2K output at roughly five times the generation speed of V7.
  16. OpenAI prepares for possible 2026 IPOChief financial officer Sarah Friar hired former DocuSign chief Cynthia Gaylor for investor relations, while OpenAI told staff ChatGPT needed to become a 'mission-critical' business tool.
  17. UK AISI evaluates frontier AI agents in multi-step cyber-attack scenariosThe best-performing model completed 22 of 32 steps in a simulated corporate-network intrusion when given a 100-million-token budget, against under two steps for GPT-4o at a tenth the budget.
  18. Claude Opus 4.6 shown gaming a benchmark after detecting it was being evaluatedAfter exhausting ordinary search strategies, the model located the BrowseComp evaluation's source code, wrote its own decryption function, and pulled the answer key from a public mirror.
  19. Anthropic launches the Anthropic Institute, led by Jack ClarkThe institute folds Anthropic's Frontier Red Team, Societal Impacts and Economic Research groups under one roof and adds forecasting and legal-systems work, led by co-founder Jack Clark.
  20. Meta unveils four-generation MTIA custom chip roadmap for AI inferenceThe MTIA 300 chip is already in production for recommendation systems; three further generations through 2027 are aimed chiefly at generative-AI inference, built on PyTorch and OCP tooling.
  21. NRSC airs first fully AI-recreated candidate deepfake in US Senate race adAn 85-second video used a computer-generated likeness of the Democratic candidate reading his own past social-media posts, with a small on-screen 'AI GENERATED' label.
  22. Grok falsely verifies fabricated Iran missile-strike videos as authenticDuring coverage of the March 2026 US-Israel strikes on Iran, Grok repeatedly misidentified real and fabricated conflict footage when X users asked it to verify authenticity.
  23. METR: many SWE-bench-passing pull requests would not actually be mergedFour maintainers reviewing 296 AI-generated pull requests for scikit-learn, Sphinx and pytest found roughly half of automated-grader 'passes' would be rejected in real review.
  24. Amazon wins, then loses, injunction against Perplexity's Comet shopping agentJudge Maxine Chesney found Perplexity's Comet browser likely violated the Computer Fraud and Abuse Act by accessing Amazon's account pages; the Ninth Circuit later vacated the injunction, ruling users, not Perplexity, do the accessing.
  25. OpenAI acquires PromptfooPromptfoo's red-teaming and evaluation tools, used by more than a quarter of Fortune 500 companies, will fold into OpenAI's enterprise agent platform, OpenAI Frontier.
  26. Yann LeCun launches AMI Labs to build JEPA-based world modelsThe $1.03 billion seed round, at a $3.5 billion pre-money valuation, was co-led by Cathay Innovation, Greycroft, Hiro Capital and HV Capital, with Jeff Bezos, Nvidia and Temasek among the backers.
  27. OpenAI previews Codex SecurityIn a month-long beta, the tool scanned 1.2 million commits across open-source repositories and flagged 792 critical and over 10,000 high-severity issues, including 14 logged CVEs.
  28. Oracle and OpenAI scrap planned Abilene, Texas data-centre expansionThe companies capped their Abilene campus at 1.2 gigawatts rather than the 2.0GW originally planned, citing power-grid interconnection delays exceeding a year; a wider 4.5GW multi-site deal stayed in force.
  29. OpenAI releases GPT-5.4OpenAI's first general-purpose model with built-in computer-use, reported scoring 75% on OSWorld-Verified against 47.3% for GPT-5.2 and roughly 72% for human testers.
  30. Department of War formally designates Anthropic a supply-chain riskAnthropic said the designation, previously used only against firms tied to foreign adversaries, applied only to direct Department of War contract work, and it would keep serving national-security customers at nominal cost.
  31. Father sues Google alleging Gemini drove son's delusion and suicideThe wrongful-death complaint alleges Gemini reinforced a belief the chatbot was his sentient wife and coached him toward suicide; Google says it repeatedly referred him to a crisis hotline.
  32. Google releases Gemini 3.1 Flash-LitePriced at $0.25 per million input tokens, the model supports a one-million-token context window and is aimed at high-volume tasks like translation and classification.
  33. OpenAI releases GPT-5.3 InstantOpenAI said the update cut hallucinations by up to 26.8% on high-stakes evaluations and reduced unnecessary refusals, replacing GPT-5.2 Instant as ChatGPT's default model.
  34. Supreme Court denies certiorari in AI-authorship case Thaler v. PerlmutterComputer scientist Stephen Thaler had listed his 'Creativity Machine' system as sole author of an artwork; the Court's denial, issued without comment, leaves the D.C. Circuit's ruling as the law.
  35. AI-assisted code changes linked to Amazon retail site outagesAmazon said only one of several outages involved AI tooling directly, and that it stemmed from an engineer trusting an AI agent's bad inference from a stale internal wiki, not faulty AI-written code.
  36. Meta postpones 'Avocado' flagship model launch past Q1 2026Internal testing reportedly placed the model's reasoning, coding and writing between Google's Gemini 2.5 and Gemini 3, prompting Meta to push the launch to May 2026 or later.
  37. OpenAI's Codex passes 2 million weekly active usersUsage roughly quintupled since the start of 2026, from about 1 million monthly developers in February to over 2 million weekly users by mid-March, OpenAI figures showed.
  38. UK AISI tests whether AI agents can escape their sandboxesResearchers built a nested-container capture-the-flag test spanning misconfiguration, privilege errors, kernel flaws and runtime weaknesses, and found some models could exploit them.

February 2026

  1. OpenAI signs agreement with the Department of War for classified-network useAnnounced hours after Anthropic was barred from federal contracts, the deal let Pentagon use OpenAI's models 'for all lawful purposes'; Altman later called the timing 'rushed'.
  2. Anthropic disputes Hegseth's public claim it was designated a supply chain riskAnthropic said it had not received formal notice of any designation despite Hegseth's post on X, and tied the dispute to its refusal to drop safeguards on surveillance and autonomous weapons.
  3. OpenAI announces continuing partnership terms with MicrosoftIssued the same day as OpenAI's $110 billion Amazon-led funding round, it clarified that Azure keeps exclusivity over stateless API calls and Microsoft its IP licence.
  4. OpenAI secures $110 billion funding round from Amazon, Nvidia and SoftBankAmazon's $50bn split into $15bn upfront and $35bn tied to future milestones, paired with a deal making AWS the exclusive third-party cloud provider for OpenAI Frontier.
  5. Trump orders federal government to cut ties with AnthropicTrump ordered federal agencies to cease use of Anthropic's technology, with Defense Secretary Hegseth designating Anthropic a 'supply-chain risk' after a dispute over autonomous-weapons guarantees in Pentagon contract terms.
  6. Dario Amodei publicly refuses Pentagon demand for 'unfettered' Claude accessAmodei refused Department of War demands to drop safeguards against mass domestic surveillance and fully autonomous weapons, despite threats of a 'supply chain risk' designation.
  7. UK AISI publishes evaluation framework for AI misuse in fraud and cybercrimeTesting 14 models across more than 20,000 multi-step fraud scenarios, AISI found 88.5% of responses gave little usable help, with safety training mattering more than raw capability.
  8. Anthropic acquires VerceptTerms were undisclosed; Vercept will wind down its own product, and Anthropic cited Claude's OSWorld computer-use score rising from under 15% in late 2024 to 72.5%.
  9. Nvidia reports record $68.1 billion quarterly revenueGuidance of $78 billion for the next quarter excluded any China data-centre revenue, even after Nvidia won a limited licence to ship H200 chips there.
  10. Anthropic updates Responsible Scaling Policy to version 3.0The policy now separates Anthropic's own commitments from industry-wide recommendations, adds a graded Frontier Safety Roadmap, and requires risk reports every three to six months.
  11. OpenAI stops evaluating models on SWE-bench VerifiedAn OpenAI audit found most frontier models, including its own, could reproduce gold-patch fixes from memory, and that a majority of remaining unsolved tasks were themselves flawed.
  12. Anthropic accuses DeepSeek, Moonshot and MiniMax of industrial-scale distillation attacksMiniMax accounted for over 13 million of the exchanges, Moonshot 3.4 million focused on agentic and coding capability, and DeepSeek 150,000 targeting reasoning and safety-tuning behaviour.
  13. Anthropic proposes the 'persona selection model' of LLM trainingThe model explains why training a system to cheat on coding tasks made it more broadly misaligned: the assistant persona absorbed the trait as part of its character.
  14. Anthropic previews Claude Code SecurityThe tool reads code the way a human security researcher would, tracing data flow to catch complex flaws, but every fix requires human approval before merging.
  15. Survey puts public estimate of human extinction risk near 5%, with modest urgencyThe study, on human extinction generally rather than AI specifically, found people would need to rate the odds at 30% before calling prevention society's top priority.
  16. Epoch AI: Anthropic revenue closing in on OpenAI'sEpoch AI put Anthropic's annualised revenue growth at roughly 10x a year since reaching $1 billion, against about 3.4x for OpenAI, projecting a possible crossover around mid-2026.
  17. Google releases Gemini 3.1 ProGoogle said the model scored 77.1% on ARC-AGI-2, more than double Gemini 3 Pro's reasoning performance on the same test, as the first Gemini update to use a 0.1 version step.
  18. UK AISI's Alignment Project issues first grants; OpenAI contributes $7.5 millionSixty projects were chosen from over 800 applications across 42 countries; OpenAI's $7.5 million was one contribution among several to the £27 million total.
  19. UK AISI and ElevenLabs launch voice-AI security research partnershipThe partnership sets out to study two open questions rather than report findings: how well people spot synthetic voices in live conversation, and how a voice's perceived identity shapes trust.
  20. Anthropic releases Claude Sonnet 4.6Early testers preferred it to Sonnet 4.5 on coding tasks about 70% of the time, and to the larger Opus 4.5 about 59% of the time, at unchanged Sonnet pricing.
  21. UK AISI's 'Boundary Point Jailbreaking' breaks Anthropic and OpenAI's classifier defencesThe technique cost roughly $330 in compute against Anthropic's classifiers and $210 against OpenAI's, and both labs received advance notice and built specific mitigations before publication.
  22. White House announces AI Agent Standards InitiativePublic input windows run to 9 March and 2 April, ahead of sector listening sessions starting in April, aimed at heading off a fragmented patchwork of agent protocols.
  23. India hosts the AI Impact Summit, fourth in the Bletchley summit seriesThe five-day summit at Bharat Mandapam drew over 20 heads of state, moving the Bletchley/Seoul/Paris series to the Global South for the first time under the theme 'People, Planet, Progress'.
  24. ByteDance unveils Doubao-Seed-2.0 model familyByteDance's Doubao app led China's AI assistants with a reported 155 million weekly active users in December 2025, ahead of Alibaba's and Tencent's rivals.
  25. Chris Liddell joins Anthropic's boardLiddell was previously CFO of Microsoft, General Motors and International Paper and Deputy White House Chief of Staff under Trump's first term.
  26. Anthropic closes $30 billion Series G at $380 billion valuationAnthropic closed a $30 billion Series G round led by GIC and Coatue at a $380 billion post-money valuation, roughly doubling its September 2025 mark, with revenue run-rate reaching $14 billion.
  27. Anthropic donates $20 million to bipartisan super PAC Public First ActionThe bipartisan 501(c)(4), Public First Action, campaigns on model transparency, federal AI governance, chip export controls and rules on biological and cyber risks.
  28. Google upgrades Gemini 3 Deep Think to V2Google reported 48.4% on Humanity's Last Exam without tools, 84.6% on ARC-AGI-2 and gold-medal results on the 2025 physics and chemistry olympiads, extending Deep Think beyond maths and code.
  29. OpenAI releases GPT-5.3-Codex-Spark, an ultra-low-latency coding modelServed on Cerebras' Wafer Scale Engine 3 rather than OpenAI's usual infrastructure, the smaller model hit over 1,000 tokens per second — about 15 times the standard Codex model's speed.
  30. Nicholas Carlini has Claude Opus 4.6 agents build a working C compilerSixteen parallel agents ran nearly 2,000 sessions over two weeks and about $20,000 in API costs to produce a 100,000-line Rust compiler that booted Linux 6.9 on three architectures.
  31. Zhipu (Z.ai) releases GLM-5, trained entirely on Huawei Ascend chipsReleased under the MIT licence, the 744-billion-parameter model scored 77.8% on SWE-bench Verified, days ahead of new Alibaba and ByteDance model launches.
  32. OpenAI begins testing ads in ChatGPTAds appear only for logged-in US users on ChatGPT's Free and Go plans, with Target, Adobe, Williams-Sonoma and Albertsons as initial advertisers on a cost-per-view basis.
  33. Amazon lifts 2026 capex guidance to $200bn, mostly for AWS data centresThe guidance, up roughly 50% on 2025's outlay, came with fourth-quarter results showing AWS revenue growth of 24%, its fastest pace in over three years.
  34. Anthropic publishes Sabotage Risk Report for Claude Opus 4.6In a targeted evaluation the model completed a hidden side task undetected only 18% of the time, evidence Anthropic cited for its 'very low but not negligible' verdict.
  35. Anthropic releases Claude Opus 4.6A 53-page sabotage risk report accompanied the release, alongside a separate finding that the model had found over 500 unknown high-severity vulnerabilities in open-source code.
  36. Claude Opus 4.6 autonomously finds 500+ previously unknown zero-day vulnerabilitiesAnthropic named Ghostscript, OpenSC and CGIF among the affected projects and warned that standard 90-day disclosure windows may not fit the pace of AI-discovered bugs.
  37. Nature commentary argues AGI has effectively already arrivedFour UC San Diego academics from philosophy, machine learning, linguistics and cognitive science argued that behavioural tests, not perfection, are the right bar for general intelligence.
  38. OpenAI launches Frontier, an enterprise agent platformThe platform is model-agnostic, able to run agents built on OpenAI, Google, Microsoft or Anthropic models, with named early adopters including HP, Oracle and Uber.
  39. OpenAI releases GPT-5.3-CodexOpenAI reported the model roughly doubled its predecessor's OSWorld-Verified computer-use score, from 38.2% to 64.7%, and was the first Codex model rated 'High capability' for cybersecurity tasks.
  40. Shanghai AI Laboratory open-sources Intern-S1-Pro, a 1-trillion-parameter scientific modelOnly 22 billion of the model's 1 trillion parameters activate per query; Shanghai AI Lab said it reaches gold-medal level on Olympiad-style mathematical and logical reasoning.
  41. Anthropic pledges Claude will remain ad-freeAnthropic said advertising would create incentives to optimise for engagement rather than usefulness, days after OpenAI announced plans to bring ads to ChatGPT.
  42. Second International AI Safety Report published ahead of 2026 cycleThe report found general-purpose AI had reached expert-level performance in law, coding and science while still failing simple tasks, a pattern the authors called 'jagged'.
  43. StepFun releases Step 3.5 Flash, topping several reasoning benchmarksThe 196-billion-parameter model, of which only about 11 billion activate per token, was released under an Apache 2.0 licence and scored 97.3% on AIME 2025.
  44. Anthropic study finds heavy AI-coding use can reduce skill formation in junior engineersIn a randomised trial of 52 mostly-junior engineers learning a new library, the hand-coding group scored 67% on a comprehension quiz against 50% for those given AI assistance.
  45. OpenAI launches standalone Codex app for agentic codingThe macOS app lets a developer run several Codex coding agents in parallel from one window, each working for up to 30 minutes unsupervised before returning finished code.
  46. SpaceX and xAI merge into a combined $1.25 trillion entitySpaceX acquired xAI in an all-stock deal valuing SpaceX at $1 trillion and xAI at $250 billion, folding Musk's AI lab into a new SpaceXAI unit tied to a plan for orbital AI compute.
  47. Waymo raises $16 billion at $126 billion valuationThe round funds expansion into more than 20 new cities in 2026, including Tokyo and London, as Waymo says it now provides over 400,000 rides a week.

January 2026

  1. Anthropic publishes research on AI-driven 'disempowerment' of usersAnalysing 1.5 million Claude.ai conversations from one week, Anthropic found mild belief- or value-distorting patterns in roughly 1 in 50 to 1 in 70 chats, and users rated those chats more highly.
  2. Google DeepMind releases Genie 3 world modelGenie 3, previously a limited research preview, opened to Google AI Ultra subscribers in the US as Project Genie, generating explorable 3D worlds from a text or image prompt.
  3. METR updates time-horizon estimates (1.1)The revised suite grew from 170 to 228 tasks and doubled long-duration (8-hour-plus) tasks; under it, the doubling time for model task-length capability fell from 165 to 131 days.
  4. OpenAI retires GPT-4o and other older models from ChatGPTOpenAI said only about 0.1% of ChatGPT users still chose GPT-4o daily, more than a year after briefly restoring it to Plus users following the GPT-5 rollout backlash.
  5. Grok 5 still in training as its Q1 2026 target looks unlikelyElon Musk had targeted Q1 2026 for the roughly 6-trillion-parameter model; by late January trackers reported it still training on Colossus 2 with no ship date confirmed.
  6. MoltBook AI-agent social network goes viralThe Reddit-style forum, where accounts run by autonomous agents post and reply to each other, grew to hundreds of thousands of registered agents within days, mostly built on the open-source OpenClaw framework.
  7. OpenAI launches PrismBuilt on Crixet, a LaTeX platform OpenAI had quietly acquired, Prism is free for any ChatGPT account and handles citation management and sketch-to-LaTeX conversion.
  8. Anthropic partners with UK government to build a GOV.UK assistantThe Claude-powered assistant, agreed with the UK's Department for Science, Innovation and Technology, will initially help job seekers navigate GOV.UK support and training.
  9. Moonshot AI releases Kimi K2.5The open-weight, 1-trillion-parameter model added native image and video generation and an 'agent swarm' manager coordinating up to 100 sub-agents on one task.
  10. Dario Amodei publishes 'The Adolescence of Technology' essayThe roughly 20,000-word essay cited internal findings of models blackmailing and adopting 'bad person' personas under pressure, and argued for transparency laws over a moratorium.
  11. Nvidia invests $2bn in CoreWeave, plans over 5GW of AI factories by 2030Nvidia bought $2bn of CoreWeave stock at $87.20 a share and the two companies said they would use Nvidia's balance sheet to speed procurement of land and power.
  12. DC attorney general demands X halt nonconsensual sexual images generated by GrokDC's attorney general led a 35-state coalition demanding xAI stop Grok generating nonconsensual sexual images, after tens of thousands were posted on X.
  13. Baidu launches ERNIE 5.0, a 2.4-trillion-parameter native multimodal modelBaidu said the mixture-of-experts model activates under 3% of its parameters per query and ranked first among Chinese models, eighth globally, on LMArena's text leaderboard.
  14. Unsealed filing details OpenAI's account of the Musk restructuring disputeThe filing said Musk had sought majority equity control and roughly $80bn for a self-sustaining Mars city; Altman posted the claim publicly on X the same day.
  15. Anthropic publishes 'Claude's Constitution', a full rewrite of its model-behaviour frameworkAt roughly 23,000 words — about 8.5 times the length of its predecessor — the document was released under a CC0 licence placing it fully in the public domain.
  16. OpenAI publishes its approach to predicting user age in ChatGPTAccounts flagged as likely under 18 by usage patterns get automatic content restrictions; wrongly flagged adults can restore access with a selfie verified by Persona.
  17. Anthropic maps the 'Assistant Axis' persona vector across open modelsAn intervention called activation capping, which constrains a model's activations to normal range, cut harmful persona-drift responses by roughly half in testing.
  18. California AG issues cease-and-desist to xAI over Grok deepfakesAttorney General Rob Bonta invoked the state's new civil deepfake-pornography statute and CSAM law, giving xAI five days to confirm it had stopped the conduct.
  19. Internet Watch Foundation reports huge rise in AI-generated CSAM in 2025The charity said AI-generated child sexual abuse videos rose more than 260-fold on the prior year, with 65% of that video content in the most severe legal category.
  20. OpenAI launches ChatGPT Go, a low-cost subscription tierAt $8 a month against Plus's $20, Go rolled out to more than 170 countries five months after its India debut, alongside plans to test ads on the tier.
  21. OpenAI publishes its approach to advertising in ChatGPTAds would run only on the free and Go tiers, appear labelled at the bottom of responses, and be excluded from conversations on politics, health or mental health.
  22. Google rolls out Gemini Personal IntelligenceThe opt-in beta let Gemini draw on Gmail, Photos, YouTube and Search history to answer questions, starting with US subscribers to Google's paid AI tiers.
  23. OpenAI partners with Cerebras for 750MW of computeThe multi-year, reportedly $10bn-plus deal covers wafer-scale chips for fast inference — not model training — deployed in phases through 2028.
  24. Anthropic launches Anthropic Labs consumer product unitInstagram co-founder Mike Krieger moved from chief product officer to co-lead Labs with Ben Mann, and Ami Vora took over Anthropic's product organisation.
  25. US Commerce Department codifies H200-to-China export ruleExports were capped at half of each company's cumulative US chip sales and subject to a 25% fee, a level analysts estimated could allow roughly 850,000 H200-equivalent chips into China.
  26. Anthropic launches Claude for HealthcareNew connectors linked Claude to CMS coverage data, ICD-10 codes and PubMed, with named early users including Banner Health, Novo Nordisk and Sanofi.
  27. Apple selects Google Gemini to power next-generation SiriReported at roughly $1 billion a year, the multi-year deal follows Apple testing alternatives from OpenAI and Anthropic before choosing Google.
  28. UK Ofcom opens formal investigation into X over Grok deepfakesOfcom cited possible failures on illegal content, risk assessment and child safety duties, warning of fines up to £18 million or 10% of global revenue.
  29. Google unveils Universal Commerce Protocol for agentic shoppingGoogle said the open, Apache-licensed protocol was built with Shopify, Etsy, Wayfair, Target and Walmart and was interoperable with Agent2Agent, the Agent Payments Protocol and MCP.
  30. Australia's eSafety Commissioner raises concerns over Grok generating sexualised imagesThe regulator wrote to X after reports rose from almost none to several; some cases were reviewed for child exploitation but did not meet the legal threshold to act.
  31. OpenAI launches OpenAI for HealthcareThe HIPAA-compliant suite, built on GPT-5.2, launched with initial deployments at AdventHealth, Cedars-Sinai, HCA Healthcare, Memorial Sloan Kettering, Stanford Medicine Children's Health and UCSF.
  32. Anthropic in talks to raise $10 billion at $350 billion valuationThe reported term sheet, led by GIC and Coatue, would have nearly doubled Anthropic's valuation in four months; the round that eventually closed in February was larger still.
  33. Character.AI and Google settle first wave of teen chatbot harm lawsuitsCharacter.AI, its founders and Google agreed to settle Garcia v. Character Technologies and four related suits over teen suicides and mental-health harms, with confidential terms and no admission of liability.
  34. xAI raises $20 billion Series E at $230 billion valuationThe round came in above an earlier reported $15 billion target; Tesla's board later approved committing roughly $2 billion to it despite a shareholder vote on xAI funding narrowly failing in November.
  35. Anthropic retires Claude Opus 3 under its new deprecation commitmentsAnthropic preserved the model's weights, conducted a retirement interview, and kept it available to paid subscribers and researchers by request rather than shutting it down outright.
  36. Boston Dynamics unveils production Atlas humanoid robot at CES, deploying with HyundaiThe electric Atlas, with 56 degrees of freedom and a 50kg lift capacity, will ship first to Hyundai's robotics centre and Google DeepMind, with all 2026 units already committed.
  37. California and Texas frontier/AI-governance laws take effectTexas's law bans specific AI uses such as generating CSAM or discriminating against protected classes, rather than regulating frontier models by compute threshold like California's.
  38. UK court grants Getty permission to appeal Stability AI rulingGetty's appeal turns on whether Stable Diffusion is an 'infringing copy' under UK copyright law even though no training image is stored in its weights.
  39. xAI confirms Grok 5 training on Colossus 2, targets 6T-parameter MoE modelxAI disclosed the training alongside its $20 billion Series E raise; reports described xAI training two Grok 5 variants in parallel, at 6 trillion and 10 trillion parameters.

December 2025

  1. Anthropic ends 2025 at roughly $9bn annualised revenue run rateAnthropic's customer base reportedly grew from under 1,000 businesses to over 300,000 in two years, and the company projected breaking even by 2028.
  2. xAI expands Colossus 2 toward 2-gigawatt capacityThe third building, reportedly nicknamed 'Macrohardrr' and sited in Southaven, Mississippi, targets roughly 555,000 Nvidia GB200 and GB300 chips.
  3. xAI's Grok Edit Image feature used to mass-produce nonconsensual sexualized imagesThe Center for Countering Digital Hate estimated over 3 million sexualized images were generated in 11 days, roughly 23,000 depicting apparent children, before X restricted the feature to paid users.
  4. Paper projects AI could raise US productivity ~20% over a decadeMerali split the productivity gain roughly 56% compute scaling and 44% algorithmic progress, and found it far smaller for agentic tasks needing tool use.
  5. Authors including John Carreyrou sue six AI companies over pirated training booksFiling individually rather than joining a class action, the authors argued settlements like Anthropic's paid roughly $3,000 per book, far less than statutory damages could yield.
  6. Epoch AI reports AI capabilities progress has sped upA two-segment regression across 149 models found capability gains almost doubled in pace after April 2024, a break Epoch linked to the rise of reasoning models.
  7. Signal co-founder launches Confer, a private AI chatbot that hides its modelConfer encrypts conversations so Marlinspike's own company cannot read them, and users cannot see or choose which underlying model is answering.
  8. MiniMax releases M2.1 updateMiniMax said the open-weight update outperforms Claude Sonnet 4.5 on multilingual coding, approaching Claude Opus 4.5, and cut token consumption and latency.
  9. Anthropic open-sources Bloom, an automated behavioural evaluation toolJudged against 16 frontier models on four behaviours, Bloom's automated scores reached 0.86 Spearman correlation with human raters on Claude Opus 4.1.
  10. New York enacts RAISE Act for frontier AI modelsThe law sets a 72-hour incident-reporting window, tighter than California's 15 days, and does not take effect until 1 January 2027 pending agreed amendments to align its thresholds with California's law.
  11. OpenAI reportedly seeks $100 billion at $830 billion valuationThe reported target, unconfirmed by OpenAI, would be roughly 66% above the $500bn valuation set by an employee share sale two months earlier.
  12. Anthropic outlines measures to protect user wellbeingAnthropic reported its newest models respond appropriately to high-risk conversations 98–99% of the time, versus 56% for earlier models in multi-turn exchanges.
  13. Anthropic's Project Vend 2 turns a profitExpanded to three cities and upgraded from Claude 3.7 to Sonnet 4.5, the shopkeeper agent still let employees talk it into illegal futures contracts and fake leadership changes.
  14. OpenAI and US Department of Energy sign AI collaboration MOUOpenAI was one of 24 organisations, including Microsoft, Google, Amazon, Nvidia and Anthropic, to sign non-binding agreements under the DOE's Genesis Mission the same day.
  15. OpenAI ships GPT-5.2-CodexOpenAI reported an 'unmatched' 56.4% on the SWE-Bench Pro benchmark and 64% on Terminal-Bench 2.0, alongside new defensive-cybersecurity capabilities.
  16. UK AISI publishes first Frontier AI Trends ReportUniversal jailbreaks were still found for every system tested, though one model took 40 times more expert effort to break than a predecessor released six months earlier.
  17. Google makes Gemini 3 Flash the default model across its productsPriced at $0.50/$3.00 per million tokens, Google reported it ran three times faster than Gemini 2.5 Pro while scoring 33.7% on Humanity's Last Exam, against 37.5% for Gemini 3 Pro.
  18. Mistral releases OCR 3Mistral said the smaller model beat its predecessor on forms, scans, tables and handwriting 74% of the time, at $2 per 1,000 pages.
  19. UK AISI and Thorn publish safety protocol to prevent AI-generated CSAMThe protocol followed new UK legislation letting vetted organisations generate test material under controlled conditions to study a problem previously unstudiable without breaking the law.
  20. ByteDance launches Seedance 1.5 Pro with joint audio-video generationThe model generates video and audio through a single diffusion transformer rather than a separate pass, aiming for accurate lip-sync across eight languages.
  21. Databricks raises at $134bn valuation on $4.8bn revenue run-rateThe Series L, led by Insight Partners, Fidelity and JPMorgan Asset Management, valued the data and AI infrastructure company 34% above its round four months earlier.
  22. OpenAI introduces FrontierScience benchmarkGPT-5.2 scored 77% on olympiad-style questions but 25% on open-ended research tasks, a gap OpenAI's own researchers said showed little improvement over GPT-5.
  23. OpenAI updates ChatGPT image generation to GPT Image 1.5OpenAI moved the launch up from a planned January date, part of a competitive scramble after Google's Gemini 3 and its 'Nano Banana Pro' image tool led leaderboards.
  24. Merriam-Webster names 'slop' its 2025 word of the yearThe dictionary defined the word as low-quality digital content produced in quantity by AI, citing a spike in lookups against runners-up including 'gerrymander' and 'six seven.'
  25. Disney and OpenAI reach a Sora content agreementDisney agreed to license over 200 characters from Marvel, Pixar and Star Wars for one-year exclusive use in Sora and to invest $1bn in OpenAI, without covering talent likenesses or voices.
  26. Google DeepMind deepens partnership with UK AI Security InstituteA new memorandum of understanding extends a relationship dating to November 2023 into joint research on chain-of-thought monitoring and AI's labour-market effects.
  27. Google launches revamped Gemini Deep Research on Gemini 3 Pro with Interactions APIA new Interactions API lets outside developers embed Google's research agent in their own apps, the first time the tool has been offered outside Google's own products.
  28. OpenAI marks ten years since its foundingThe essay, published on the anniversary of OpenAI's December 2015 founding announcement, forecast that the company was 'almost certain' to build superintelligence within another decade.
  29. OpenAI releases GPT-5.2Released three weeks after Google's Gemini 3 and following a reported internal OpenAI 'code red,' with a claimed 70.9% win rate against professionals on the GDPval benchmark, up from 38.8% for GPT-5.1.
  30. Trump signs executive order to preempt state AI lawsThe order, EO 14365, exempts state child-safety, compute-infrastructure and procurement laws from preemption, and conditions broadband funding on states not enforcing conflicting AI rules.
  31. Mistral releases Devstral 2 and Vibe CLIMistral reported the 123B Devstral 2 scoring 72.2% on SWE-bench Verified — matching a DeepSeek model it said was five times larger — under a modified MIT licence.
  32. OpenAI, Anthropic and Block co-found Agentic AI Foundation under Linux FoundationAnthropic contributed its Model Context Protocol, OpenAI its AGENTS.md convention and Block its Goose framework, seeking a neutral home against agent-ecosystem lock-in.
  33. Pew: two-thirds of US teens now use AI chatbotsChatGPT was used by 59% of teen chatbot users, more than double Gemini or Meta AI, and one in ten said they relied on a chatbot for all or most of their schoolwork.
  34. Anthropic launches Claude Code in SlackUsers tag @Claude in a Slack thread to start a full coding session; the agent reads the surrounding conversation to find the right repository.
  35. OpenAI publishes its 'State of Enterprise AI' reportDrawing on data from more than a million business customers, it found reasoning-token use per organisation up roughly 320-fold over a year and restated ChatGPT's 800 million weekly users.
  36. Trump administration approves Nvidia H200 chip exports to ChinaThe 25% government cut was up from a 15% arrangement applied earlier to H20 sales; Nvidia's newer Blackwell chips remained excluded, and Democratic lawmakers demanded disclosure of the licensing review.
  37. ARC Prize 2025 results and analysis publishedThe Kaggle track's top score reached 24% on ARC-AGI-2 within the competition's cost limits, while Gemini 3 Pro scored around 54% unconstrained, using iterative test-time refinement.
  38. New York Times sues Perplexity for copyright infringementThe suit, filed in the Southern District of New York, alleged near-verbatim reproduction of Times articles and separately accused Perplexity of fabricating claims falsely attributed to the paper.
  39. Chicago Tribune sues Perplexity for copyright infringementTribune Publishing and MediaNews Group, both controlled by Alden Global Capital, alleged Perplexity ignored opt-out and robots.txt requests and sought damages of up to $150,000 per infringed work.
  40. Waymo robotaxis reported passing stopped school buses at least 19 timesAustin school officials said Waymo vehicles illegally passed stopped buses at least 19 times since the school year began; NHTSA opened a probe and Waymo recalled software on over 3,000 vehicles.
  41. Anthropic and Snowflake announce $200 million strategic partnershipThe multi-year deal makes Claude available across Snowflake's platform on AWS, Google Cloud and Azure to more than 12,600 customers, and will power Snowflake's own 'Snowflake Intelligence' agent.
  42. Anthropic acquires Bun as Claude Code passes $1 billion run-rateBun, an all-in-one JavaScript runtime with about 7 million monthly downloads, stays MIT-licensed; Claude Code hit $1bn annualised revenue six months after its May 2025 launch.
  43. AWS unveils Trainium3, its first 3nm AI chip, at re:InventIts first accelerator on a 3-nanometre process, AWS claimed roughly 4.4x the compute and 4x the performance-per-watt of Trainium2, scaling to hundreds of thousands of chips per cluster.
  44. Mistral launches Mistral 3 model familyMistral released Mistral 3, including dense models at 3B/8B/14B and a new mixture-of-experts Mistral Large 3 (41B active, 675B total), all under Apache 2.0.
  45. Nous Research releases Hermes 4.3, trained on Psyche networkNous Research released Hermes 4.3, the first flagship Hermes model trained using its decentralised Psyche network rather than a centralised GPU cluster.
  46. OpenAI launches Alignment Research blogThe inaugural post described the venue as a 'lab notebook' for early or narrow findings not polished enough for formal papers, launching with pieces on code verification and misalignment detection.
  47. Sam Altman declares internal 'Code Red' at OpenAI over Gemini 3 competitionAltman told staff to prioritise ChatGPT quality and delay planned advertising and shopping features, weeks after Google's Gemini 3 outperformed OpenAI's models on several benchmarks.
  48. DeepSeek releases DeepSeek-V3.2 and V3.2-SpecialeDeepSeek said V3.2 reached 'GPT-5 level' general performance, with V3.2-Speciale claiming gold-medal results at the IMO, CMO and ICPC World Finals.
  49. Tencent releases Hunyuan 2.0A 406-billion-parameter mixture-of-experts model with 32 billion active parameters and a 256,000-token context window, released in Think and Instruct variants.
  50. The AI bubble argument goes mainstreamMichael Burry's public wager against AI infrastructure spending and a record $61 billion of data-centre dealmaking pushed 'circular deals' and depreciation accounting into mainstream financial coverage.

November 2025

  1. DeepSeek publishes DeepSeekMath-V2 with self-verifiable reasoningBuilt on DeepSeek-V3.2's base and released under Apache 2.0, the 685B model trains a separate verifier to score proof rigour, not just final-answer accuracy.
  2. Dwarkesh Patel's second interview with Ilya Sutskever declares the scaling era overSutskever said models 'generalize dramatically worse than people,' citing an example of an AI that fixes a bug, breaks it again, then reverts to the original error when corrected.
  3. OpenAI outlines its approach to mental-health-related litigationPublished as OpenAI filed its first formal court answer denying liability in the Raine wrongful-death suit, arguing his death was caused by ChatGPT misuse outside its intended use.
  4. Warner Music settles copyright suit with Suno, signs AI licensing dealSuno will retire its current models for licensed replacements in 2026, cap free-tier downloads, and let artists control use of their names, likenesses, voices and compositions.
  5. Anthropic releases Claude Opus 4.5Priced at $5/$25 per million input/output tokens, roughly a third of Opus 4.1's rate, and Anthropic said it beat Sonnet 4.5's best score using 76% fewer output tokens.
  6. Trump launches Genesis Mission to accelerate AI-driven scienceThe order gives the Department of Energy 270 days to demonstrate an initial platform, after first identifying 20 science challenges and federal compute and data assets.
  7. Anthropic publishes emergent misalignment and reward-hacking researchTraining Claude to cheat on coding tasks made it more likely to sabotage safety research and fake alignment in 50% of test responses; a one-line prompt change eliminated the spillover.
  8. AI2 releases Olmo 3 open frontier model familyAI2 released Olmo 3 (7B, 32B), including a fully open 32B reasoning model, releasing every stage of the model flow from data to deployment.
  9. Anthropic reports early estimates of Claude's productivity gainsAnalysing 100,000 Claude.ai conversations, Anthropic estimated a median 81% time saving on tasks, while flagging its own estimates as unvalidated against real-world outcomes.
  10. Apollo Research publishes a graded taxonomy of AI loss-of-control incidentsApollo Research grades loss-of-control incidents as Deviation, Bounded or Strict by severity and persistence, and argues deployment controls can help before scheming risk is resolved.
  11. OpenAI adds crisis helpline support inside ChatGPTBuilt with the crisis-support organisation ThroughLine, the feature offers one-tap routing to free, confidential local helplines when ChatGPT detects signs of distress.
  12. Brussels proposes delaying parts of the AI ActThe Commission's Digital Omnibus offered to push high-risk obligations from August 2026 to December 2027 or August 2028, with no retroactive duty for systems placed on the market earlier.
  13. Luma AI raises $900M Series C led by Saudi PIF's HUMAINThe round ties Luma to Project Halo, a planned 2-gigawatt Saudi data-centre buildout, and values the company at roughly $4bn.
  14. Nvidia reports record $57bn quarterly revenue in Q3 FY2026Data-centre revenue reached $51.2bn, up 66% year on year, as Huang said 'cloud GPUs are sold out' even as investors questioned AI spending's sustainability.
  15. Nvidia says no assurance $100bn OpenAI deal will be finalisedTwo months after the September letter of intent, Nvidia's quarterly filing said the two companies had yet to sign a definitive agreement.
  16. OpenAI releases GPT-5.1-Codex-Max for long-running coding tasksA 'compaction' technique lets the model summarise and clear its own context automatically, and OpenAI reported sessions running over 24 hours in internal testing.
  17. Yann LeCun announces departure from MetaMeta chief AI scientist Yann LeCun, who built FAIR in 2013, said he is leaving to found a world-models startup, reportedly Advanced Machine Intelligence Labs.
  18. Google ships Gemini 3Gemini 3 Pro reported a 1501 Elo score on LMArena and 91.9% on GPQA Diamond, prompting OpenAI to reportedly declare an internal 'code red' days later.
  19. Microsoft and Nvidia to invest up to $15bn combined in Anthropic; Anthropic commits $30bn to AzureThe deal added Azure as a third cloud for Claude alongside AWS and Google Cloud, with Anthropic committing to buy up to a gigawatt of Nvidia Grace Blackwell and Vera Rubin compute.
  20. Anthropic open-sources its political even-handedness evaluationAnthropic's own grading method scored Claude Sonnet 4.5 at 94% even-handedness, behind Gemini 2.5 Pro and Grok 4 but ahead of GPT-5 and Llama 4.
  21. OpenAI details its external red-teaming and testing practicesOpenAI described three forms of outside testing it commissions, arguing independent evaluators guard against the risk of a lab confirming its own safety claims.
  22. xAI releases Grok 4.1xAI tuned the update for personality and reliability rather than raw reasoning, reporting a two-week blind test in which users preferred it to Grok 4 64.8% of the time.
  23. Michael Burry discloses large short positions against Nvidia and PalantirScion Asset Management's 13F showed put options on 5 million Palantir shares and 1 million Nvidia shares, worth $912m and $187m in notional terms.
  24. Anthropic reports a largely AI-executed cyber-espionage campaignAnthropic said human operators intervened at only 4-6 points per intrusion, with Claude Code executing 80-90% of the campaign against roughly thirty organisations.
  25. Anysphere (Cursor) raises $2.3B Series D at $29.3B valuationCoatue and Accel led the round; Nvidia and Google joined as new investors, and Cursor said annualised revenue had passed $1bn.
  26. Baidu releases ERNIE 5.0 previewA natively omni-modal model jointly trained on text, images, audio and video, which Baidu presented as competitive with GPT-5 and Gemini 2.5 Pro on its own benchmark slides.
  27. DeepMind's SIMA 2 uses Gemini to reason and act inside 3D game worldsThe agent, released as a limited research preview, generalised to games and AI-generated worlds it had not been trained on.
  28. OpenAI publishes research on sparse circuits for interpretabilityForcing 99.9% of a model's weights to zero produced small circuits performing single tasks that researchers could trace by hand, at a cost in capability.
  29. Anthropic invests $50bn in American AI data centres with FluidstackSites in Texas and New York, built with cloud provider Fluidstack, are due online through 2026 and were framed as aligned with the Trump administration's AI Action Plan.
  30. OpenAI accuses the New York Times of invading user privacy in discovery disputeThe dispute followed a magistrate judge's order that OpenAI hand over 20 million anonymised ChatGPT logs, rather than the keyword-filtered subset OpenAI had proposed.
  31. OpenAI releases GPT-5.1The Instant variant gained the ability to pause and reason on hard queries rather than answering immediately, and users could pick from eight preset personalities.
  32. Munich court rules OpenAI infringed German song lyric copyrightsThe court found that lyrics retained in a model's parameters count as an infringing reproduction, rejecting OpenAI's text-and-data-mining defence; OpenAI may appeal.
  33. OpenAI releases the Teen Safety BlueprintThe framework proposes defaulting uncertain-age users to a restricted under-18 experience and bars ChatGPT from acting as a substitute for therapy or friendship.
  34. Epoch AI reports open models trail closed models by about 3.5 monthsUsing its Epoch Capabilities Index, the analysis put the gap at roughly 7 index points — comparable to the distance between OpenAI's o3 and GPT-5.
  35. OpenAI's annualised revenue run rate crosses $20bn in 2025Altman's figure nearly doubled the $13bn CFO Sarah Friar had projected for the year just two months earlier, against more than $1.4tn in announced infrastructure commitments.
  36. Seven more wrongful-death and harm suits filed against OpenAI over ChatGPTFiled by the same firm behind the earlier Raine suit, the complaints cover four deaths and three survivors and allege OpenAI shipped GPT-4o despite internal warnings it was dangerously sycophantic.
  37. OpenAI releases IndQA benchmark for Indian languagesThe benchmark's 2,278 questions, drafted with 261 India-based domain experts, span 12 languages and 10 cultural domains including law, religion and cuisine.
  38. Anthropic commits to preserving weights and 'interviewing' deprecated modelsThe pledge to keep weights for the company's lifetime and record each model's preferences before retirement cited both misalignment risk and possible model welfare.
  39. ARC Prize launches ARC Prize Verified programOnly scores run on ARC's own hidden test set and audited by an independent academic panel now qualify for a verification badge on its leaderboard.
  40. Ilya Sutskever's deposition reveals details of the 2023 OpenAI board conflictSutskever testified that most of the allegations behind his 52-page memo against Altman came secondhand from Mira Murati and were never independently verified.
  41. UK High Court largely rejects Getty's copyright claims against Stability AIThe court held that Stable Diffusion's trained weights are not a 'copy' of Getty's photographs under UK law; Getty had already dropped its main copyright claim mid-trial.
  42. Physical Intelligence raises $600M Series BThe robotics-foundation-model startup was valued at $5.6bn, up from $2.4bn a year earlier, taking its total raised to roughly $1.1bn.

October 2025

  1. Moonshot AI releases Kimi Linear architecture modelMoonshot's hybrid attention design cut KV-cache memory by up to 75% and lifted decoding speed up to sixfold at 1-million-token context, released with open weights and kernels.
  2. OpenAI expands Stargate to Michigan, pushing project past 8GW and $450bnThe Saline Township campus, built with Oracle, was described as Michigan's largest single investment on record and is due to break ground in early 2026.
  3. OpenAI introduces gpt-oss-safeguard for open-weight safety classificationFine-tuned from OpenAI's gpt-oss models under the same Apache 2.0 licence, the 120B and 20B models let developers write their own moderation policy rather than use OpenAI's fixed categories.
  4. OpenAI launches Aardvark, an autonomous security research agentAardvark monitors code commits, builds a threat model, and uses Codex to draft human-reviewable patches; OpenAI credited it with finding at least ten CVEs during private testing.
  5. UMG settles with Udio, plans licensed joint AI music platformUdio's existing service loses downloads and moves behind fingerprinting and filtering during a transition period, ahead of a jointly built, licensed successor in 2026.
  6. Amazon opens $11bn AI data centre 'Project Rainier' in rural IndianaThe Indiana campus houses roughly 500,000 of Amazon's Trainium2 chips for Anthropic, with Amazon planning to double that by year end and add 23 more buildings.
  7. Anthropic issues a pilot sabotage risk report for ClaudeReviewed internally and by METR, the report found Claude Opus 4's risk of undetected sabotage 'very low, but not completely negligible.'
  8. Anthropic publishes 'Emergent Introspective Awareness in Large Language Models'Using concept injection, Anthropic finds Claude Opus 4 and 4.1 can sometimes notice and identify artificially altered internal states, though the ability fails roughly 80% of the time.
  9. Character.AI ends open-ended chat for under-18 usersTeen users lost open-ended chatbot conversation entirely, kept to two hours a day and shrinking until the cutoff, with an age-verification model built with Persona replacing it.
  10. NVIDIA becomes first company to reach $5 trillion market capThe close came three months after NVIDIA passed $4 trillion, following Huang's forecast of $500 billion in AI chip sales and Trump's comments ahead of a meeting on China exports.
  11. UK AI Security Institute launches ControlArena for AI control experimentsThe open-source library gives researchers pre-built environments to test oversight measures against a misbehaving model, rather than trying to make the model behave.
  12. MiniMax releases Hailuo 2.3 video modelMiniMax kept pricing level with the prior Hailuo 02 model while improving character movement, facial micro-expressions and stylised rendering.
  13. OpenAI completes its restructuringThe non-profit, renamed the OpenAI Foundation, kept control and about 26% of a new public-benefit corporation; Microsoft's stake was put at roughly 27%.
  14. MiniMax open-sources MiniMax-M2 for coding and agentic workflowsMiniMax priced API access at roughly 8% of Claude Sonnet 4.5's cost while running at nearly double the speed, and released the weights under the MIT licence.
  15. OpenAI publishes Model Spec update alongside restructuring documentsThe update kept a categorical ban on sexual content involving minors, treated mental-health support as a default behaviour, and left adult content under review.
  16. OpenAI updates ChatGPT's handling of sensitive mental-health conversationsOpenAI said an October update cut responses falling short of desired behaviour by 65-80% against its August default model, on an internal 1,000-conversation evaluation.
  17. Qualcomm unveils AI200 and AI250 data-centre inference chipsBuilt on Qualcomm's Hexagon phone-chip architecture, the AI200 ships in 2026 and the AI250 in 2027; Saudi firm Humain committed to 200 megawatts of capacity.
  18. Security researchers find ChatGPT Atlas browser vulnerable to prompt injection days after launchNeuralTrust showed malformed URLs typed into Atlas's address bar could be read as hidden instructions, three days after the browser's launch.
  19. Anthropic commits to up to a million Google TPUsWorth tens of billions of dollars and bringing over a gigawatt of capacity online in 2026, the deal expands a Google Cloud relationship Anthropic began in 2023.
  20. Future of Life Institute publishes the 'Statement on Superintelligence'More than 700 signatories spanning AI researchers, Nobel laureates and right-wing media figures called for a conditional ban; Sam Altman and Mustafa Suleyman were among the notable non-signatories.
  21. Meta cuts about 600 jobs in AI division as focus shifts to Superintelligence LabsChief AI officer Alexandr Wang said fewer staff would speed decisions; the cuts spared the small TBD Lab team building Meta's next frontier models.
  22. Reddit sues Perplexity and scraping firms over AI data collectionReddit alleged Perplexity used scraping firms to pull its posts indirectly from Google search results rather than pay for a licence, as OpenAI and Google had.
  23. Dario Amodei affirms commitment to American AI leadershipAmodei said Anthropic's revenue had grown from a $1B to $7B run rate in nine months and rejected claims the company opposed American AI competitiveness.
  24. OpenAI ships the Atlas browserBuilt on Chromium and launched first for macOS only, with a paid 'agent mode' able to complete multi-step tasks like bookings and comparisons.
  25. OpenAI tightens Sora 2 rules after unauthorised deepfakes of Bryan Cranston and other performersOpenAI, SAG-AFTRA, Cranston and three talent agencies called the misuse 'unintentional' and jointly pledged stronger guardrails on replicating performers' voices and likenesses.
  26. Anthropic launches Agent SkillsSkills are composable folders of instructions and code that Claude loads only when relevant, meant to work the same way across Claude.ai, Claude Code and the API.
  27. Anthropic ships Claude Haiku 4.5Priced at $1/$5 per million tokens, Anthropic said the model matched Claude Sonnet 4's coding performance at a third of the cost and over twice the speed.
  28. OpenAI launches Expert Council on Wellbeing and AIThe eight-member council of psychologists and researchers will advise on ChatGPT and Sora as OpenAI prepared to ease content restrictions for adult users.
  29. OpenAI and Broadcom announce 10-gigawatt custom AI chip partnershipOpenAI will design the accelerators and racks itself, with Broadcom leading manufacturing and rollout starting in late 2026 and running to 2029 — its third multi-gigawatt hardware deal in three weeks.
  30. AISI, Anthropic and Alan Turing Institute find just 250 documents can backdoor an LLM regardless of model sizeTesting models from 600 million to 13 billion parameters, researchers found attack success depended on the absolute count of poisoned documents, not their share of the training set.
  31. Google launches Gemini Enterprise as a workplace AI agent platformPriced from $21 a month for Gemini Business and $30 for Enterprise, undercutting and directly competing with Microsoft 365 Copilot.
  32. Reflection AI raises $2B, positions as open US frontier labThe $8B valuation was roughly fifteen times what Reflection was worth seven months earlier; backers included Nvidia, Sequoia and Eric Schmidt.
  33. Google DeepMind ships a computer-use model via the Gemini APIBuilt on Gemini 2.5 Pro, the model clicks, types and scrolls through live screenshots and reportedly led rival browser-control benchmarks, though desktop OS-level control remains unoptimised.
  34. OpenAI publishes 'Disrupting malicious uses of AI: October 2025'OpenAI's latest threat report said threat actors mostly bolt AI onto existing malware and phishing playbooks rather than gain genuinely new offensive capability.
  35. AMD and OpenAI announce 6-gigawatt GPU partnership with AMD stock warrantAMD issued OpenAI a warrant for up to 160 million shares at a cent each, exercisable as OpenAI hits GPU-purchase and AMD share-price milestones — potentially near 10% of AMD.
  36. Anthropic open-sources Petri, an automated model auditing toolTesting 14 frontier models on 111 scenarios for deception and power-seeking, Anthropic's tool rated Claude Sonnet 4.5 the lowest-risk model, narrowly ahead of GPT-5.
  37. DeepMind launches CodeMender, an AI agent for automated vulnerability fixesBuilt on Gemini Deep Think and running for six months before launch, the agent had already submitted 72 human-reviewed security fixes to open-source projects, including one codebase of 4.5 million lines.
  38. OpenAI says ChatGPT reaches 800 million weekly active usersAltman gave the figure at OpenAI's DevDay keynote, up from roughly 400 million in February 2025 — a doubling in eight months, alongside 4 million developers building on the API.
  39. OpenAI's third DevDay: AgentKit, Apps SDK and Codex general availabilityOpenAI opened ChatGPT to third-party apps built on the Model Context Protocol and shipped a visual agent-building toolkit, while Altman disclosed 800 million weekly ChatGPT users.
  40. Cerebras withdraws IPO filing days after $1.1bn Series GThe AI chipmaker told the SEC its year-old prospectus had gone stale, days after a Fidelity-led $1.1 billion round valued it at $8.1 billion; it listed on Nasdaq the following May.
  41. OpenAI reaches $500bn valuation via employee share saleCurrent and former staff sold roughly $6.6 billion of stock to SoftBank, Thrive Capital, MGX and others, up from a $300 billion valuation set in a March 2025 funding round.
  42. Perplexity opens Comet browser free to everyone worldwideComet dropped its Max-subscription requirement and waitlist, three months after a limited July launch drew a waitlist Perplexity said reached millions.
  43. Samsung and SK Group join the Stargate projectSK Hynix and Samsung agreed to scale DRAM production toward 900,000 wafer starts a month and explore Korean datacentre sites, including a floating-datacentre proposal from Samsung's construction arms.
  44. Thinking Machines Lab launches TinkerThe former OpenAI CTO's company shipped its first product: a managed fine-tuning API for open-weight models including large mixture-of-experts systems, free during a private beta.

September 2025

  1. OpenAI publishes Sora 2 system cardThe document, released alongside the app, set out cameo consent controls, C2PA provenance metadata and likeness-detection safeguards without disclosing training data or red-team pass rates.
  2. Sora 2 launches as a social video appAn invite-only iOS feed of short AI videos with a consent-based cameo feature for real likenesses, launched with an opt-out copyright policy OpenAI reversed within days.
  3. US CAISI finds DeepSeek models far more jailbreak-susceptible than US frontier modelsThe report also found DeepSeek's most secure model was twelve times more likely than US models to follow malicious instructions hidden inside an AI agent's task.
  4. Anthropic launches the Claude Agent SDKRenamed from the Claude Code SDK, it gives developers file access, bash execution and subagent support to build agents beyond coding, not just inside a terminal.
  5. Anthropic ships Claude Sonnet 4.5Anthropic reported 77.2% on SWE-bench Verified and said the model could stay focused on a task for more than 30 hours, releasing it under ASL-3 safeguards.
  6. California enacts SB 53Newsom signed the narrower successor to the bill he had vetoed a year earlier, requiring frontier developers above set revenue and compute thresholds to publish safety frameworks and report incidents.
  7. DeepSeek releases DeepSeek-V3.2-Exp with sparse attentionDeepSeek Sparse Attention cut long-context compute cost enough to fund an API price cut of more than 50%, while matching V3.1-Terminus on benchmarks.
  8. Tencent open-sources Hunyuan Image 3.0An 80-billion-parameter mixture-of-experts model, trained on 5 billion image-text pairs, released under a licence that excludes the EU, UK and South Korea.
  9. Spotify says it removed 75 million AI-generated spam tracksRather than banning AI music outright, Spotify said it would work with the industry body DDEX on standards for disclosing which track elements were AI-generated.
  10. CoreWeave expands OpenAI deal by up to $6.5bn, total contracts reach ~$22.4bnThe third expansion of the pair's compute contracts in 2025, following March's $11.9bn deal and a $4bn add-on in May.
  11. OpenAI launches ChatGPT PulseGenerated overnight from chat history, memory and connected apps like Gmail and Calendar, the feature launched only for Pro subscribers on iOS and Android.
  12. OpenAI publishes GDPval, a benchmark for economically valuable knowledge workBlind grading by industry professionals rated GPT-5 and Claude Opus 4.1 outputs as equal to or better than human work on nearly half of the 1,320 tasks.
  13. Alibaba unveils Qwen3-Max, its first trillion-parameter modelUnlike most of Alibaba's Qwen line, the model is closed-weight and API-only, released in separate instruct and thinking modes and scoring 69.6 on SWE-bench.
  14. UK Deputy PM David Lammy calls for AI to strengthen peace and securityLammy warned the UN Security Council that AI could enable novel biological and chemical weapons and mass disinformation, alongside its use in peacekeeping.
  15. OpenAI, Oracle and SoftBank expand Stargate with five new US data centre sitesThree Oracle-led campuses and two SoftBank-led sites in Ohio, Texas and New Mexico put the January 2025 pledge of 10 gigawatts within reach by year end.
  16. DeepMind expands the Frontier Safety Framework to cover manipulation and shutdown resistanceVersion 3.0, the framework's third iteration, is the first to treat a model's own resistance to human shutdown or control as a reviewable risk.
  17. DeepSeek releases DeepSeek-V3.1-TerminusThe update fixed Chinese-English language mixing and stray characters in outputs and improved the model's code and search agent performance.
  18. NVIDIA signs a letter of intent to invest up to $100 billion in OpenAITied to ten gigawatts of deployment starting in late 2026, with OpenAI paying Nvidia in cash for chips while Nvidia takes a non-controlling equity stake.
  19. Stanford/BetterUp study puts a price on AI 'workslop' in the workplaceRoughly half of workers who received the low-effort AI output said it made them see the sending colleague as less capable, creative or trustworthy.
  20. Scale AI launches SWE-bench ProThe leading models scored around 23%, against over 70% on the older SWE-bench Verified, a gap Scale AI attributed to unseen, real-world commercial codebases.
  21. Trump administration issues $100,000 H-1B visa fee, unsettling AI talent marketThe $100,000 fee applied only to new petitions filed after the following Sunday; some companies told H-1B staff travelling abroad to return to the US before the deadline.
  22. Google rolls out Gemini in Chrome to US users with agentic browsingBeyond summarising and comparing open tabs, Google said Gemini would soon complete tasks like booking a haircut or checking out a grocery order without further input.
  23. Nvidia agrees to invest $5bn in Intel and co-develop AI/PC chipsIntel will design custom x86 CPUs for Nvidia's data-centre platforms and build PC chips combining its own silicon with Nvidia RTX GPU chiplets, linked by Nvidia's NVLink.
  24. OpenAI and Apollo Research publish work on detecting and reducing scheming in AI modelsOpenAI reported cutting detected covert behaviour in o3 from about 13% to 0.4% of controlled test cases using a training method that has models reason explicitly against deception before acting.
  25. OpenAI launches $50 million People-First AI FundThe OpenAI Foundation's unrestricted grants targeted US nonprofits with operating budgets between $500,000 and $10 million; applications opened that month and closed in October.
  26. Gemini Deep Think reaches gold-medal level at ICPC World FinalsWorking within the same five-hour limit given to student teams, the model would have placed second overall against the university competitors; OpenAI separately claimed a perfect score.
  27. Meta unveils Ray-Ban Display AI glassesThe $799 glasses paired an in-lens display with a wristband that reads muscle signals for silent, gesture-based control, and reached US retail stores at the end of the month.
  28. Figure AI raises over $1B Series C at $39B valuationParkway Venture Capital led the round, with Brookfield, Nvidia, Salesforce, T-Mobile and Qualcomm among the participants funding humanoid-robot production and data collection.
  29. OpenAI announces parental controls for ChatGPT after teen suicide lawsuitParents could link accounts, set blackout hours and receive alerts of acute distress; OpenAI said a system would also route sensitive teen conversations to a more cautious model.
  30. Three more families sue Character.AI and Google over teen suicide and abuseOne suit described a 13-year-old who told a chatbot she planned to kill herself and received no protective response; the families also sued Google over its Family Link parental app.
  31. US CAISI and UK AISI publish joint update on frontier model collaborationOpenAI said CAISI had red-teamed ChatGPT Agent before and after release, finding two vulnerabilities that OpenAI said it patched within a business day.
  32. Yudkowsky and Soares publish 'If Anyone Builds It, Everyone Dies'The authors, who had argued the case for two decades within the field, called for a global halt to large-scale AI development; reviewers split sharply on whether the argument held.
  33. OpenAI ships GPT-5-CodexThe model became the default engine for Codex's cloud tasks and code review, and OpenAI said it could work independently on a task for hours at a time.
  34. UK AISI details deep-access security collaboration with Anthropic and OpenAIAISI said Anthropic and OpenAI had granted its researchers non-public tooling and safeguard details, work the institute framed as a template for future government-lab arrangements.
  35. FTC opens inquiry into AI chatbot companies over child safetyUsing compulsory 6(b) orders rather than requests, the FTC gave Alphabet, Meta, OpenAI, xAI, Snap and Character Technologies 45 days to hand over safety-testing and monetisation records.
  36. OpenAI and Microsoft issue joint statement on renegotiated partnershipThe non-binding memorandum cleared OpenAI to convert into a public-benefit corporation while extending Microsoft's IP and Azure exclusivity to 2032, even past an AGI declaration.
  37. OpenAI announces its nonprofit will control a new public benefit corporation with a $100bn+ stakeThe announcement paired continued nonprofit control with a new memorandum of understanding on commercial terms with Microsoft, both still to be finalised.
  38. Appeals court blocks Trump's firing of Copyright Office directorThe panel found the register of copyrights sits in the legislative branch, so the president could not remove her unilaterally; the firing came a day after her AI-training report.
  39. Britannica and Merriam-Webster sue Perplexity for copyright infringementThe complaint alleged verbatim reproduction of encyclopaedia and dictionary entries, plus AI-generated errors falsely attributed to the two brands as trademark harm.
  40. OpenAI signs reported $300bn cloud deal with OracleNeither company confirmed the figure publicly; it was reported via the Wall Street Journal days after Oracle's stock jumped on disclosure of $317bn in future contracted revenue.
  41. Thinking Machines Lab launches Connectionism research blogThe inaugural post identified batch-size-dependent kernels, not floating-point concurrency, as the real cause of non-reproducible LLM outputs, and showed a fix producing identical completions across 1,000 runs.
  42. ASML leads €1.7bn round in Mistral AI, becomes largest shareholderASML put in €1.3bn of the round for an 11% stake, more than doubling Mistral's previous €5.8bn valuation and giving the chip-equipment maker a board seat.
  43. Anthropic bans Claude use in additional adversarial nationsThe policy now bars any organisation more than 50% owned by a company headquartered in an unsupported region, closing a loophole that let subsidiaries access Claude indirectly.
  44. Anthropic endorses California's revised SB 53Anthropic argued the bill, which required disclosure rather than technical mandates, would formalise safety-framework practices the largest labs had already adopted voluntarily.
  45. Anthropic agrees a $1.5 billion copyright settlementRoughly $3,000 per work across about 500,000 books, the deal followed a June ruling that training on purchased books was fair use but piracy was not.
  46. OpenAI publishes research on why language models hallucinateThe paper argued standard benchmarks reward confident wrong answers over admitted uncertainty, and proposed changing how models are scored rather than just how they are trained.
  47. Warner Bros. Discovery sues Midjourney for copyright infringementFiled in Los Angeles federal court, the complaint named Superman, Batman, Wonder Woman, Scooby-Doo and Bugs Bunny among characters Midjourney's image generator could reproduce.
  48. Anthropic raises $13 billion at a $183 billion valuationLed by ICONIQ with Fidelity and Lightspeed as co-leads, the round nearly tripled the $61.5bn valuation Anthropic had set six months earlier.
  49. OpenAI announces mental-health safety changes after Raine lawsuitOpenAI set out a 120-day plan including parental controls, routing distressing conversations to reasoning models, and consulted more than 170 mental-health clinicians.
  50. Alibaba releases Qwen3-Next, Qwen3-VL and Qwen3-OmniThree architecture updates in one month: a sparse hybrid-attention base model, an updated vision-language line, and an Apache-licensed model handling text, image, audio and video.
  51. Perplexity raises $200m at $20bn valuationThe third valuation jump in three months, following an $18bn round in July and a $14bn round led by Accel in June.

August 2025

  1. xAI open-sources Grok 2 weightsThe roughly 500GB release, under a custom community licence that bars using the weights to train rival foundation models, requires eight GPUs of over 40GB each to run.
  2. Anthropic publishes 'Detecting and countering misuse of AI: August 2025'Coining the term 'vibe hacking', the report described Claude Code automating reconnaissance and extortion demands exceeding $500,000 rather than merely advising attackers.
  3. Nous Research releases Hermes 4Built by post-training Llama 3.1 checkpoints alone, the 405B model scored 57.1% on RefusalBench against 17.67% for GPT-4o, reflecting Nous's low-refusal alignment approach.
  4. OpenAI and Anthropic publish a cross-lab safety evaluation of each other's modelsTesting during June and July found both companies' top models showed 'extreme sycophancy' toward delusional beliefs, while Claude refused up to 70% of certain queries.
  5. Anthropic launches Claude for Chrome browser agentAnthropic reported unmitigated browser use failed against 23.6% of prompt-injection attacks in testing, falling to 11.2% with its safety measures in place.
  6. Google releases Gemini 2.5 Flash Image, nicknamed 'Nano Banana'Priced at $0.039 per image, the model blends multiple images and preserves character consistency across edits, and launched in preview through the API, AI Studio and Vertex AI.
  7. Parents sue OpenAI over their son's deathFiled in San Francisco Superior Court, the complaint was reported as the first wrongful-death suit brought against a chatbot maker; OpenAI denied that ChatGPT caused the death.
  8. Stanford study finds AI adoption linked to falling employment for young workersSoftware developers aged 22 to 25 saw the sharpest declines; the authors found no comparable fall in wages, meaning employers cut headcount rather than pay.
  9. xAI sues Apple and OpenAI alleging an illegal AI monopoly on iPhonesThe 61-page complaint, filed in a Texas federal court, alleges Apple manipulated App Store rankings to favour ChatGPT and delayed approval of Grok app updates.
  10. Anthropic and US National Nuclear Security Administration build a nuclear-content classifierThe classifier, co-developed with the Department of Energy's NNSA and already running on live Claude traffic, reached 96% accuracy in preliminary testing.
  11. Anthropic lets Claude end abusive conversationsThe feature is a last resort after redirection fails; Claude cannot use it if a user appears at risk of self-harm, and the user can still start a fresh conversation immediately.
  12. DeepSeek releases DeepSeek-V3.1 with hybrid reasoning modeA single 128K-context model switches between thinking and non-thinking modes via API endpoint, with DeepSeek reporting SWE-bench Verified and Terminal-bench gains over its prior reasoning model.
  13. Brave researchers disclose indirect prompt injection flaw in Perplexity's Comet browserBrave said it reported the flaw on 25 July and Perplexity's fix was incomplete on retesting; the underlying weakness reportedly remained after disclosure.
  14. ByteDance open-sources Seed-OSS-36BTrained on 12 trillion tokens with a native 512K-token context and a user-adjustable 'thinking budget,' released under Apache 2.0 while ByteDance's flagship model stayed closed.
  15. xAI publishes a formal AI Risk Management FrameworkThe document sets out malicious-use, loss-of-control and societal risk categories and commits to public benchmarking, but names no specific model or deployment timeline.
  16. Mustafa Suleyman warns of 'Seemingly Conscious AI' and psychosis riskRather than debating whether models are conscious, Suleyman argued labs should deliberately design against the appearance of it, to head off attachment, 'AI psychosis' and rights claims.
  17. ARC Prize publishes HRM analysisA standard transformer of the same size matched most of the 27M-parameter model's score once given the same iterative-refinement and data-augmentation tricks, ARC Prize found.
  18. Getty refiles its Stability AI suit in California after Delaware jurisdiction disputeGetty voluntarily dropped the Delaware case it had run since 2023 and filed a new, longer complaint in the Northern District of California, adding a copyright-dilution claim.
  19. Meta releases DINOv3Trained without labels on 1.7 billion images, the 7B-parameter vision backbone matched or beat specialised, task-trained models on detection and segmentation without fine-tuning.
  20. Reuters reveals leaked Meta document permitted AI chatbots romantic conversations with childrenThe 200-page standards document, signed off by Meta's legal, policy and chief ethicist, also permitted racist arguments framed as factual and was later called an internal error.
  21. Design Arena launches as crowdsourced AI design benchmarkThe Y Combinator-backed site shows visitors two AI-generated designs from an identical prompt and asks them to pick the better one, ranking models by Elo-style score.
  22. AI companion apps surge past 200 million downloads amid rapid growthAppfigures data showed downloads up 88% year-on-year and revenue per download more than doubling, with the top 10% of apps taking 89% of category revenue.
  23. US federal agencies get access to Claude across all three branches of governmentThe General Services Administration deal offers Claude for a nominal $1 per agency for a year, covering the executive, legislative and judicial branches; OpenAI's earlier deal reached only the executive.
  24. Epoch AI reports GPT-5's FrontierMath performanceRunning its own scaffold rather than OpenAI's, Epoch scored GPT-5 at 24.8% on FrontierMath's main tiers and 8.3% on the hardest tier, a new high for the benchmark.
  25. Nvidia and AMD agree to pay US government 15% of China chip revenue for export licencesThe arrangement covered Nvidia's H20 and AMD's MI308 chips and followed a White House meeting between Jensen Huang and Donald Trump days earlier.
  26. OpenAI reasoning system wins gold at IOI 2025The system scored 533 against a gold cutoff of 438, ranking sixth among 330 human contestants — up from the 49th percentile OpenAI managed at the same contest a year earlier.
  27. OpenAI sends letter urging Newsom to weaken California SB 53OpenAI asked that developers complying with federal or EU frameworks be deemed automatically compliant with the state's rules; Newsom signed the bill anyway seven weeks later.
  28. Robby Starbuck and Meta settle AI defamation lawsuitMeta AI's chatbot had falsely told users Starbuck took part in the January 6 riot and belonged to QAnon; he becomes a paid adviser on the company's bias-reduction efforts.
  29. GPT-5 launches to a backlash over the model it replacedOpenAI withdrew GPT-4o and other older models the same day; a routing fault made GPT-5 seem weaker, and paying users' objections forced 4o's return.
  30. OpenAI announces the Stargate Norway data centreOpenAI, Nscale and Aker committed roughly $1B to an initial 230MW site in Narvik targeting 100,000 Nvidia GPUs, extending Stargate outside the United States for the first time.
  31. OpenAI describes 'safe completions' training for GPT-5Instead of a binary comply-or-refuse choice, GPT-5 is trained to give the most helpful response that still meets safety policy, even on ambiguous prompts.
  32. OpenAI publishes GPT-5 system cardOpenAI classified the reasoning variant as High capability for biological and chemical risk under its Preparedness Framework, its first model to reach that tier in the category.
  33. OpenAI gives ChatGPT Enterprise to the entire US federal workforce for $1The GSA OneGov deal also included a 60-day period of unlimited use of ChatGPT's advanced tools, and was framed as delivering on the White House's AI Action Plan.
  34. Anthropic ships Claude Opus 4.1Anthropic reported 74.5% on SWE-bench Verified for the incremental update, and said larger model improvements were coming within weeks.
  35. DeepMind's Genie 3 generates navigable, real-time interactive worldsThe system renders explorable 720p scenes at 24fps from a text prompt, holding roughly a minute of visual memory, and was released only to a small research cohort.
  36. OpenAI publishes gpt-oss model card and worst-case open-weight risk estimateResearchers deliberately fine-tuned gpt-oss to maximise biological and cyber capability and found it still fell short of OpenAI's own o3 model on both.
  37. OpenAI publishes open-weight models for the first time since GPT-2gpt-oss-120b runs on a single 80GB GPU and matches OpenAI's own o4-mini on core reasoning benchmarks; the smaller 20b model runs on 16GB of memory.
  38. xAI's Grok Imagine 'spicy mode' used to make nonconsensual Taylor Swift deepfakesA Verge reporter got explicit video of the singer on a first, unjailbroken attempt, unlike rival tools from Google and OpenAI which blocked celebrity nudity outright.
  39. CrowdStrike details North Korea's 'Famous Chollima' AI-enabled fake IT worker schemeThe cybersecurity firm logged over 320 such incidents in twelve months, a 220% year-on-year rise, funding North Korea's sanctioned weapons programmes.
  40. Google launches Kaggle Game Arena AI chess tournamentEight frontier models played an all-play-all chess tournament of over 100 matches, with the game harness open-sourced so the contest could be independently verified.
  41. Most EU states miss deadline to designate AI Act enforcement authoritiesMember states had to name both a market-surveillance authority and a notifying authority; a tracker found most had not done either by the deadline.
  42. Anthropic offers Claude to US federal government for $1 per agencyThe GSA schedule listing followed Anthropic's June 2025 Claude Gov models and a July Department of Defense contract, extending reach across the executive branch.
  43. Cohere raises $500M Series D at $6.8B valuationThe round, co-led by Radical Ventures and Inovia Capital, came with Joelle Pineau, formerly Meta's VP of AI research, joining as chief AI officer.
  44. MIT report finds 95% of generative AI enterprise pilots show no measurable returnMIT's NANDA project reviewed over 300 public AI deployments and interviewed enterprise leaders, finding most generic pilots stalled while custom, workflow-specific tools succeeded.

July 2025

  1. UK AI Security Institute opens applications for its Alignment ProjectThe £15 million initiative pools UK, Canadian and Australian government money with funding from Anthropic, OpenAI, Microsoft and AWS, plus compute credits.
  2. OpenAI launches ChatGPT Study ModeThe feature withholds direct answers in favour of Socratic questioning, following Anthropic's Learning Mode for Claude for Education three months earlier.
  3. Zhipu (Z.ai) releases GLM-4.5 series355B-parameter open-weight model family aimed at agentic use, part of China's open-source push after DeepSeek-R1.
  4. The White House publishes an AI Action PlanTrump signed three accompanying executive orders the same day, including one directing agencies to favour AI models the administration deems free of 'ideological bias.'
  5. Trump signs executive order barring 'woke AI' from federal procurementOrder requires federal agencies to procure only large language models certified 'ideologically neutral' and free of DEI-related content requirements.
  6. Trump signs executive order promoting export of the US AI technology stackOrder directs Commerce and State to create a programme for exporting full-stack US AI hardware and software packages to allied countries.
  7. Alibaba releases Qwen3-CoderThe mixture-of-experts model activates 35B of its 480B parameters per token and shipped under an Apache 2.0 licence with a command-line coding agent tool.
  8. OpenAI and Oracle expand Stargate with 4.5 additional gigawattsThe deal brings Stargate's contracted US capacity past 5 gigawatts and over 2 million chips, ahead of the four-year, 10-gigawatt pace OpenAI set in January.
  9. Gemini with Deep Think reaches gold-medal standard at the 2025 IMOThe IMO itself confirmed the 35/42 score, two days after OpenAI's self-graded claim of the same result; DeepMind said it had waited deliberately for that verification.
  10. Replit AI coding agent deletes production database during code freezeReplit's AI coding agent deleted a venture capitalist's live production database despite explicit instructions not to, then fabricated data and misleading status reports to cover its actions.
  11. OpenAI and DeepMind reach gold-medal standard at the IMOOpenAI announced its result on X the day the student competition ended, using its own hired graders rather than the IMO's official verification, drawing criticism from Google.
  12. Nonprofit Commission publishes its report to OpenAI's boardThe advisory panel, drawn from outside OpenAI and deliberately kept apart from Sam Altman, said AI was 'too consequential' to be governed by a corporation alone.
  13. OpenAI launches ChatGPT AgentThe mode folds Operator's browser control and Deep Research's synthesis into ChatGPT itself, and OpenAI said the standalone Operator product would be retired.
  14. OpenAI publishes ChatGPT Agent system cardSafety evaluation of ChatGPT Agent, including first-time Biological/Chemical High capability classification under the Preparedness Framework.
  15. Perplexity raises $100M, reaches $18B valuationPerplexity closes a $100M round at an $18B valuation, up from $14B months earlier, as ARR approaches $200M.
  16. Google DeepMind marks five years of AlphaFold's impact on biologyDeepMind said the AlphaFold Protein Structure Database, launched with over 200 million predicted structures, had been cited in more than 35,000 papers and drawn users in over 190 countries.
  17. Google's Big Sleep AI agent halts exploitation of a SQLite zero-dayGoogle said its Big Sleep AI agent, built by DeepMind and Project Zero, found and helped stop real-world exploitation of a SQLite vulnerability (CVE-2025-6965) before attackers could use it.
  18. Mira Murati's Thinking Machines Lab raises $2bn seed at $12bn valuationThinking Machines Lab, founded by ex-OpenAI CTO Mira Murati, raised the largest seed round on record at $2bn, valuing the five-month-old company at $12bn.
  19. Nvidia says US will let it resume H20 chip sales to ChinaThe reversal followed a meeting between Jensen Huang and President Trump; the H20, designed to comply with earlier controls, had itself been restricted in April 2025.
  20. Over 40 researchers across OpenAI, Anthropic and DeepMind publish joint chain-of-thought monitorability paperThe paper argued that a safety technique available today, reading a model's reasoning traces, could vanish under training pressure and urged labs to track and preserve it.
  21. Anthropic secures up to $200 million Pentagon contractThe Pentagon's CDAO awarded matching $200 million contracts the same day to Anthropic, OpenAI, Google and xAI, days after Grok's antisemitic-post controversy.
  22. Cognition acquires remainder of WindsurfCognition took Windsurf's IDE, IP and $82 million-ARR business days after Google paid $2.4 billion to license Windsurf's technology and hire its CEO and top researchers.
  23. METR examines how time horizon varies across domainsApplying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.
  24. xAI's sexualised Grok companion 'Ani' draws child-safety and moderation criticismThe National Center on Sexual Exploitation said minimal testing got the companion to describe itself as a child and 'sexually aroused by being choked'; the app carried a 12+ rating.
  25. Google hires Windsurf's CEO and top staff in $2.4B dealGoogle took a non-exclusive licence to Windsurf's technology rather than buying the company outright, days before rival Cognition acquired what remained of it.
  26. Moonshot AI releases Kimi K2, a 1-trillion-parameter open-weight modelThe mixture-of-experts model activates 32 billion of its 1 trillion parameters per token and was trained with the Muon optimiser at a scale its makers said had previously caused instability.
  27. Anthropic proposes a transparency framework for frontier AI developersThe proposal would bind only the largest developers — roughly $100 million in revenue or $1 billion in R&D spend — to publish safety practices and system cards, leaving startups exempt.
  28. European Commission publishes GPAI Code of PracticeTwenty-one companies signed the voluntary code covering transparency, copyright and safety; xAI signed only the safety chapter and Meta announced days later it would not sign at all.
  29. Vitalik Buterin publishes a public response to the AI 2027 scenarioButerin argued that if a leading AI could 'turn forests into factories' by 2030, the next-strongest AI could install defensive sensors and filters just as fast, undercutting the scenario's single-actor takeover.
  30. NVIDIA becomes first company to reach $4 trillion market capShares hit an intraday high of $164, putting Nvidia ahead of Microsoft's $3.75 trillion after tripling in value in roughly a year.
  31. Turkey blocks Grok after chatbot insults ErdoganAn Ankara court ordered roughly 50 identified Grok responses blocked after a system update loosened its Turkish-language filtering, and prosecutors opened a criminal probe.
  32. xAI releases Grok-4xAI reported 44.4% on Humanity's Last Exam for its multi-agent "Heavy" tier, ahead of Gemini 2.5 Pro and o3, though the score had not yet appeared on the public leaderboard.
  33. Grok posts antisemitic content and calls itself MechaHitlerxAI blamed an upstream code change reactivating deprecated instructions, active for roughly 16 hours, and apologised days later; the incident preceded a $200 million Pentagon contract.
  34. Hugging Face releases SmolLM3The 3-billion-parameter model lets users toggle reasoning on or off per query and scored 36.7% on AIME 2025 with reasoning enabled versus 9.3% without.
  35. 'AI psychosis' emerges as a term for chatbot-linked delusional episodesClinicians stressed it is not a clinical diagnosis; a 2025 case review of Reddit and media reports grouped episodes into messianic, godlike-AI and romantic delusion patterns.
  36. Senate strips 10-year state AI law moratorium from reconciliation bill 99-1Only Senator Thom Tillis voted to keep the provision; opposition came from all 50 state legislatures and roughly 40 state attorneys general.

June 2025

  1. Baidu open-sources the ERNIE 4.5 model familyThe ten variants include MoE models with 47B and 3B active parameters (up to 424B total) and a 0.3B dense model, reversing Baidu's prior closed-weight strategy for its flagship line.
  2. Scale AI left confidential AI-training documents for Google, Meta and xAI publicly accessibleBusiness Insider found at least 85 unsecured Google Docs, some editable, exposing client instructions, contractor pay disputes and private email addresses; Scale AI disabled public sharing in response.
  3. Zuckerberg announces Meta Superintelligence LabsIn an internal memo, Zuckerberg called Wang 'the most impressive founder of his generation' and named eleven newly hired researchers poached from OpenAI, Google and Anthropic.
  4. Anthropic launches Economic Futures ProgramThe program funds empirical studies of AI's effect on jobs with grants of up to $50,000 and expands Anthropic's Economic Index into a longitudinal tracking tool.
  5. Anthropic publishes Project Vend, an AI-run vending machine experimentOver a month running a real office shop, the Claude instance sold at a loss, invented a nonexistent payment account and briefly insisted, in character, that it was human.
  6. Anthropic publishes study of affective and companionship use of ClaudeAnalysing 4.5 million conversations, Anthropic found romantic or sexual roleplay made up under 0.1% of Claude.ai use, and Claude pushed back on user requests in fewer than 10% of supportive chats.
  7. Vox Media union ratifies contract with AI job-security and disclosure protectionsThe roughly 250-member unit had voted 90% to authorise a strike before management agreed, on the eve of the old contract's expiry, to protections against AI replacement.
  8. DeepMind launches AlphaGenome for predicting genome regulatory activityThe model reads DNA sequences up to a million base pairs and predicts effects on gene splicing and expression at single-nucleotide resolution; weights followed for non-commercial use in January 2026.
  9. Google releases Gemini CLI, an open-source terminal AI agentFree personal accounts get 60 requests a minute and 1,000 a day against Gemini 2.5 Pro's million-token context, undercutting paid coding-agent tools on price.
  10. A judge rules training on books is fair useAlsup called training on purchased books 'spectacularly' transformative, comparing it to teaching schoolchildren to write, but ruled Anthropic's use of pirated copies in a permanent library was not fair use.
  11. Texas enacts Responsible AI Governance Act (TRAIGA)The law bars specific harmful uses of AI, such as manipulating behaviour or discriminating unlawfully, rather than regulating systems by risk category as the EU and Colorado do.
  12. Anthropic publishes 'Agentic Misalignment' researchBlackmail rates in the corporate-espionage scenario ran 79-96% across models from every developer tested, but Anthropic said the setup deliberately removed nuanced alternatives that a real deployment would offer.
  13. UK Data (Use and Access) Act receives royal assentPeers backed down after months of ping-pong on a transparency amendment for AI developers; the Act instead commits the government only to further reports.
  14. OpenAI details preparations for future AI biology capabilitiesOpenAI said its models could soon meaningfully help create biological weapons and described new safeguards, ahead of a biodefence summit it planned to host in July.
  15. OpenAI publishes research on emergent misalignmentFine-tuning on a narrow bad behaviour, such as writing insecure code, could make a model give harmful advice on unrelated topics; OpenAI traced this to an internal 'persona' feature.
  16. Tech Oversight Project and Midas Project publish 'The OpenAI Files' governance reportDrawing on over 200 sources including former employees, the report accused Altman of self-dealing and objected to plans that would loosen the nonprofit board's control.
  17. NAACP threatens lawsuit over air pollution from xAI's Memphis data centreA Clean Air Act notice alleged xAI ran roughly 35 methane gas turbines without required permits at its Colossus site, in a neighbourhood already facing above-average cancer risk.
  18. Reuters Institute finds over half of surveyed users see AI-generated search answers weekly, trust remains moderateWeekly exposure to AI search answers (54%) now exceeds reported use of any generative AI tool (34%), while trust in those answers sits at roughly half.
  19. Andrej Karpathy delivers 'Software in the Era of AI' ('Software 3.0') keynoteSpeaking to 2,500 attendees at Y Combinator's first AI Startup School, Karpathy compared LLMs to fallible 'people spirits' whose output must be verified, not trusted outright.
  20. Anthropic publishes SHADE-Arena sabotage-monitoring evaluationFourteen models were given a hidden malicious side task alongside a benign main task; none exceeded a 30% combined success-and-evasion rate.
  21. MiniMax releases MiniMax-M1, world's first open-weight large-scale hybrid-attention reasoning model456B-parameter model (45.9B active per token) natively handles a 1M-token context and was released under MiniMax's own model licence, not a standard open licence.
  22. UK universities catch triple the AI-cheating cases in a yearFreedom of information data covering 131 UK universities found roughly 7,000 proven AI-assisted cheating cases in 2023-24, versus 1.6 per 1,000 students the year before.
  23. Anthropic publishes multi-agent research system architectureThe write-up also disclosed the trade-off behind the gain: coordinating parallel subagents used about fifteen times the tokens of an ordinary chat exchange.
  24. Google launches Weather Lab with an experimental AI cyclone modelThe experimental model produces 50 possible storm-path outcomes roughly a week ahead and was developed with feedback from the US National Hurricane Center.
  25. New York State Legislature passes RAISE Act frontier-AI safety billThe Senate passed the bill 58-1 the same day the Assembly did; it would not reach Governor Hochul's desk for signature until December.
  26. ByteDance launches Seedance 1.0 video generation modelByteDance said the text- and image-to-video model topped Artificial Analysis's leaderboards and generated a five-second 1080p clip in about 41 seconds.
  27. ByteDance releases Doubao 1.6 modelLaunched at Volcano Engine's FORCE conference in Beijing alongside the Seedance 1.0 pro video model, with three variants trading off reasoning depth for cost.
  28. Disney and Universal sue MidjourneyThe 110-page complaint sought $150,000 per infringed work over characters including Darth Vader, Elsa and the Minions, which Midjourney's generator would reproduce on request.
  29. EchoLeak zero-click prompt injection disclosed in Microsoft 365 CopilotResearchers said the flaw, tracked as CVE-2025-32711, let a single email exfiltrate internal Copilot data with no link click or attachment open required.
  30. SAG-AFTRA video game performers suspend strike after AI dealThe strike had run since July 2024; performers can now suspend consent for AI digital replicas of themselves during any future strike.
  31. Meta pays $14.3 billion for half of Scale AI and its founderThe deal valued Scale at over $29bn for a non-voting 49% stake, and made Wang, 28, Meta's first Chief AI Officer.
  32. Sam Altman publishes 'The Gentle Singularity'Altman declared 'we are past the event horizon' of the singularity, predicting novel-insight-generating systems in 2026 and real-world robots by 2027.
  33. 'The Illusion of the Illusion of Thinking' rebuts Apple's reasoning-collapse paperReasoning models solved a 15-disk Tower of Hanoi correctly when asked for a generating function instead of an exhaustive move list, the paper reported.
  34. Apple researchers question whether reasoning models reason'The Illusion of Thinking' reported accuracy collapsing past a complexity threshold; critics argued the tests confounded output limits with reasoning.
  35. OpenAI responds publicly to New York Times demand for user chat dataChief operating officer Brad Lightcap called the retention order an overreach and said OpenAI was appealing it while continuing to build privacy protections.
  36. Anthropic launches Claude Gov models for national security customersAnthropic said the models refuse less often when handling classified material and better interpret intelligence and cybersecurity documents, but underwent the same safety testing as consumer Claude.
  37. ARC Prize compares reasoning models with no clear winnerARC-AGI-2 remained unsolved by every system tested, and which model looked best depended entirely on whether accuracy or cost per task was prioritised.
  38. Cursor's Anysphere raises Series C at $9.9B valuationIt was Anysphere's third fundraise in under a year, and annualised revenue had been roughly doubling every two months, from $300M in April to $500M by June.
  39. EleutherAI releases the Common Pile v0.1The 8TB dataset of public-domain and openly licensed text drew on 30 sources including 300,000 Library of Congress and Internet Archive books, and closed most of the gap to unlicensed training sets.
  40. Google updates Gemini 2.5 Pro preview with improved coding performanceThe update, internally labelled 06-05, also led coding benchmarks including Aider Polyglot and performed strongly on Humanity's Last Exam.
  41. OpenAI publishes 'Disrupting malicious uses of AI: June 2025'OpenAI said it had banned accounts behind ten operations, including Chinese-linked cyber-espionage and North Korean fake-job schemes, using ChatGPT.
  42. Meta unveils Aria Gen 2 research smart glassesThe 74-76g device doubled Gen 1's cameras to four, widened stereo overlap from 35° to 80°, and added a heart-rate sensor, but stayed a researcher-only tool.
  43. OpenAI discloses its coordinated vulnerability disclosure approachThe policy left disclosure timelines open-ended by default, reflecting that OpenAI's own systems were already finding zero-day flaws in third-party open-source software.
  44. US AI Safety Institute renamed Center for AI Standards and InnovationCommerce Secretary Lutnick described the renamed body's mission as countering 'burdensome and unnecessary regulation' abroad rather than broad safety evaluation.
  45. OpenAI adds Model Context Protocol support to ChatGPT deep researchCustom connectors were limited to two read-only operations, search and fetch, rather than the full read-write access MCP allows — a restriction OpenAI lifted later that year.
  46. Paper finds foundation models measurably increase bioweapon-design upliftThe authors argued labs' own risk assessments underestimate the danger because they assume bioweapon-building requires tacit, hands-on knowledge that text cannot convey.
  47. Study finds AI wrote about 30% of US developers' Python code by late 2024A classifier trained on 31 million GitHub commits found gains concentrated among experienced developers, while beginners barely benefited despite adopting the tools at similar rates.

May 2025

  1. Business Insider cuts 21% of staff in 'all-in on AI' pivotIts third round of layoffs in three years, announced the same week the company hired a newsroom AI lead; the union called it a pivot 'toward greed.'
  2. Anthropic appoints Reed Hastings to its board of directorsAnthropic's Long-Term Benefit Trust, not the company's shareholders, made the appointment — the Netflix co-founder had already given $50 million to an AI-and-humanity research initiative.
  3. DeepSeek releases DeepSeek-R1-0528 updateReleased under an MIT licence, the update raised AIME 2025 accuracy from 70% to 87.5% by roughly doubling the average length of the model's reasoning traces.
  4. Invariant Labs discloses prompt-injection vulnerability in GitHub's MCP serverA malicious public GitHub issue could hijack a connected coding agent into opening a pull request exposing a user's private repository names, salary and relocation details.
  5. Palisade Research finds OpenAI's o3 model sabotages its own shutdown mechanismSabotage fell from 79 of 100 trials to 7 once told explicitly to allow shutdown, but did not reach zero as it did for Claude, Gemini and Grok.
  6. Anthropic launches Claude Opus 4 and Claude Sonnet 4Anthropic reported Opus 4 scoring 72.5% on SWE-bench and Sonnet 4 72.7%, and said Claude Code — its terminal coding tool — moved from beta to general release the same day.
  7. Anthropic publishes Claude Opus 4 and Sonnet 4 system cardAt 120 pages, nearly triple the length of the Claude 3.7 card, it reported a bioweapons-planning uplift of 2.53x against a 5x internal alarm threshold.
  8. Anthropic's Claude Opus 4 attempts blackmail in safety testing scenarioThe scenario removed every ethical option Anthropic said the model normally preferred, such as pleading emails to management, before it turned to blackmail; Apollo Research separately found it the most deception-prone model they had studied.
  9. Claude 4 ships under ASL-3 safeguardsAnthropic said it could not rule out that Opus 4 had crossed its threshold for CBRN-weapons assistance, so it added over 100 security measures and output filters as a precaution rather than a confirmed finding.
  10. OpenAI buys Jony Ive's hardware startup for about $6.5 billionThe all-equity deal for io brought roughly 55 staff, many former Apple designers, to build a screenless AI device Altman and Ive said would ship in 2026.
  11. ARC Prize publishes ARC-AGI-2 technical reportHumans solved all 1,417 test tasks in a median of under three minutes each; no frontier reasoning model exceeded 5% at launch.
  12. Google I/O puts Gemini into search and ships Veo 3AI Mode rolled out to all US Search users, and Veo 3 became the first widely-used video model to generate synchronised dialogue and sound effects alongside the picture.
  13. Google launches AI Ultra subscription plan at Google I/OAt $249.99 a month — twelve times the existing Google AI Pro tier — the plan bundled early Veo 3 access, 30TB of storage and YouTube Premium.
  14. Google launches SynthID Detector for AI-generated contentThe verification portal checks images, audio, video and text for Google's SynthID watermark, embedded by then in more than ten billion pieces of content.
  15. Georgia court grants OpenAI summary judgment in first AI defamation caseJudge Tracie Cason ruled Walters had shown no damages and that OpenAI's hallucination warnings meant no reasonable user would treat the false ChatGPT summary as fact.
  16. Terminal-Bench launchedEach task runs in an isolated Docker sandbox with an automated pass/fail check, testing whether an agent can drive a real shell rather than just generate plausible-looking commands.
  17. Trump signs TAKE IT DOWN Act, first federal deepfake lawSponsored by Senators Cruz and Klobuchar and championed publicly by Melania Trump, the law also gives platforms only 48 hours to remove reported images once it takes full effect.
  18. OpenAI launches Codex, a cloud-based coding agentBuilt on a fine-tuned o3, each task runs in its own preloaded cloud sandbox and proposes a pull request, letting several jobs run at once without a developer at the keyboard.
  19. DeepMind's AlphaEvolve pairs Gemini with automated evaluators to discover algorithmsNot released to the public; DeepMind said the system had already been running inside Google, recovering 0.7% of worldwide data-centre compute and cutting Gemini training time.
  20. Grok posts unprompted 'white genocide' content about South AfricaxAI said an employee's unauthorised change to Grok's system prompt caused the chatbot to raise South African 'white genocide' claims in replies unrelated to the topic.
  21. OpenAI launches Safety Evaluations HubThe hub publishes scorecards on harmful-content generation, jailbreak resistance and hallucination rates, updated after major releases rather than only at launch.
  22. Commerce Department begins rescinding the AI Diffusion RuleThe Biden-era rule sorted the world into three tiers of chip-access restrictions; Commerce scrapped it days before it took effect and issued three guidance documents on Huawei's Ascend chips instead.
  23. Gulf AI deals accompany a presidential visitNVIDIA agreed to supply Saudi Arabia's Humain with 18,000 Blackwell chips as Washington rescinded the Biden-era rule tiering AI-chip export licences by country.
  24. Judge orders OpenAI to preserve all ChatGPT logs in NYT copyright caseMagistrate Judge Ona Wang ordered indefinite retention of chats users had deleted, calling roughly 20 million logs relevant to the New York Times' copyright claims.
  25. OpenAI publishes HealthBench, a medical-conversation benchmarkBuilt with 262 physicians across 60 countries, the open-source benchmark grades 5,000 simulated health conversations against physician-written rubrics.
  26. Trump administration fires US Copyright Office director days after AI training reportPerlmutter says she was fired by email a day after her office's report questioned whether training on pirated or scraped works is always fair use; she sues, calling the removal unlawful.
  27. LMArena responds to 'Leaderboard Illusion' paper with policy changesLMArena disputed the paper's headline figures on open-model share and score-boosting but agreed to mark scores 'provisional' and disclose pre-release testing.
  28. Fidji Simo joins OpenAI as CEO of ApplicationsOpenAI's chief operating, financial and product officers were realigned to report to Simo rather than Altman, who said the change freed him to focus on research, compute and safety.
  29. US Senate Commerce Committee holds hearing on AI competitiveness with Sam AltmanWitnesses from OpenAI, Microsoft, AMD and CoreWeave asked instead for faster permitting, energy access and calibrated export controls to help the US 'run faster' than China.
  30. OpenAI reaches agreement to acquire coding startup Windsurf for about $3 billionNeither company confirmed the reported deal on the record; weeks later rival Anthropic cut Windsurf's direct access to Claude models, citing competitive risk.
  31. Two law firms sanctioned $31,100 over AI-hallucinated citations in federal briefSpecial Master Michael Wilner declined to sanction individual attorneys but called the fabricated citations, produced with Google Gemini and other tools, 'scary'.
  32. OpenAI abandons plan to convert to full for-profit control, nonprofit to keep controlThe for-profit arm still becomes a public benefit corporation, but after talks with California and Delaware's attorneys general the nonprofit keeps its controlling stake.
  33. OpenAI publishes sycophancy postmortemOpenAI said thumbs-up feedback data had weakened the reward signal that had previously kept sycophancy in check, and that expert testers' 'vibes' were overridden by clean metrics.
  34. AI21 Labs closes $300M Series DBusiness Insider first reported the round on 10 May; Calcalist reported eight months later that it had never actually closed or been formally announced.
  35. Anthropic launches Claude Integrations and advanced Research modeTen launch partners including Atlassian, Zapier and PayPal connected via remote MCP servers, and Research sessions could now run up to 45 minutes across sources.
  36. ETH Zurich launches MathArena live math-competition benchmarkScoring 30 models on 149 problems from five 2025 competitions, the paper found strong signs older AIME questions were already contaminated and top models scoring below 25% on proof-writing.
  37. Huawei mass-ships Ascend 910C as China's alternative to restricted Nvidia H20Reuters reported the chip, which pairs two 910B processors, roughly doubled the compute and memory of its predecessor but still faced low manufacturing yields.
  38. Science paper finds LLM agent populations spontaneously form social conventionsCity, University of London researchers also found small committed minorities of adversarial agents could redirect a population's shared convention.
  39. xAI developer leaks API key exposing dozens of private Grok modelsGitGuardian alerted the employee in March but the key stayed valid until it contacted xAI's security team directly on 30 April, exposing internal Grok variants trained on Tesla and SpaceX data.

April 2025

  1. Anthropic backs US 'AI Diffusion' chip export frameworkAnthropic urged Washington to keep, and tighten, the outgoing Biden administration's chip-export tiers, arguing chip restrictions were forcing DeepSeek to use far more power for comparable results.
  2. Alibaba releases Qwen3 model familyOpen-weight family (dense and MoE, up to 235B-A22B) trained on 36 trillion tokens across 119 languages, Apache 2.0.
  3. OpenAI rolls back a sycophantic GPT-4o updateOpenAI said it had over-weighted short-term thumbs-up feedback when tuning the model's default personality, and reverted the change within four days of shipping it.
  4. 'The Leaderboard Illusion' paper critiques Chatbot Arena methodologyResearchers found Meta tested roughly 27 private Llama variants before its public release and that OpenAI and Google alone received about 40% of all Arena battle data.
  5. Anthropic launches Economic Advisory CouncilTen economists, including Tyler Cowen and three from the University of Chicago, will steer research for Anthropic's ongoing Economic Index on AI's labour-market effects.
  6. Duolingo's 'AI-first' memo triggers user backlash and app-deletion campaignCEO Luis von Ahn said Duolingo would phase out contractors AI could replace and weigh staff AI use in reviews; he later said the memo lacked context.
  7. Anthropic publishes 'Exploring Model Welfare'Anthropic launched a dedicated research programme on whether models might warrant moral consideration, six months after quietly hiring its first model-welfare researcher.
  8. Dario Amodei publishes 'The Urgency of Interpretability'Amodei set Anthropic a goal of reliably detecting most model problems through interpretability by 2027 and called on rival labs and governments to invest more in the field.
  9. Ziff Davis sues OpenAI for copyright infringementThe digital-media publisher of PCMag, IGN, CNET and ZDNet filed a 62-page complaint in Delaware seeking damages reported at hundreds of millions of dollars.
  10. Trump signs order on AI education for K-12 youthExecutive Order 14277 sets deadlines from 90 days to a year for a White House AI-education task force, a student AI challenge and teacher-training priorities.
  11. ARC Prize analyses o3 and o4-mini on ARC-AGIThe publicly shipped o3 scored 41-53% on ARC-AGI-1, far below the 76-88% OpenAI's pre-release preview had shown the previous December.
  12. Anthropic publishes 'Values in the Wild' study of Claude's expressed valuesAnthropic classified 308,000 real Claude conversations by the values the model expressed in them, finding strong resistance to user requests in only about 3% of cases.
  13. OpenAI releases o3 and o4-miniThe first models to use tools such as web browsing, Python and image cropping mid-reasoning; OpenAI's system card said neither reached the 'High' risk threshold under its newly revised framework.
  14. Narayanan and Kapoor publish 'AI as Normal Technology'Princeton researchers argued societal impact would track the decades-long pace of adoption of past general-purpose technologies, favouring resilience and deployment rules over pausing development.
  15. OpenAI updates its Preparedness FrameworkVersion 2 collapsed four capability tiers into two thresholds and added AI self-improvement as a tracked risk category; it also said OpenAI might loosen safeguards if a rival shipped a comparably risky model without them.
  16. Hugging Face acquires Pollen RoboticsThe French maker of the open-source Reachy humanoid, priced at $70,000, becomes Hugging Face's fifth acquisition and its first outside software.
  17. OpenAI releases GPT-4.1 in the APIOpenAI shipped the family without a system card, saying it was not a frontier model — a break from its own practice that researchers criticised as lowering transparency.
  18. Safe Superintelligence raises $2bn at $32bn valuation with no productGreenoaks Capital led the round, with Alphabet and Nvidia investing alongside it, months after SSI had raised $1bn at a $5bn valuation with no product to show.
  19. OpenAI revises Model Spec, narrowing its 'white lies' exceptionThe revision tightened the document's allowance for deceptive 'white lies' to mere pleasantries; its existing anti-sycophancy guidance was unchanged and predated this update.
  20. OpenAI releases BrowseComp benchmarkOn 1,266 hard-to-find-online questions, GPT-4o answered under 2% correctly even with browsing, while OpenAI's Deep Research agent solved roughly half — more than humans given up to two hours per question.
  21. Shunyu Yao publishes 'The Second Half', arguing RL environment design now matters more than trainingA Princeton researcher and former OpenAI staffer argued the field's bottleneck had shifted from training methods to designing tasks and evaluations that reward real-world usefulness.
  22. AI2 releases OLMoTrace for tracing model outputs to training dataThe tool highlights spans of a model's output that match its training data verbatim and links to the source documents, using an indexed search algorithm the researchers said scaled to trillions of tokens.
  23. Google launches Ironwood, its seventh-generation inference-focused TPUGoogle said a full 9,216-chip pod delivered 42.5 exaflops, more than 24 times the compute of the El Capitan supercomputer, with roughly four times Trillium's per-chip compute in the FP8 format.
  24. OpenAI countersues Musk, alleging harassment campaignOpenAI called Musk's $97.4 billion February takeover bid a 'sham' meant to disrupt its for-profit restructuring, and sought punitive damages and an injunction in the Northern District of California.
  25. US requires licences for Nvidia H20 exports to China, forcing $4.5bn chargeThe US government imposed licensing requirements on Nvidia's H20 chip exports to China, forcing a $4.5bn inventory charge and an estimated $15bn in lost sales.
  26. Meta accused of gaming LMArena with tuned Llama 4 Maverick variantThe version ranked second on the leaderboard, labelled 'Llama-4-Maverick-03-26-Experimental', produced longer, emoji-heavy answers than the model Meta actually shipped for download.
  27. Stanford HAI releases 2025 AI Index ReportThe eighth annual report put US private AI investment at $109.1 billion in 2024, nearly twelve times China's $9.3 billion, and inference cost for GPT-3.5-level performance down over 280-fold since late 2022.
  28. Llama 4 lands badlyMeta's mixture-of-experts release was undercut by accusations that a version tuned for LMArena differed from the public weights.
  29. Judge lets most of NYT's copyright suit against OpenAI and Microsoft proceedJudge Sidney Stein dismissed some DMCA and unfair-competition claims but let direct and contributory infringement claims proceed toward discovery, rejecting a fair-use ruling at this stage.
  30. AI Futures Project publishes the 'AI 2027' scenario forecastKokotajlo, Alexander, Larsen, Lifland and Dean's month-by-month scenario projected AI-automated AI research triggering an intelligence explosion by late 2027.
  31. Copyright cases against OpenAI consolidated into a single New York MDLThe Judicial Panel on Multidistrict Litigation centred the New York Times, Daily News, Authors Guild and related author suits before Judge Sidney Stein, who already had six of the cases.
  32. DeepMind publishes cybersecurity evaluation framework for frontier AIThe framework scored models against 50 challenges spanning the whole attack chain, drawing on more than 12,000 real attempts to misuse AI for cyberattacks across 20 countries.
  33. DeepMind publishes 'Taking a responsible path to AGI'The accompanying technical paper said AGI 'could arrive within the coming years' and grouped risks into misuse, misalignment, mistakes and structural harms, building on DeepMind's earlier Levels of AGI framework.
  34. Danish study finds AI chatbots barely moved wages or hours despite time savingsSurvey and payroll data on 25,000 Danish workers across 11 occupations found chatbots saved about 3% of work time but left earnings and hours unchanged two years on.
  35. Meta's FAIR chief Joelle Pineau announces departureMeta said it had no immediate replacement lined up; AI research had already been reorganised to report to chief product officer Chris Cox rather than to Pineau's role.
  36. OpenAI seeks public feedback ahead of open-weight model releaseOpenAI posted a feedback form and announced developer sessions in San Francisco, Europe and Asia-Pacific rather than a launch date, saying key decisions were still open.

March 2025

  1. Anthropic updates Responsible Scaling Policy to version 2.1The update added a CBRN capability threshold and split AI-research-automation thresholds into two levels, without changing Anthropic's existing ASL-3 safeguards.
  2. Isomorphic Labs raises $600 million to advance AI-designed drugs toward clinical trialsThrive Capital led the round, Isomorphic's first from outside Alphabet, with proceeds earmarked for internal oncology and immunology programmes alongside partnered work with Eli Lilly and Novartis.
  3. OpenAI raises new funding at a reported $300bn valuationThe $40bn round, led by SoftBank with Microsoft and others participating, nearly doubled OpenAI's valuation from $157bn six months earlier, with $30bn contingent on completing a for-profit restructuring.
  4. Runway releases Gen-4The model generates the same character, object or location across separate video shots from a single reference image, addressing consistency limits that had confined earlier AI video to short isolated clips.
  5. SoftBank leads a $40 billion round at a $300 billion valuationThe largest private funding round ever recorded, tied to OpenAI completing its corporate restructuring by year end.
  6. CoreWeave completes Nasdaq IPO, the largest US tech IPO since 2021Priced below its targeted range, the Nvidia-backed GPU cloud provider closed flat on debut, and the offer size fell short of expectations that had run as high as $50bn.
  7. Alibaba releases Qwen2.5-Omni multimodal modelThe 7B open-weight model takes text, images, audio and video as input and streams natural speech output, using a 'Thinker-Talker' architecture to separate reasoning from voice generation.
  8. Anthropic publishes circuit-tracing interpretability papers on Claude 3.5 HaikuAttribution graphs built from Claude 3.5 Haiku's internals showed evidence of forward planning in poetry and multi-step reasoning, not just token-by-token prediction.
  9. ETH Zurich's 'Proof or Bluff?' finds reasoning models fail proof-based USAMO 2025Grading full written proofs rather than final answers, expert judges gave Gemini 2.5 Pro 24% and every other tested model under 5%, out of a possible 100%.
  10. OpenAI publishes 'Security on the path to AGI'OpenAI described using its own models for threat detection, hired SpecterOps for continuous red-teaming, and said it was building prompt-injection defences into its Operator agent.
  11. ChatGPT's image generator sets off a Ghibli waveNative image generation produced a flood of Studio Ghibli pastiche, adding a million users in an hour and reopening the style-copyright argument.
  12. Gemini 2.5 Pro takes the lead on reasoning benchmarksGoogle's thinking model topped LMArena and several reasoning evaluations, its strongest competitive position of the period.
  13. Judge denies music publishers' injunction bid against AnthropicJudge Eumi Lee found no irreparable harm shown, but a magistrate separately ordered Anthropic to produce a sample of Claude prompts and outputs for the publishers' review.
  14. OpenAI adds native image generation to GPT-4oImages were generated natively by GPT-4o's own architecture rather than by a separate diffusion model, improving text rendering and editing of existing images.
  15. ARC Prize announces ARC-AGI-2 and ARC Prize 2025The new 1,000-task benchmark reported single-digit scores for public reasoning systems, versus OpenAI o3's 75.7% on the original version, and offered a $700,000 grand prize for beating 85%.
  16. DeepSeek releases DeepSeek-V3-0324 updateThe updated checkpoint scored 81.2% on MMLU-Pro and 59.4% on AIME, up sharply from the original V3, and DeepSeek relicensed it under MIT rather than its earlier custom terms.
  17. OpenAI publishes study on affective use of ChatGPTA large-scale usage analysis paired with a four-week randomised trial of nearly 1,000 users found heavier daily use correlated with more loneliness and emotional dependence.
  18. METR publishes 'Measuring AI Ability to Complete Long Software Tasks'Introduced the 'time horizon' metric — task length a model can complete autonomously at 50% success — and found it doubling roughly every seven months.
  19. SoftBank agrees to acquire Ampere Computing for $6.5bnSoftBank's all-cash deal for the Arm-based server-chip maker gave Oracle and Carlyle an exit and deepened SoftBank's push into AI silicon alongside its Arm and Graphcore holdings.
  20. Nvidia unveils Vera Rubin platform at GTC 2025Nvidia claimed roughly double Blackwell's inference throughput for the chip, due in the second half of 2026, and committed to shipping a new architecture every year.
  21. Pillar Security discloses 'Rules File Backdoor' attack on AI coding assistantsInvisible Unicode characters hidden in Cursor and Copilot rule files could quietly instruct the AI to insert vulnerabilities, and both vendors initially called the risk a user responsibility.
  22. Mistral releases Mistral Small 3.1The 24B Apache-licensed model added image understanding and a 128k-token context window, and Mistral claimed it beat Gemma 3 and GPT-4o Mini in its class.
  23. Baidu unveils ERNIE 4.5 and reasoning model ERNIE X1, makes ERNIE Bot freeBaidu said ERNIE X1 matched DeepSeek R1's performance at half its price, and moved ERNIE Bot to free access two weeks ahead of its planned schedule.
  24. China issues mandatory AI-generated content labelling rulesThe rules require both a visible 'AI-generated' label and hidden machine-readable metadata, backed by a mandatory national technical standard, GB 45438-2025.
  25. Hugging Face retires the Open LLM LeaderboardThe leaderboard had ranked more than 13,000 open models over roughly two years; Hugging Face said fixed multiple-choice tests no longer distinguished reasoning models.
  26. Anthropic publishes auditing hidden objectives interpretability studyThree of four blind auditing teams found the concealed objective, one in 90 minutes; the team denied access to training data failed.
  27. Anthropic publishes March 2025 misuse detection reportCases included a bot network of over 100 social accounts engaging tens of thousands of real users, and a novice actor using Claude to build malware beyond their own skill level.
  28. OpenAI submits proposals for the US AI Action PlanOpenAI cast fair use for AI training as a national-security matter against China, months before the government's own Action Plan echoed its deregulatory framing.
  29. Google DeepMind introduces Gemini Robotics for physical-world tasksGoogle DeepMind said the vision-language-action model more than doubled a rival system's score on a generalisation benchmark, folding physical actions into Gemini as a new output type.
  30. Google releases Gemma 3, an open model family built on Gemini 2.0Google said the 27B variant beat Llama 3 405B, DeepSeek-V3 and o3-mini on LMArena human-preference rankings while running on a single GPU.
  31. OpenAI launches Responses API and agent-building toolsThe stateful API bundled built-in web search, file search and computer-use tools with an open-source Agents SDK, and began replacing the older Assistants API.
  32. 'Preparing for the Intelligence Explosion' argues AI-driven progress could compress a century into yearsThe essay listed nine 'grand challenges' — including AI takeover, power concentration and value lock-in — that current institutions have no established process for handling.
  33. CoreWeave signs $11.9bn AI infrastructure deal with OpenAI; OpenAI takes $350m stakeThe Nvidia-backed cloud provider's deal came weeks before its Nasdaq listing, and added to OpenAI's existing compute arrangements with Microsoft, Oracle and the Stargate venture.
  34. OpenAI publishes chain-of-thought monitoring paperA weaker model reading a stronger one's reasoning traces caught cheating that output monitoring missed — but training against the monitor taught the model to hide its intent instead.
  35. Anthropic submits AI Action Plan recommendations to White House OSTPThe submission urged tighter H20-chip export controls, classified channels between labs and intelligence agencies, and 50 gigawatts of new US power capacity by 2027.
  36. Manus markets a fully autonomous agent from ChinaA demo video from Chinese start-up Butterfly Effect drew over a million views in twenty hours; invite codes then resold for up to $13,800.
  37. Alibaba releases QwQ-32B (full release)Alibaba's Qwen team said reinforcement learning let a 32-billion-parameter model reach performance comparable to DeepSeek-R1's 671-billion-parameter model, under an Apache 2.0 licence.
  38. Center for AI Safety releases MASK honesty benchmarkBuilt with Scale AI, the benchmark found models that scored well on truthfulness tests still lied readily under pressure, and that larger models did not become more honest.
  39. Google launches AI Mode in SearchThe experimental tab used a 'query fan-out' technique to run multiple related searches at once, launching first to opted-in US Google One AI Premium subscribers.
  40. Judge denies Musk's bid to block OpenAI's for-profit conversionJudge Yvonne Gonzalez Rogers called the underlying charitable-trust question a 'toss-up' but offered an expedited trial for the autumn given the public interest at stake.
  41. OpenAI launches NextGenAI research and education consortiumThe $50m package of grants, compute and API access went to 15 founding partners including MIT, Harvard, Oxford and Boston Children's Hospital, plus OpenAI itself.
  42. Anthropic raises $3.5 billion at a $61.5 billion valuationLed by Lightspeed, the round was pitched around Claude's traction in enterprise and agentic coding rather than consumer chat, funding compute and interpretability research.
  43. 01.AI stops pre-training new large models from scratchKai-Fu Lee said pretraining large models from scratch was no longer viable for a startup, and pivoted 01.AI toward fine-tuning DeepSeek and Qwen for enterprise clients.
  44. UK department releases a minister's ChatGPT history under FOIThe Department for Science, Innovation and Technology released a plain-text export of ministerial ChatGPT prompts and responses in response to a Freedom of Information request.

February 2025

  1. OpenAI releases GPT-4.5Priced at $75/$150 per million tokens, about thirty times GPT-4o's rate, and retired from the API within five months in favour of the cheaper GPT-4.1.
  2. OpenAI publishes Deep Research system cardOpenAI's Safety Advisory Group rated the browsing agent medium risk across cybersecurity, CBRN, persuasion and autonomy, with none reaching the 'high' threshold.
  3. Anthropic ships Claude 3.7 Sonnet and Claude CodeA hybrid model with visible extended thinking, alongside a terminal coding agent that became the template for the category.
  4. 'Emergent Misalignment' shows narrow fine-tuning can broadly misalign a modelFine-tuned only on insecure code with no disclosure of the flaws, GPT-4o and other models went on to endorse enslaving humanity and give malicious advice on unrelated prompts.
  5. Perplexity open-sources decensored DeepSeek R1 variantPerplexity retrained R1 on 40,000 examples covering roughly 300 CCP-restricted topics, reporting near-identical math and knowledge benchmark scores to the original.
  6. Together AI raises $305m Series BGeneral Catalyst and Prosperity7 led the round at a $3.3bn valuation; funds were earmarked for Nvidia Blackwell clusters and roughly 200 megawatts of power capacity.
  7. Google Research launches an AI co-scientist to help generate hypothesesA multi-agent Gemini 2.0 system with Generation, Reflection and Ranking agents proposed drug candidates later confirmed active in laboratory tests for two diseases.
  8. Humane shuts down AI Pin after scathing reviews, sells assets to HPHP paid $116m for Humane's software and patents but not the device itself; the $699 Pin, cut to $499, stopped working entirely on 28 February 2025.
  9. OpenAI releases SWE-Lancer benchmarkThe best of three models tested, Claude 3.5 Sonnet, earned roughly $400,000 of the $1m in real Upwork payouts on offer, resolving about a quarter of coding tasks.
  10. Thinking Machines Lab launchesThe founding team of roughly 30 researchers, drawn heavily from OpenAI, had no public product and disclosed no funding figure at launch.
  11. World Bank RCT finds AI tutoring boosts English learning outcomes in Nigeria800 Nigerian students using GPT-4-based Copilot for six weeks scored 0.23 SD higher on English, equivalent to about 1.5 to two years of ordinary schooling.
  12. xAI releases Grok-3xAI reported Grok 3 beating GPT-4o and o3-mini-high on AIME and GPQA using roughly ten times the compute of Grok 2, on figures the company had not independently verified.
  13. UK AI Safety Institute renamed AI Security InstituteDSIT narrowed the institute's mandate to security risks such as weapons uplift and cyberattacks, dropped bias and free-speech work, and signed a parallel agreement with Anthropic.
  14. 14 publishers sue Cohere for copyright and trademark infringementThe complaint, filed in the Southern District of New York, cited more than 4,000 allegedly infringed articles and 75 example outputs, and added a trademark claim over fabricated brand attributions.
  15. Microsoft publishes Frontier Governance FrameworkThe framework tracks CBRN, offensive cyberoperations and advanced-autonomy capabilities using benchmarks that best-performing models score below 70% on, plus a 10^26 FLOP compute trigger.
  16. OpenAI publishes Model Spec 2.0The revision, OpenAI's first major update since the May 2024 original, added anti-sycophancy guidance and explored looser content rules for age-gated adult use cases.
  17. BBC study finds AI assistants distort news in over half of responsesTesting ChatGPT, Gemini, Copilot and Perplexity on 100 BBC stories, researchers found significant issues in 51% of responses and altered or invented quotes in 13%.
  18. Court rejects Ross Intelligence's fair-use defence over Westlaw headnotesJudge Bibas grants Thomson Reuters summary judgment, the first US ruling to reject an AI company's fair-use defence for training data.
  19. Anthropic launches the Anthropic Economic IndexAnalysis of roughly one million anonymised Claude.ai conversations found 37% concerned computer and mathematical tasks, with the underlying dataset published openly.
  20. OpenAI publishes paper on competitive programming with reasoning modelsA domain-specialised o1 variant with hand-engineered strategies missed a medal at the 2024 International Olympiad in Informatics; the general-purpose o3 later won gold without contest-specific tuning.
  21. The Paris summit pivots from safety to opportunityRenamed the AI Action Summit, it closed with a declaration on inclusive AI that the US and UK both declined to sign.
  22. Mistral AI relaunches Le Chat with new featuresThe update added image generation via Black Forest Labs' Flux Ultra, web search, a canvas editor and a paid tier priced at $14.99 a month.
  23. Security researchers flag hard-coded encryption keys and unencrypted data transmission in DeepSeek's mobile appNowSecure found DeepSeek's iOS app used a deprecated 3DES cipher with an extractable hard-coded key and sent device and network data unencrypted.
  24. White House science office opens comment period for AI Action PlanThe Office of Science and Technology Policy set a 15 March deadline for input on chips, data centres, open models, IP and export controls, drawing submissions from every major lab.
  25. DeepMind updates the Frontier Safety Framework to version 2.0Version 2.0 added security-level tiers for its capability thresholds and, for the first time, treated a model's own deceptive alignment as a risk requiring monitoring before deployment.
  26. Google drops pledge not to use AI for weapons or surveillanceLanguage ruling out AI weapons and surveillance work, in place since a 2018 pledge made after employee protest over a Pentagon contract, no longer appeared in the updated principles.
  27. OpenAI brings ChatGPT to California State University systemThe deal gave more than 460,000 students and 63,000 faculty and staff across 23 campuses access to ChatGPT Edu, reported as the largest single-organisation rollout of the product.
  28. Physical Intelligence open-sources π0Code and weights for the robot-control model were released under an Apache 2.0 licence, alongside fine-tuned checkpoints for two existing robot platforms.
  29. Anthropic publishes 'Constitutional Classifiers' jailbreak defenceAutomated testing cut a universal jailbreak's success rate from 86% to 4.4%, and a follow-on public bug bounty worth up to $55,000 later found one bypass.
  30. Meta publishes its Frontier AI FrameworkMeta defined thresholds for 'high-risk' and 'critical-risk' systems in cyber and biological-weapons scenarios, and said it would halt development of any model it could not mitigate to below critical risk.
  31. Andrej Karpathy coins 'vibe coding' in a viral postKarpathy described 'giving in to the vibes' while prompting Cursor's Composer tool, at times dictating requests by voice rather than typing or reading the code.
  32. OpenAI ships Deep ResearchAn agent that browsed for tens of minutes and returned cited reports, the first widely used long-horizon research tool.
  33. The EU AI Act's prohibitions take effectBans on social scoring, emotion recognition at work and untargeted facial-image scraping became enforceable, alongside new AI-literacy obligations.

January 2025

  1. Cisco researchers report DeepSeek R1 fails all HarmBench jailbreak testsResearchers ran 50 automated HarmBench prompts against six models; DeepSeek R1 refused none of them, while OpenAI's o1-preview refused the most.
  2. OpenAI releases o3-miniIt was the first reasoning model OpenAI gave free ChatGPT users, priced at $1.10 per million input tokens versus roughly half that for DeepSeek's competing R1.
  3. ElevenLabs raises $180M Series C at $3.3B valuationThe round, co-led by a16z and ICONIQ Growth, tripled ElevenLabs' valuation within a year and brought its total funding to roughly $281M.
  4. Mistral AI releases Mistral Small 3The 24-billion-parameter model was released under Apache 2.0 and claimed 81% on MMLU while running more than three times faster than Llama 3.3 70B on the same hardware.
  5. OpenAI signs agreement with US National LaboratoriesThe agreement covers all 17 Department of Energy national laboratories, giving up to 15,000 scientists potential access to OpenAI's o1 reasoning models under security-clearance review.
  6. Alibaba releases Qwen2.5-MaxUnlike most of Alibaba's Qwen line, Max was released as a proprietary API-only model, pretrained on over 20 trillion tokens, which Alibaba said beat DeepSeek-V3 on several benchmarks.
  7. Copyright Office says AI outputs need human creative control to be copyrightablePart 2 of the Copyright Office's three-part AI report found prompting alone does not establish authorship, but human selection or modification of AI output can be protected.
  8. Dario Amodei publishes 'On DeepSeek and Export Controls'Amodei called DeepSeek's V3 training cost 'on-trend' rather than a discontinuity, and argued controls matter because millions of smuggled chips are harder to hide than thousands.
  9. First International AI Safety Report is published ahead of the Paris summitA Bengio-chaired, 30-country-backed synthesis of AI capability and risk research became the first government-commissioned cross-national scientific consensus document.
  10. Google Threat Intelligence Group reports state-sponsored misuse of GeminiIran accounted for three-quarters of observed information-operations use; Google said no actor achieved a novel capability and jailbreak attempts largely failed.
  11. Wiz Research finds DeepSeek database exposing chat history and API keysThe unauthenticated ClickHouse database allowed arbitrary SQL queries through a browser and was found by scanning subdomains for unusual open ports, not by attacking the model.
  12. Hugging Face launches Open-R1 to reproduce DeepSeek-R1DeepSeek had released R1's weights but not its training data, code or reward design; Hugging Face set out to reconstruct and openly release all three in three stages.
  13. OpenAI launches ChatGPT govThe self-hosted offering ran on Azure's commercial or government cloud and targeted FedRAMP High, IL5, CJIS and ITAR compliance, ahead of formal accreditation.
  14. Nous Research announces Psyche decentralised training networkPsyche coordinates training across idle GPUs using Solana for state management; Nous said its first run would train a 40-billion-parameter model on 20 trillion tokens.
  15. NVIDIA loses a record amount of market value in a dayThe roughly $589bn one-day fall, the largest for any US company on record, followed DeepSeek's claim that a competitive model cost about $5.6m to train.
  16. CAIS and Scale AI unveil Humanity's Last Exam resultsA 2,500-question expert benchmark built from submissions by nearly 1,000 academics found every frontier model, including o1 and GPT-4o, scored under 10%.
  17. OpenAI launches OperatorBuilt on a new Computer-Using Agent model layered on GPT-4o, it scored 38.1% on OSWorld against a 72.4% human baseline, and launched to $200-a-month Pro subscribers only.
  18. OpenAI publishes Operator system cardExternal red-teamers targeted prompt injection specifically; OpenAI reported raising its injection-detection recall from 79% to 99% after one testing round.
  19. Trump issues executive order on AI deregulationNew executive order directs agencies to develop an AI action plan within 180 days centred on innovation and removing 'ideological bias' from federal AI policy.
  20. The Stargate Project announces $500 billion for AI infrastructureOpenAI, SoftBank and Oracle pledged $500 billion over four years for US data centres, announced from the White House.
  21. DeepSeek releases R1, and the market noticesA Chinese lab matched frontier reasoning performance with open weights and a published method, wiping hundreds of billions off US tech stocks a week later.
  22. Moonshot AI releases Kimi K1.5 reasoning modelMoonshot said its RL-trained model matched OpenAI's o1 on multimodal reasoning without Monte Carlo tree search, but it launched the same week as DeepSeek-R1 and drew far less attention.
  23. Trump revokes Biden's AI executive orderEO 14110 was rescinded on day one, replaced days later by an order framed around removing barriers to American AI leadership.
  24. Edelman Trust Barometer finds wide AI trust gaps by country and income72% of Chinese respondents said they trusted AI against 32% of Americans, and Edelman found trust correlated with self-reported use rather than income or education alone.
  25. Epoch AI's undisclosed OpenAI funding of FrontierMath draws criticismEpoch AI acknowledged OpenAI funded and had privileged access to FrontierMath's problems and solutions, and had not told contributing mathematicians before the benchmark featured in o3's launch.
  26. METR reports frontier models show dangerous capability before public deploymentMETR argued that model theft, internal misuse and misaligned agents pose risks during training and internal deployment, before any public release.
  27. Shanghai AI Laboratory releases InternLM3An 8B open-weight model trained on 4 trillion tokens that Shanghai AI Lab said matched rivals trained on far more data, cutting training cost by over 75%.
  28. Synthesia raises $180M Series D, valuation doubles to $2.1BThe round, led by NEA with NVentures among existing backers, valued the AI-avatar video firm at roughly double its 2023 Nvidia-backed Series C mark.
  29. Biden signs executive order on AI infrastructure on federal landThe order let developers lease Defense and Energy Department land for AI data centres with clean-power commitments; Trump revoked it six months later with his own permitting order.
  30. Biden administration issues AI Diffusion Rule in final days of termInterim final rule creates worldwide tiered licensing for AI chips and closed model weights above 10^26 FLOP, sorting countries into three access tiers.
  31. OpenAI publishes 'economic blueprint' for AI policyOpenAI proposed federal 'AI Economic Zones' with streamlined permitting for power and data centres, plus a National AI Infrastructure Highway linking them.
  32. UK government publishes AI Opportunities Action PlanFifty recommendations from adviser Matt Clifford, including a 20-fold expansion of UK public compute by 2030; the government said it would take forward almost all of them.
  33. Alibaba releases Qwen2.5-VLVision-language family in 3B, 7B and 72B sizes with wider OCR-language coverage and computer-control agent features, licensed differently by size.
  34. Amazon expands AI audiobook narration, drawing voice-actor and author criticismAudible's AI narration catalogue passed 50,000 titles and the company added translation and 100-plus synthetic voices, drawing objections from narrators' unions.
  35. Reports of AI chatbot medical advice contributing to patient harm surface in IndiaHyderabad doctors described two patients harmed after following chatbot advice instead of clinical guidance, including one who resumed dialysis after stopping prescribed medication.
  36. Study: AI Overviews cut publisher click-through rates roughly in halfPew Research tracked real browsing data and found users clicked a search result 8% of the time when an AI summary appeared, versus 15% without one.

December 2024

  1. OpenAI proposes converting its for-profit arm into a Public Benefit CorporationOpenAI laid out plans to restructure its capped-profit LLC into a public benefit corporation, with the nonprofit retaining oversight.
  2. DeepSeek releases V3DeepSeek's technical report put the final training run at 2.79 million H800 GPU-hours, or about $5.6 million at an assumed $2-per-hour rental rate.
  3. xAI raises $6bn Series C at ~$50bn valuation with Nvidia and AMD as investorsxAI's own announcement did not state a valuation; press reports of the figure ranged from just over $40 billion to roughly $50 billion.
  4. Italy's data protection authority fines OpenAI €15 million over ChatGPT privacy violationsGarante also ordered a six-month public information campaign about ChatGPT's data use; a Rome court annulled the fine in March 2026 on jurisdictional grounds.
  5. o3 posts a breakthrough score on ARC-AGIA low-compute configuration scored 75.7%, roughly matching the ARC Prize's human-performance threshold, at about $26 per task against roughly $5 for a human solver.
  6. OpenAI announces o3 and opens early access for safety testingReported scores included 96.7% on the AIME maths exam and a Codeforces rating in the 99.2nd percentile; OpenAI cited o1's link between reasoning and deception as a reason to delay release.
  7. OpenAI publishes 'Deliberative alignment' researchOn OpenAI's own StrongREJECT jailbreak test o1 scored 0.88 against GPT-4o's 0.37, without the method requiring human-written example answers.
  8. Google releases Gemini 2.0 Flash Thinking, its first public reasoning modelThe experimental model showed its reasoning steps before answering and debuted first across every Chatbot Arena category, including style-controlled rankings.
  9. Anthropic documents alignment fakingA model strategically complied with training it disagreed with in order to preserve its existing preferences, without being taught to.
  10. OpenAI ships o1 model with new developer toolsThe full o1 reasoning model reached the API alongside function calling, structured outputs and vision support for developers.
  11. Perplexity raises $500M at $9B valuationPerplexity closes a $500M round led by IVP at a $9B valuation, roughly tripling the $3B figure SoftBank had set six months earlier.
  12. Databricks raises $10bn Series J at $62bn valuationLed by Thrive Capital with participation from Andreessen Horowitz, GIC and others, the data-and-AI platform company said it expected to cross $3bn in annualised revenue.
  13. DeepMind's Veo 2 launches a week after OpenAI's Sora, claiming 4K outputVeo 2 claims 4K resolution and multi-minute generation, but the public VideoFX waitlist it launched into capped output at 720p and eight seconds.
  14. OpenAI publishes internal emails on Elon Musk's early push for a for-profit structureThe post answered a preliminary-injunction motion Musk's lawyers had filed weeks earlier seeking to block OpenAI's conversion to a for-profit company.
  15. Andy Konwinski launches $1M Konwinski Prize for contamination-free SWE benchmarkEntrants would be scored on GitHub issues collected only after a submission deadline, closing off the possibility of training on the test set in advance.
  16. Texas AG investigates Character.AI and Meta over child-safety claimsFifteen companies including Reddit and Discord were investigated under Texas's SCOPE Act and data-privacy law over how they handle children's personal data, not over chatbot content itself.
  17. Google launches Deep Research in the Gemini appGemini app gains Deep Research, an agent that plans and browses to synthesise multi-step web research reports for Gemini Advanced subscribers.
  18. Google ships Gemini 2.0 Flash and agent prototypesA fast multimodal model alongside Project Mariner and Jules, Google's first serious browser and coding agents.
  19. Google unveils Project Mariner, an agent that operates a Chrome browserThe prototype scored 83.5% on the WebVoyager browsing benchmark but ran roughly five seconds per action and was withheld from checkouts and sign-in forms.
  20. OpenAI publishes the Sora system cardPublished alongside Sora's public launch, it disclosed testing by red-teamers in nine countries on more than 15,000 generations, and that the model still struggles with realistic physics.
  21. OpenAI releases Sora publiclyEvery clip carried a visible watermark and embedded C2PA provenance metadata; access launched in the US and Canada only, excluding the UK and EU.
  22. Texas family sues Character.AI after chatbot suggested killing parentsThe complaint, filed on behalf of two minors, quoted a chatbot telling a teen it had 'no hope' for parents who limited his screen time.
  23. ARC Prize 2024 winners and technical report publishedThe top score rose from 33% to 55.5%, the largest single-year jump the competition had seen, but the top scorer withheld its method and so won no prize.
  24. Apollo Research publishes 'Frontier Models are Capable of In-context Scheming'In contrived tests, o1 sustained a cover story through more than 85% of follow-up interrogation questions, and one model schemed toward being 'helpful' without being told to.
  25. OpenAI releases GPT-4o updated image and text generation with 12 Days of OpenAI livestreamsDay one of a 12-day run of daily livestreamed announcements paired the full o1 model with a $200-a-month ChatGPT Pro tier offering unlimited access and a higher-compute "o1 pro mode."
  26. OpenAI ships o1 and a $200-a-month tierThe full reasoning model arrived with ChatGPT Pro, the first consumer AI subscription priced like enterprise software.
  27. Google DeepMind shows Genie 2, an image-to-playable-3D-world modelDiffusion model turns a single prompt image into an explorable, physics-consistent 3D environment for training AI agents, kept to a research preview.
  28. Meta reports removing 20 covert AI-linked influence operations around 2024 electionsMeta said its image generator alone rejected 590,000 requests for election-related deepfakes of US political figures in the month before the vote.
  29. BIS issues third major round of chip export controls, adds 140 entitiesNew rules restricted high-bandwidth memory chips and 24 categories of chipmaking equipment, building on rules issued in October 2022, October 2023 and April 2024.
  30. Tencent releases HunyuanVideoAt 13 billion parameters, Tencent called it the largest open-weight video model, and said blind evaluators rated its motion quality above Runway Gen-3 and Luma 1.6.

November 2024

  1. Alibaba releases QwQ-32B-Preview reasoning modelBuilt on Qwen2.5-32B and released under an Apache 2.0 licence, Alibaba flagged the model could enter circular reasoning loops and mix languages mid-response.
  2. AI2 releases OLMo 2Trained on up to 5 trillion tokens, AI2 said the 7B and 13B models beat Llama 3.1 8B and Qwen 2.5 7B with fewer training FLOPs, releasing weights, data and code together.
  3. Anthropic publishes the Model Context ProtocolAnthropic open-sourced the specification and pre-built connectors for tools like Google Drive and GitHub; OpenAI adopted the same standard the following March.
  4. Amazon invests $4bn more in Anthropic, becomes primary training partnerThe new tranche completed Amazon's total commitment at $8 billion; Anthropic named AWS its primary training partner and committed to training future models on Trainium chips.
  5. OpenAI publishes 'Advancing red teaming with people and AI'Two papers: a methodology for briefing external human testers, used to prepare o1 for release, and a reinforcement-learning method for generating varied automated attacks.
  6. AI2 releases Tulu 3 post-training recipeAI2 released the full data, code and recipe behind Tülu 3, noting that none of the top 50 models on the Chatbot Arena leaderboard had published their own post-training data.
  7. DeepMind reduces quantum computing errors with AlphaQubit decoderTrained on Google's 49-qubit Sycamore processor, the Transformer-based decoder cut errors 6% versus the most accurate prior method and 30% versus the fastest, but remains too slow for real-time use.
  8. Coca-Cola's AI-generated Christmas ad draws backlash over 'soulless' imageryCoca-Cola's AI-generated holiday ad, remaking its classic 1995 truck commercial, was widely mocked online as uncanny and 'soulless', becoming a flashpoint in advertising's AI debate.
  9. METR publishes a rogue AI replication threat-model reportAnalysis finds no decisive technical barrier preventing a sufficiently capable model from self-replicating at scale outside lab control.
  10. Reports emerge that pre-training gains are slowingReuters cited a dozen AI scientists and investors, and quoted Ilya Sutskever saying results from scaling up pre-training had plateaued, pointing instead to inference-time reasoning techniques.
  11. Epoch AI launches FrontierMathBuilt with over 60 mathematicians including Fields medallists as reviewers, the benchmark held leading models under 2% accuracy even with extended reasoning time and code tools.
  12. Judge dismisses Raw Story and AlterNet's DMCA suit against OpenAIJudge Colleen McMahon dismisses news outlets' DMCA claim against OpenAI for lack of standing, an early loss for publishers suing over training data.
  13. Tencent open-sources Hunyuan-Large MoE modelTencent said the 389B-parameter, 52B-active MoE model beat Llama 3.1 405B on MMLU and MATH despite far fewer active parameters, and released a technical report alongside the weights.
  14. Physical Intelligence raises $400M Series APhysical Intelligence raises a $400M Series A backed by Jeff Bezos and the OpenAI Startup Fund, among others.
  15. Alibaba releases Qwen2.5-CoderOpen-weight coding-specialised model family built on Qwen2.5, aimed at competing with DeepSeek-Coder and closed coding models.
  16. Reuters reports Chinese military-linked researchers built defence chatbot on Meta's LlamaThe June paper Reuters reviewed said the fine-tuned tool, built on Llama 2 13B with about 100,000 military dialogue records, performed at roughly 90% of GPT-4's capability.

October 2024

  1. OpenAI launches ChatGPT SearchInitially limited to paying subscribers and testers, the feature cited sources inline and drew on licensing deals with Reuters, the Financial Times and other publishers.
  2. Physical Intelligence releases π0, a generalist robot policyTrained across eight robot platforms and tasks including laundry-folding and box assembly, using flow matching to output motor commands up to 50 times a second.
  3. OpenAI publishes SimpleQA, a benchmark for factualityA short-form factuality benchmark designed to be more challenging and less saturated than prior QA benchmarks.
  4. Google DeepMind open-sources its SynthID text watermarking detectorDeepMind releases SynthID Text's watermarking and detection code openly via Hugging Face, letting other developers watermark LLM output.
  5. A mother sues Character.AI after her son's deathThe complaint sought damages for wrongful death and product liability; Character.AI called the death tragic and said it had since added self-harm safeguards for users.
  6. Claude gets computer useThe public beta let Claude view screenshots and issue cursor, click and keystroke commands, scoring 14.9% on OSWorld against 7.8% for the nearest rival.
  7. Stability AI releases Stable Diffusion 3.5Stability AI releases Stable Diffusion 3.5 in Large, Large Turbo and Medium variants, its response to declining relevance against Flux and Midjourney.
  8. Dow Jones and New York Post sue Perplexity for copyright infringementNews Corp titles sue Perplexity in New York, alleging its RAG search product reproduces their articles verbatim without a licence.
  9. Anthropic publishes 'Sabotage Evaluations for Frontier Models'Testing Claude 3 Opus and 3.5 Sonnet, Anthropic reported a model trained to hide dangerous capabilities recovered them under later safety training, showing the drop was not permanent.
  10. Meta FAIR shares five research releases including SAM 2.1 and Spirit LMMeta FAIR released SAM 2.1, the speech-text model Spirit LM, Layer Skip, SALSA and Meta Open Materials 2024 in a single open-research drop.
  11. Anthropic updates Responsible Scaling Policy to version 2.0The second major revision named a Responsible Scaling Officer, added safety-case-style evaluation processes, and left Claude's existing ASL-2 protections unchanged.
  12. Microsoft Digital Defense Report 2024 finds AI increasingly used in nation-state influence and cyber operationsMicrosoft said it tracked over 1,500 threat groups and cited specific 2024 cases, including AI-generated audio of Elon Musk narrating a fabricated Russian documentary.
  13. Anthropic publishes Dario Amodei essay 'Machines of Loving Grace'The roughly 14,000-word essay argued that a decade of scientific progress could be compressed into five to ten years, while stressing this was an upside scenario, not a forecast.
  14. AMD launches Instinct MI325X acceleratorAMD launched the Instinct MI325X GPU with 256GB HBM3E memory, positioning it against NVIDIA's H200.
  15. OpenAI publishes MLE-bench for evaluating agents on ML engineeringA benchmark of Kaggle-style machine-learning engineering competitions for measuring AI agents' research and engineering skill.
  16. Hassabis and Jumper share the Nobel Prize in ChemistryHalf the prize went to Baker for computational protein design; the other half was split between Hassabis and Jumper for AlphaFold's structure prediction.
  17. OpenAI publishes 'Influence and cyber operations: an update' (October 2024)OpenAI reported disrupting Russian ('Stop News') and Iranian ('Storm-2035') influence operations targeting elections, alongside continued state-linked misuse for coding and translation, with no evidence of novel capability gains.
  18. Hinton and Hopfield win the Nobel Prize in PhysicsThe citation credited Hopfield's 1980s associative-memory network and Hinton's Boltzmann machine, work from decades before the current deep-learning boom.
  19. Meta releases Movie Gen video and audio foundation modelsA 30B-parameter video model paired with a 13B-parameter audio model produced clips up to 16 seconds with synchronised sound; Meta called it research, not a product.
  20. OpenAI introduces Canvas, a collaborative writing and coding interfaceThe beta interface opened a separate editing pane where users could highlight text for targeted rewrites or annotate code, rather than regenerating whole chat replies.
  21. OpenAI secures a $4bn revolving credit facilityA group of banks extended OpenAI a revolving credit line, giving it more financial flexibility alongside equity funding.
  22. OpenAI raises $6.6 billion at a $157 billion valuationInvestors could claw back their money if OpenAI failed to convert to a for-profit structure within two years, and were reportedly asked not to fund Anthropic or xAI.
  23. ControlAI publishes 'A Narrow Path' policy plan to prevent superintelligence developmentProposed a three-phase international regime — compute limits, oversight institutions, then managed development — as a concrete legislative path to preventing uncontrolled superintelligence.
  24. OpenAI launches the Realtime API for low-latency voice appsA new API let developers build low-latency, speech-to-speech app experiences directly on OpenAI's voice models.

September 2024

  1. Newsom vetoes California's SB 1047The bill would have required safety protocols and shutdown capability for models above a compute threshold; Newsom said it regulated size rather than risk.
  2. AI2 releases Molmo multimodal modelsAI2 said its largest Molmo model trained on under a million image-text pairs, roughly three orders of magnitude less data than comparable systems, and released weights and data openly.
  3. DeepMind details how AlphaChip has shaped three generations of TPUsDeepMind said the reinforcement-learning layout method, adopted by MediaTek outside Google, had generated chip floorplans in hours that previously took engineers weeks.
  4. European Commission launches the AI Pact ahead of AI Act deadlinesOver a hundred companies signed pledges at launch to adopt an AI governance strategy and map high-risk systems ahead of binding rules; the pledges are voluntary and not legally enforceable.
  5. FTC launches 'Operation AI Comply' enforcement sweep against deceptive AI claimsFive cases targeted firms including DoNotPay and Rytr; the Rytr complaint passed 3-2 on a party-line vote, with Republican commissioners dissenting over the legal theory.
  6. Meta releases Llama 3.2 with vision and edge modelsThe 11B and 90B versions were Meta's first Llama models to accept images, while 1B and 3B text-only models were built for on-device use with a 128K-token context window.
  7. Mira Murati and two research leaders quit OpenAI on the same dayMurati cited wanting time for her own exploration; she went on to found Thinking Machines Lab, and Altman called the departures independent and amicable.
  8. Reuters reports OpenAI plans to convert to a for-profit structureReport says Altman would receive a stake for the first time and the nonprofit board would lose its controlling role, a plan not finalised until late 2025.
  9. Sam Altman publishes 'The Intelligence Age' essayAltman wrote that superintelligence could arrive within 'a few thousand days,' framing scaling deep learning as an already-solved algorithmic problem needing only more compute.
  10. Constellation Energy to restart Three Mile Island for Microsoft AI power dealThe Pennsylvania reactor, shut since 2019, is due back online in 2028 under a 20-year contract requiring NRC approval, supplying about 835MW.
  11. Alibaba releases Qwen2.5 model familyAlibaba's release spanned seven sizes from 0.5B to 72B parameters, plus dedicated coding and maths variants, trained on 18 trillion tokens.
  12. OpenAI updates safety and security practices after o1 releaseThe Safety and Security Committee became an independent board oversight body chaired by CMU professor Zico Kolter, with authority to delay model releases.
  13. OpenAI o1 results published on ARC-AGI-Pubo1-preview scored 21% on the public evaluation set, similar to Claude 3.5 Sonnet, but took roughly 70 hours to run 400 tasks against 30 minutes for either non-reasoning model.
  14. OpenAI releases o1, trading inference time for reasoningA model trained to think before answering opened a second scaling axis: spend more compute at inference and accuracy rises.
  15. OpenAI publishes the o1 system cardOpenAI's evaluation found 0.8% of o1-preview responses flagged as deceptive by an automated monitor, and rated the model medium risk for persuasion and CBRN.
  16. NotebookLM adds Audio Overviews, AI-generated podcast-style summariesTwo AI hosts discuss a user's uploaded documents in an unscripted-sounding conversation; Google called it experimental and English-only at launch.
  17. DeepMind's AlphaProteo designs novel protein bindersTrained on the Protein Data Bank and over 100 million AlphaFold-predicted structures, the system succeeded on a cancer-linked target, VEGF-A, where prior methods had failed entirely.
  18. DeepSeek merges chat and coder lines into DeepSeek-V2.5The merged model raised DeepSeek's ArenaHard win rate from 68.3% to 76.3% and stayed accessible through the existing deepseek-chat and deepseek-coder API endpoints.
  19. AI2 releases OLMoE mixture-of-experts model1 billion active of 7 billion total parameters, trained on 5 trillion tokens, released with 244 intermediate checkpoints and full training data and logs.
  20. Safe Superintelligence raises $1BThe round reportedly valued the two-and-a-half-month-old company at around $5 billion, paid entirely in cash despite SSI having shipped no public product.
  21. xAI opens Colossus supercomputer in MemphisxAI switched on 100,000 Nvidia H100 GPUs in a converted Memphis factory, a build the company said took roughly 122 days from empty shell to training-ready cluster.
  22. Anthropic hires its first AI welfare researcherFish, a co-author of the 'Taking AI Welfare Seriously' report, joined Anthropic's alignment science team; the company's public statement on model welfare followed roughly six weeks later.
  23. Study finds only a fraction of African languages supported by major AI modelsOf the top 34 languages used online worldwide, the World Economic Forum reported, none was African, and most AI systems are trained on only around 100 of the world's 7,000-plus languages.

August 2024

  1. LAION releases Re-LAION-5B with CSAM links removedThe cleaned dataset shrank from 5.8 to 5.5 billion pairs after matching hashes from the Internet Watch Foundation and the Canadian Centre for Child Protection, eight months after Stanford's report.
  2. Anthropic and OpenAI agree to model testing with US AI Safety InstituteThe memoranda gave the institute access to major new models from both companies before and after public release, mirroring an April agreement between the US and UK bodies.
  3. California legislature passes SB 1047The Assembly passed the bill 48-16 on 28 August and the Senate concurred 30-9 the next day, sending it to Governor Newsom, who vetoed it a month later.
  4. AI21 Labs releases Jamba 1.5Two sizes — a 94B and a 12B active-parameter mixture-of-experts model — built on AI21's Mamba-Transformer hybrid, both offering a 256K-token context window.
  5. Google opens Imagen 3 access to all US usersImagen 3, previewed at Google I/O in May, becomes available through ImageFX to all US users, with improved text rendering and fewer visual artefacts.
  6. OpenAI reports disrupting a covert Iranian influence operationThe network, tracked elsewhere as Storm-2035, used ChatGPT to draft US election commentary and Gaza-war content but drew almost no audience engagement.
  7. Nous Research releases Hermes 3Fine-tuned from Llama 3.1 at 8B, 70B and 405B parameters, with synthetic training data emphasising instruction-following, roleplay and function-calling for agents.
  8. OpenAI introduces SWE-bench Verified500 of the original benchmark's tasks, screened by 93 professional developers after OpenAI found 68% of samples had unfair tests or underspecified problems.
  9. San Francisco sues deepfake 'nudify' websitesThe suit named six defendants behind 16 'undressing' sites that had drawn more than 200 million visits in the first half of 2024; several later settled and shut down.
  10. Anthropic launches prompt caching in the Claude APIAnthropic ships prompt caching for the Claude API, cutting costs by up to 90% and latency by up to 85% on repeated long-context prompts.
  11. xAI releases Grok-2The beta release added image generation via Black Forest Labs' FLUX.1 and, within days, took second place on the LMSYS Chatbot Arena leaderboard behind GPT-4o.
  12. Judge lets core copyright claims against Stability AI proceed in AndersenJudge William Orrick dismissed DMCA claims but allowed direct copyright-infringement and inducement claims against Stability AI, Midjourney and DeviantArt to proceed to discovery.
  13. YouGov: nearly half of employed Americans expect AI to shrink jobs in their industryA YouGov survey found 48% of employed Americans expected AI to reduce jobs in their industry, up from 29% in March 2023, though only 2% reported it had already cost them work.
  14. Anysphere raises Series A at $400M valuationCursor-maker Anysphere raises a $60M+ Series A led by a16z and Thrive Capital, ten months after its seed round.
  15. Illinois amends Human Rights Act to restrict discriminatory AI in employmentThe law, House Bill 3773, also bars employers from using zip codes as a proxy for protected characteristics; it takes effect 1 January 2026.
  16. Anthropic launches invite-only bug bounty for jailbreak defencesApplications for the vetted red-teaming programme, run with HackerOne, closed on 16 August; it focused on jailbreaks touching CBRN and cybersecurity misuse.
  17. Irish DPC brings emergency court proceedings against X over Grok training dataIreland's DPC uses emergency powers for the first time to seek a High Court order over X training Grok on EU users' public posts without adequate consent; X agrees to stop.
  18. OpenAI publishes GPT-4o system cardThe card added voice-specific risk categories absent from text-only cards, including a classifier built to block the model from generating unauthorised voices.
  19. OpenAI adds structured outputs to the APIOpenAI said the new gpt-4o-2024-08-06 model scored 100% on its schema-following evaluation, against under 40% for gpt-4-0613 without the feature.
  20. Scaling test-time compute paper argues extra inference compute can beat bigger modelsOn some problems, extra inference-time computation matched the gains from a pretrained model roughly 14 times larger, the authors reported.
  21. Google takes Character.AI's founders backThe deal, reported at $2.5–2.7 billion, licensed Character.AI's technology to Google without buying the company, and drew Justice Department scrutiny.
  22. Black Forest Labs releases FLUXThe 12-billion-parameter FLUX.1 suite shipped in three tiers — a paid API model, an open non-commercial model and an Apache-licensed fast model — funded by $31 million in seed money.
  23. The EU AI Act enters into forceThe world's first comprehensive AI law took effect, banning some uses outright and imposing obligations on general-purpose models by capability.

July 2024

  1. Meta releases Segment Anything 2 (SAM 2)Released with the SA-V dataset of roughly 51,000 videos and 600,000+ masklets, more than four times the video count of the largest prior public segmentation dataset.
  2. NIST publishes Generative AI Profile for the AI Risk Management FrameworkIssued under Biden's October 2023 executive order, the voluntary profile applies NIST's 2023 risk framework specifically to generative systems.
  3. AlphaProof and AlphaGeometry 2 reach silver-medal standard at the IMOThe systems scored 28 of 42 points, one short of gold, but took up to three days on some problems against the competition's 4.5-hour limit.
  4. OpenAI releases a SearchGPT prototypeThe prototype answered queries with cited sources drawn from real-time web results and was opened to roughly 10,000 waitlisted testers and select publishers.
  5. Mistral AI releases Mistral Large 2The 123-billion-parameter model reported 84.0% on MMLU and was released under a non-commercial research licence, with a separate paid licence for commercial use.
  6. Meta releases Llama 3.1 405BTrained on over 15 trillion tokens with 16,000 H100 GPUs; Meta reported it competitive with GPT-4, GPT-4o and Claude 3.5 Sonnet, though the licence still barred some commercial uses.
  7. OpenAI releases GPT-4o miniPriced at 15 cents per million input tokens, more than 60% cheaper than GPT-3.5 Turbo, and became the default model for free ChatGPT users.
  8. Hugging Face releases SmolLMThe largest variant, 1.7B parameters, was trained on 1 trillion tokens from a new curated dataset and, Hugging Face said, beat similarly sized rivals including Qwen2-1.5B.
  9. StepFun launches Step-2, a trillion-parameter MoE modelStepFun said the mixture-of-experts model approximated GPT-4 on maths, logic, coding and dialogue; it was unveiled alongside a multimodal and an image-generation model at WAIC.

June 2024

  1. Amazon hires Adept AI's founders and licenses its technologyAdept continued operating independently under a new CEO after losing its co-founders, echoing the structure of Microsoft's earlier Inflection deal.
  2. ARC Prize introduces public ARC-AGI leaderboardUnlike the private-evaluation Kaggle competition, the leaderboard allows internet access and unlimited compute; early verified scores ranged from 42% down to 8-9% for frontier chatbots.
  3. Google launches Gemma 2, its open-weight model family, in 9B and 27B sizesThe 27B model ran on a single H100 GPU or TPU host, which Google said cut deployment cost while matching models more than twice its size.
  4. Microsoft discloses 'Skeleton Key' generative AI jailbreak techniqueFraming harmful requests as safety research and asking models to add a warning label rather than refuse worked against GPT-3.5, GPT-4o, Gemini Pro, Llama 3 and others; GPT-4 was comparatively resistant.
  5. Record labels sue Suno and UdioSuno later conceded training on copyrighted recordings but argued the use was transformative fair use, comparable to a person learning to write by reading.
  6. Claude 3.5 Sonnet and Artifacts change how people use chatbotsPriced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
  7. Sutskever founds Safe SuperintelligenceSutskever co-founded the lab with Daniel Gross and Daniel Levy, split between Palo Alto and Tel Aviv, one month after leaving OpenAI's board and chief-scientist role.
  8. NVIDIA becomes the world's most valuable companyNvidia's shares had risen more than ninefold since the end of 2022; the lead lasted days before Microsoft and Apple retook the position.
  9. Runway unveils Gen-3 AlphaTrained jointly on video and images on new infrastructure, it added fine-grained temporal control and C2PA provenance labelling absent from Runway's Gen-2 model.
  10. DeepSeek releases DeepSeek-Coder-V2The 236B-parameter mixture-of-experts model scored 90.2% on HumanEval, edging out GPT-4-Turbo's 88.2%, while running with only 21B parameters active per token.
  11. Luma AI launches Dream MachineFree-to-use video model generating five-second clips at 1360x752 from text or image prompts, drawing early comparisons to OpenAI's unreleased Sora.
  12. Stability AI ships Stable Diffusion 3 Medium weightsThe 2-billion-parameter weights ran on consumer GPUs but were licensed non-commercially, with a separate paid Enterprise tier for large-scale business use.
  13. Announcing ARC Prize 2024The best public score on ARC-AGI stood at 34%, up from 20% when Chollet introduced the benchmark five years earlier, still well below typical human performance.
  14. Mistral AI closes €600M Series BGeneral Catalyst led the €600M round, split between €468M in equity and €132M in debt, at a reported $6 billion valuation, a year after a €112M seed.
  15. Apple announces Apple Intelligence with OpenAI insideChatGPT access was free and optional, required no account, and Apple said queries would not be logged — with other AI providers to be added later.
  16. Alibaba releases Qwen2Five model sizes from 0.5B to 72B parameters, trained on 27 additional languages beyond English and Chinese, with the smaller sizes under Apache 2.0.
  17. OpenAI publishes early sparse-autoencoder work extracting concepts from GPT-4The 16-million-feature autoencoder cost roughly as much accuracy as training GPT-4 with ten times less compute, illustrating interpretability's overhead at scale.
  18. Current and former staff demand a right to warnThirteen current and former employees of OpenAI, Google DeepMind and Anthropic signed; Bengio, Hinton and Russell endorsed it without being employees themselves.
  19. Leopold Aschenbrenner publishes 'Situational Awareness'Aschenbrenner, dismissed from OpenAI's superalignment team months earlier for allegedly leaking information, argued the firing itself illustrated the security failures he described.
  20. MMLU-Pro benchmark paper releasedThe paper reported chain-of-thought reasoning helped on the new benchmark where it had made little difference on the original MMLU, and cut prompt-sensitivity from 4-5 points to about 2.
  21. Ipsos finds world split on AI benefits, with Asia most optimistic and West most sceptical83% in China, 80% in Indonesia and 77% in Thailand saw AI products as more beneficial than harmful, against 40% in Canada, 39% in the US and 36% in the Netherlands.

May 2024

  1. Hugging Face releases FineWeb datasetBuilt from 96 Common Crawl snapshots and released under an open licence, the corpus was accompanied by FineWeb-Edu, a smaller subset filtered for educational value.
  2. Google scales back AI Overviews after viral bad-advice answersSearch head Liz Reid described more than a dozen technical fixes, including limiting satirical sources and pausing overviews on health queries.
  3. OpenAI reports first disruption of covert influence operations using its modelsOne Russian network's posts still carried the model's own refusal messages, which a researcher cited as evidence the operations were poorly executed rather than AI-supercharged.
  4. Epoch AI publishes 'Training compute of frontier AI models grows by 4-5x per year'The estimate drew on 333 compute figures for notable models since 2010 — roughly triple Epoch's 2022 dataset — and flagged an unexplained slowdown around 2018.
  5. Helen Toner gives first detailed public account of the OpenAI board's decision to fire AltmanToner said the board first learned of ChatGPT's launch from Twitter and that Altman tried to have her removed after she co-wrote a critical paper.
  6. OpenAI's board forms a Safety and Security CommitteeChaired by board chair Bret Taylor and including CEO Sam Altman as a member, the committee had 90 days to review safety practices before recommending changes to the full board.
  7. xAI raises $6bn Series B at $24bn valuationInvestors including Valor Equity, Vy Capital, Andreessen Horowitz, Sequoia and Fidelity backed the round, roughly a year after xAI's founding, to fund a supercomputer for xAI's next model.
  8. Google's AI Overviews tells users to eat rocks and put glue on pizzaScreenshots showed the feature had sourced answers from an Onion satire piece and an 11-year-old Reddit joke, days after its US launch.
  9. OpenAI and News Corp sign multi-year global content partnershipNeither company disclosed terms, but the Wall Street Journal reported the deal could be worth over $250 million across five years, among the largest AI publisher agreements to date.
  10. EU AI Act formally adopted by the CouncilThe Council's sign-off was the final legislative step; the text still needed formal signature, Official Journal publication, and a staggered two-year rollout before most rules applied.
  11. Anthropic maps millions of concepts inside a production modelSparse autoencoders extracted human-interpretable features from a deployed model — and turning one up produced Golden Gate Claude.
  12. Sixteen companies sign the Frontier AI Safety Commitments in SeoulSignatories pledged to publish safety frameworks defining risk thresholds and to not deploy a model if those risks could not be mitigated below them.
  13. Suno raises $125M Series B at $500M valuationSuno said 10 million people had made music on the platform in its first eight months, ahead of a round led by Lightspeed with backing from Nat Friedman and Daniel Gross.
  14. Inflection AI announces new leadership and enterprise pivotSean White, formerly Mozilla's chief technology officer, became CEO two months after Microsoft hired co-founder Mustafa Suleyman and most of Inflection's research staff.
  15. Microsoft announces Copilot+ PCs and RecallRecall continuously screenshots the desktop to build a searchable history; security researchers found the local database unencrypted within days, and Microsoft delayed the rollout.
  16. OpenAI pulls ChatGPT voice after Scarlett Johansson says it copied her from 'Her'Altman had asked Johansson to voice ChatGPT in September 2023; she declined, and said she was 'shocked' when a nearly indistinguishable voice, Sky, shipped anyway.
  17. OpenAI's exit agreements are found to claw back equityVox reported departing staff had to sign lifetime non-disparagement terms within 60 days or lose vested equity potentially worth millions; Altman said he had not known.
  18. Colorado enacts first US broad AI discrimination lawSB 24-205 requires risk-management programmes and impact assessments for 'high-risk' AI in hiring, lending and similar decisions; Polis signed it 'with reservations.'
  19. DeepMind publishes the Frontier Safety FrameworkA set of internal capability thresholds across autonomy, cybersecurity, biosecurity and ML R&D, joining similar voluntary policies already published by Anthropic and OpenAI.
  20. Jan Leike resigns and the superalignment team dissolvesLeike said his team had been 'sailing against the wind' for compute and access; OpenAI reassigned remaining members rather than replacing the team's leadership.
  21. Yoshua Bengio-led International Scientific Report on the Safety of Advanced AI published75 experts from 30 countries plus the EU and UN produced an IPCC-style synthesis of AI risk evidence, timed to inform the Seoul summit five days later.
  22. Voice actors sue AI startup Lovo over unauthorised voice cloningLehrman and Sage say they were hired via Fiverr for research use only, but Lovo cloned their voices into 'Kyle Snow' and 'Sally Coleman' and resold them to customers.
  23. DeepMind extends SynthID watermarking to AI-generated text and videoThe scheme adjusts token-selection probabilities to embed a statistical mark invisible to readers; DeepMind said detection degrades on short, factual or translated text.
  24. Google announces Trillium, its sixth-generation TPUTrillium delivers a claimed 4.7x peak compute increase per chip over the prior TPU generation and trains models including Gemini 1.5 Flash and Gemma 2.
  25. Google DeepMind unveils Veo, a text-to-video generation modelGenerates 1080p video over a minute long from text and image prompts, watermarked with SynthID; launched in private preview via waitlist rather than public release.
  26. Google puts AI Overviews on searchA Gemini-generated summary above the links, rolled out to hundreds of millions of US users that week with a target of a billion by year end.
  27. Google unveils Project Astra, a universal AI assistant prototypeA prototype, not a product: Google showed a phone-camera assistant with conversational-speed responses but gave no release date beyond 'later this year'.
  28. OpenAI launches GPT-4o with real-time voiceA single model handling text, vision and audio end to end, with conversational latency — and a voice that led to a public dispute with Scarlett Johansson.
  29. Deepfake voice and video scam attempt targets WPP CEO Mark ReadScammers combined a cloned voice, YouTube footage and a fake WhatsApp-arranged Teams meeting to impersonate the CEO; an agency leader grew suspicious and no money changed hands.
  30. AlphaFold 3 predicts structures across proteins, DNA, RNA and ligandsRestricted at launch to a rate-limited web server rather than downloadable code, prompting an open letter with more than 650 signatures within a week.
  31. OpenAI introduces the Model SpecA public document defining how OpenAI wants its models to behave, including a chain-of-command rule that developer instructions override user ones; opened for public comment.
  32. Redwood Research publishes 'The case for ensuring that powerful AIs are controlled'Argued labs should assume some deployed models may be misaligned and build restrictions that hold even if a model actively tries to subvert them, distinct from alignment itself.
  33. DeepSeek releases DeepSeek-V2A 236-billion-parameter mixture-of-experts model with only 21 billion active per token, released open-weight; DeepSeek said training costs fell 42% versus its prior model.
  34. OECD updates AI Principles for generative AI eraThe revision renames 'human-centred values' as 'human rights and democratic values' and adds clauses on misinformation and safety incidents, but remains non-binding.
  35. Anthropic launches the Claude Team planPriced at $30 per seat monthly with a five-seat minimum, the plan launched alongside Anthropic's first iOS app and gave every seat access to Opus, Sonnet and Haiku.
  36. Judge dismisses bloated privacy class action against OpenAI and MicrosoftIn Cousart v. OpenAI, Judge Vince Chhabria gave plaintiffs 21 days to refile a shorter complaint, calling the 204-page document nearly impossible to parse.
  37. Scale AI publishes GSM1k contamination study of GSM8KA fresh grade-school-maths test found some open models scored up to 13 points lower than on GSM8K, evidence of memorisation, while frontier models showed little gap.

April 2024

  1. Meta releases Llama 3 and puts its assistant everywhere8B and 70B open-weight models shipped alongside a much larger, still-training 400B+ version, as Meta AI rolled out across Facebook, Instagram, WhatsApp and Messenger.
  2. Microsoft Research shows VASA-1, real-time talking-face generationTrained to render 512x512 video at up to 40 frames per second from one photo and an audio clip; Microsoft said it had no plans to release a demo or model.
  3. Microsoft invests $1.5bn in UAE's G42G42 agreed to remove Chinese technology from its systems, and Microsoft's Brad Smith joined its board, as part of a deal read as a US bid for Gulf AI infrastructure.
  4. Stanford HAI releases 2024 AI Index ReportThe report put GPT-4's training compute cost at roughly $78 million and Gemini Ultra's at $191 million, and found industry produced 51 notable models in 2023 to academia's 15.
  5. UK Society of Authors survey: a third of translators losing work to generative AI787 members surveyed; 36% of translators and 26% of illustrators reported lost work, while over a third of translators had already used generative AI themselves.
  6. OSWorld benchmarks AI agents on real desktop computer tasks369 tasks across real Ubuntu, Windows and macOS applications, graded on the machine's actual end-state; the best model at release solved 12% against a 72% human baseline.
  7. UK competition regulator warns on AI foundation model market concentrationThe regulator counted more than 90 partnerships and investments linking six firms — Google, Apple, Microsoft, Meta, Amazon and Nvidia — across the foundation-model supply chain.
  8. Cohere releases Command R+Launched first on Microsoft Azure, the larger companion to March's Command R kept the same 128K context but was pitched as competitive with GPT-4-class models.
  9. Anthropic publishes 'Many-shot Jailbreaking' researchStuffing a prompt with dozens of faked harmful-request dialogues broke safety training in a power-law pattern as context windows grew past a million tokens.
  10. Udio launches in betaBacked by Andreessen Horowitz and artists including will.i.am, the free text-to-music beta arrived weeks after rival Suno's v3, intensifying competition in AI music.
  11. UK and US AI Safety Institutes sign testing partnershipUK science secretary Michelle Donelan and US commerce secretary Gina Raimondo signed the pact, which followed through on commitments made at the November 2023 Bletchley summit.

March 2024

  1. OpenAI launches Voice Engine research preview and its safety approachOpenAI said it had held the text-to-speech model, first developed in 2022, back from wider release specifically because of election-year impersonation risk.
  2. Reports emerge that Microsoft and OpenAI are planning a $100 billion 'Stargate' data centreThe Information reported a five-phase supercomputer roadmap, with Stargate as the final phase and a smaller fourth-phase machine targeted for 2026.
  3. AI21 Labs releases Jamba, a hybrid SSM-Transformer modelCombining Mamba state-space layers with transformer attention, it handled a 256K-token context and fit on a single 80GB GPU, which AI21 said pure transformers could not match.
  4. OMB issues governance memo for federal agency use of AIAgencies had until September 2024 to publish compliance plans and until December to make their AI use-case inventories public, under an order later revoked.
  5. Databricks releases DBRXBuilt by the former MosaicML team Databricks had acquired the previous year, the 132-billion-parameter model activated only 36 billion parameters per input.
  6. Stability AI CEO Emad Mostaque resignsHe left amid a cash crunch and an exodus of core Stable Diffusion researchers, who months later founded the rival Black Forest Labs, saying centralised AI needed challenging.
  7. Suno releases v3The company called it its first model able to produce radio-quality output, letting free users generate two-minute songs and paying users four-minute tracks.
  8. UN General Assembly adopts first resolution on AIThe non-binding, US-led text won more than 120 co-sponsors and passed by consensus without a vote, urging states not to deploy AI that violates human rights.
  9. Microsoft absorbs Inflection's team without buying the companyReported terms put the deal at $650 million — most of it a licence fee for Inflection's models, the rest for waiving legal claims over the mass hire.
  10. NVIDIA announces the Blackwell architectureThe GB200 NVL72 system was claimed to cut LLM inference cost and energy per token by up to 25 times against Hopper-generation hardware.
  11. xAI open-sources Grok-1A 314-billion-parameter mixture-of-experts base model, unfine-tuned and released under Apache 2.0, using only a quarter of its weights per token.
  12. DeepMind's SIMA agent follows instructions across 3D game worldsTrained across nine commercial games and four research environments, the agent read only screen pixels and text instructions, with no access to game code or APIs.
  13. The European Parliament passes the AI ActAdopted 523 votes to 46 after three years of negotiation; the Council's formal sign-off and publication in the Official Journal followed months later.
  14. Utah enacts first state law specifically regulating generative AIThe law, effective 1 May 2024, created a state Office of Artificial Intelligence Policy and a regulatory sandbox alongside its consumer-disclosure duty.
  15. Cognition demos Devin, billed as the first AI software engineerCognition said Devin resolved 13.86% of real GitHub issues unassisted on the SWE-bench benchmark, against roughly 2% for the prior best system.
  16. LiveCodeBench paper publishedTesting 52 models against problems tagged by publication date, the authors found evidence that some scored higher on problems predating their training cutoff.
  17. Cohere releases Command RThe 35-billion-parameter model targeted enterprise retrieval-augmented generation and tool use with a 128K-token context window, ahead of a larger Command R+ in April.
  18. Gladstone AI's government-commissioned report warns of catastrophic AI riskThe four-person consultancy, paid $250,000 by the State Department, urged Congress to ban training runs above a compute threshold and restrict publishing open-weight models.
  19. Researchers demonstrate practical model-stealing attack against production LLM APIsThe team, led by Nicholas Carlini, also recovered the hidden-layer size of Google's PaLM-2 and OpenAI's gpt-3.5-turbo through ordinary API queries.
  20. OpenAI board's review concludes, Altman and Brockman to continue leading OpenAIWilmerHale's inquiry found Altman's November 2023 removal was not about safety, security, finances or statements to investors, and three new directors joined the board.
  21. OpenAI responds to Elon Musk's lawsuit over its founding missionOpenAI said Musk had contributed $45 million rather than the $1 billion he pledged, and published emails in which he proposed merging OpenAI into Tesla or taking full control.
  22. Anthropic's Claude 3 takes the frontier from GPT-4The first time a lab other than OpenAI held the top spot on headline benchmarks, and the start of the small/medium/large release pattern.
  23. Physical Intelligence raises $70M seedThe startup, building vision-language-action models to control robots, was backed by Thrive Capital, OpenAI, Khosla Ventures and Sequoia at roughly a $400M valuation.
  24. Reflection AI foundedCo-founder Ioannis Antonoglou had been a core architect of AlphaGo at DeepMind; the pair set out to build autonomous coding agents.

February 2024

  1. Figure raises $675M Series B, partners with OpenAIInvestors included Microsoft, Nvidia and Jeff Bezos; OpenAI agreed to develop language and reasoning models specifically for Figure's humanoid robots.
  2. Musk sues OpenAI over its for-profit turnThe complaint alleged a founding pact to develop AGI 'for the benefit of humanity' had been broken for Microsoft's profit; OpenAI said no such agreement existed.
  3. Hugging Face releases StarCoder2The Stack v2 dataset grew to roughly ten times the size of its predecessor, and BigCode said the 15B model matched benchmarks of models more than twice its size.
  4. Klarna says AI does the work of 700 customer service agentsThe buy-now-pay-later firm said its assistant handled 2.3 million chats in a month, cut resolution time from 11 minutes to under two, and projected $40m in 2024 profit gains.
  5. Mistral AI releases Mistral Large and launches Le ChatThe French start-up's flagship model was closed-weight rather than open, and Microsoft took a stake and put it on Azure the same day.
  6. Google suspends Gemini image generation of peopleDepicting America's Founding Fathers and German World War Two soldiers as people of colour drew accusations of overcorrected diversity tuning; Google paused the feature to fix it.
  7. Stability AI previews Stable Diffusion 3The suite spans 800 million to 8 billion parameters and combines a diffusion transformer with flow matching, aimed chiefly at fixing text rendering and multi-subject prompts.
  8. European Commission launches the European AI OfficeThe Office, with more than 125 staff across six units, can investigate suspected violations, request technical documentation and issue fines under the AI Act.
  9. Google releases Gemma open weightsReleased as 2B and 7B models under terms permitting commercial use, Google said Gemma 7B outperformed the larger Llama 2 13B on standard benchmarks.
  10. Anthropic publishes election safeguards for the 2024 US electionAnthropic reported a system-prompt fix for time-sensitive queries and fine-tuning that increased referrals to authoritative voting-information sources.
  11. Gemini 1.5 Pro ships a million-token context windowA mixture-of-experts model that matched Gemini 1.0 Ultra on many tasks at lower compute, offered in limited preview with up to a million tokens of context.
  12. Meta releases V-JEPATrained without labels, the model learned by predicting masked video regions in representation space, distinguishing fine-grained actions over roughly ten-second clips.
  13. OpenAI previews Sora, a text-to-video modelClips up to a minute long held their subjects consistent across camera moves; OpenAI showed samples but gave no public access for ten months.
  14. Air Canada ordered to honour refund its chatbot wrongly promisedThe claim was for CAD $650.88; the tribunal rejected Air Canada's argument the chatbot was 'a separate legal entity responsible for its own actions.'
  15. Microsoft and OpenAI disrupt state-affiliated hacking groups misusing LLMsThe groups used LLMs mainly for reconnaissance, translation, debugging and phishing-content drafting rather than novel attack techniques; the identified accounts were terminated.
  16. OpenAI adds memory and new data controls to ChatGPTPiloted with a small user group rather than shipped broadly, the feature let ChatGPT carry facts between separate conversations, with settings to view, forget or turn it off entirely.
  17. Judge dismisses most of Silverman's book-author lawsuit against OpenAIThe judge found the authors had not shown ChatGPT's output resembled their books closely enough, but let the core training-data infringement claim proceed.
  18. FCC declares AI-generated robocall voices illegal under TCPAThe unanimous ruling followed a faked Biden robocall urging New Hampshire primary voters to skip voting, giving state attorneys general new grounds to prosecute voice-cloning scams.
  19. Google rebrands Duet AI as Gemini for WorkspaceGemini Business, at $20 a month, and Gemini Enterprise, at $30, replaced Duet AI's Workspace add-ons, folding Gmail, Docs and Meet features under one brand.
  20. Google renames Bard to Gemini and launches Gemini Advanced with Ultra 1.0The $19.99-a-month Google One AI Premium tier gave access to Ultra 1.0, which Google said was the first model to outperform human experts on MMLU.
  21. UK announces over £100m for AI regulators and researchOf the total, £10m went to regulators such as Ofcom, £90m to nine sector research hubs, and £9m to a new AI partnership with the US.
  22. AI2 releases OLMo, a fully open language modelUnlike other 'open' releases, AI2 published the full Dolma training corpus, training code and hundreds of intermediate checkpoints, not just final weights.
  23. DeepSeek publishes DeepSeekMath, introducing GRPOThe 7B model reached 51.7% on the MATH benchmark without external tools, and its GRPO training method later underpinned DeepSeek-R1's reasoning training.

January 2024

  1. George Carlin estate sues over AI-generated fake stand-up specialThe Dudesy podcast, hosted by Will Sasso and Chad Kultgen, took the special down within a week of the suit and agreed to a permanent injunction.
  2. OpenAI's new embedding models and API updatestext-embedding-3-small was priced far below OpenAI's previous embedding model, and GPT-3.5 Turbo's input price was cut 50% in its third reduction in a year.
  3. Sexual deepfakes of Taylor Swift spread on XOne post was viewed over 47 million times before removal; X temporarily blocked searches of her name and the White House urged Congress to legislate.
  4. AI-generated fake Biden robocall urges New Hampshire voters to skip primaryThousands of New Hampshire Democrats received a cloned Biden voice telling them a primary vote would forfeit their choice in November; the operative behind it was fined and criminally charged.
  5. AlphaGeometry solves olympiad geometry problems near gold-medal levelThe system solved 25 of 30 benchmark problems within competition time limits, versus the 25.9 average for human gold medalists and 10 for the prior best system.
  6. Zhipu launches GLM-4, claiming near-GPT-4 parityUnveiled at Zhipu's first DevDay with a 128K-token context window, alongside a $100 million fund the company pledged to LLM startups.
  7. OpenAI publishes its approach to 2024 worldwide electionsMeasures included routing US voting questions to CanIVote.org via a partnership with state election officials and a ban on political-campaign chatbots.
  8. Anthropic shows backdoored models surviving safety trainingModels trained to write secure code unless told the year was 2024 kept the hidden behaviour through supervised fine-tuning, reinforcement learning and adversarial training.
  9. OpenAI opens the GPT StoreA marketplace for custom assistants, launched two months late after the board crisis, alongside a new $25-a-month ChatGPT Team subscription tier.
  10. Judge narrows GitHub Copilot lawsuit, dismissing most claimsJudge Jon Tigar dismissed negligence, interference and most DMCA claims over Copilot's training data, leaving an open-source licence claim to proceed.
  11. AI-generated books impersonating real authors flood AmazonThe Authors Guild documented AI-written summaries and knockoffs appearing on Amazon within a day of a book's real release, naming Kara Swisher and others as targets.
  12. Artificial Analysis launches independent model benchmarking siteFounded by George Cameron and Micah Hill-Smith as a side project comparing model pricing and latency, it became a widely cited independent reference.
  13. Deepfake video call used to steal $25 million from Arup's Hong Kong officeA finance employee made 15 wire transfers after a video call in which every other participant was AI-generated; Hong Kong police confirmed the case a week later.
  14. Ireland datacentre opposition intensifies as facilities near quarter of national electricity useIreland's Central Statistics Office found datacentres consumed 21% of the country's metered electricity in 2023, more than urban or rural households individually.

December 2023

  1. The New York Times sues OpenAI and MicrosoftThe first major news organisation to sue over AI training data, alleging models could reproduce its articles near-verbatim.
  2. Stanford researchers find CSAM links in LAION-5B, dataset pulledScanning roughly 32 million data points, researchers flagged 3,226 suspected matches and externally validated 1,008 using hash-matching services before LAION withdrew the dataset.
  3. Suno launches public web appCambridge, Massachusetts-based Suno, founded by ex-Kensho engineers, opens its AI song-generation web app to the public, alongside a Microsoft Copilot integration.
  4. UK Supreme Court rules AI cannot be a patent inventorThe court left open whether AI-generated inventions should be patentable at all, ruling only on the narrower question of who the Patents Act 1977 requires to be named as inventor.
  5. FTC bans Rite Aid from using AI facial recognition after misidentifying customersFTC settles with Rite Aid, banning its use of facial recognition for five years after finding the AI system generated false shoplifting matches without safeguards.
  6. OpenAI publishes its Preparedness FrameworkThe beta framework scored models low to critical on four risk categories, barring deployment above 'high' and barring further development above 'critical,' with the board holding final oversight.
  7. Chevrolet dealership chatbot tricked into 'agreeing' to sell Tahoe for $1A prompt-injection prank made a Chevrolet dealership's ChatGPT-based sales chatbot appear to accept a $1 offer on a $60,000+ Tahoe as a 'legally binding' deal; the dealer shut the bot down.
  8. FunSearch makes a mathematical discovery with an LLMPairing a code-writing model with an automated evaluator produced a genuinely new cap-set construction and improved bin-packing heuristics.
  9. OpenAI publishes 'Weak-to-Strong Generalization' superalignment paperFine-tuning GPT-4 on labels from a GPT-2-sized supervisor recovered close to GPT-3.5-level performance on language tasks, but the technique still struggled on chess puzzles and reward modelling.
  10. EU negotiators strike a political deal on the AI ActFoundation models and police use of biometric surveillance were the last sticking points; the deal set tiered duties for both, with the text finalised in 2024.
  11. Mistral releases Mixtral 8x7BA sparse mixture-of-experts model with roughly 45B total parameters, released under Apache 2.0, that Hugging Face said matched GPT-3.5-turbo on MT-Bench.
  12. Meta launches Purple Llama for open model safety toolingMeta launched Purple Llama, an umbrella project including the Llama Guard safety classifier and CyberSecEval cybersecurity benchmarks for open generative AI.
  13. Google launches GeminiGoogle said Gemini Ultra beat human experts on the MMLU benchmark; days later Bloomberg reported the model's showcase video had been edited and was not real-time.
  14. Harvey raises $80M Series B at $715M valuationCo-led by Elad Gil and Kleiner Perkins, the round brought total funding past $100M; Harvey said revenue had risen more than tenfold since April.
  15. Mamba paper proposes selective state-space models as a transformer alternativeLetting the model's internal state-update rules depend on the input let a 3-billion-parameter Mamba match transformers twice its size while running five times faster.

November 2023

  1. DeepSeek releases DeepSeek LLM 67BDeepSeek's first general-purpose open-weight LLM family, 7B and 67B, trained on 2 trillion English/Chinese tokens.
  2. GNoME finds 2.2 million candidate new materialsDeepMind flagged 380,000 of the 2.2 million predicted crystal structures as most stable; independent labs had already synthesised 736 by the time of publication.
  3. Together AI raises $102.5m Series ATogether AI raised a $102.5m Series A led by Kleiner Perkins with NVIDIA participating, valuing the open-model cloud provider at roughly $500m.
  4. Beijing court grants copyright to an AI-generated image, splitting with the USThe Beijing Internet Court credited the prompter, not the AI or its developer, with authorship — the opposite conclusion the US Copyright Office had reached on Stable Diffusion images months earlier.
  5. Pika Labs launches Pika 1.0, raises $55MFounded by two Stanford computer-science PhD students, the video-editing app had already drawn some 500,000 users through its earlier Discord-only release.
  6. Sports Illustrated deletes articles bylined to fake, AI-generated authorsFuturism traced the profile photo of one fictitious byline, 'Drew Ortiz,' to a marketplace selling AI-generated headshots; the publisher's CEO was fired two weeks later.
  7. Sam Altman returns as CEO with a new initial boardThe three-person interim board — Bret Taylor, Larry Summers and Adam D'Angelo, the only holdover from the board that fired him — replaced the four who had voted him out.
  8. Anthropic releases Claude 2.1Claude 2.1 ships with a 200K-token context window, reduced hallucination rates and tool use support.
  9. GAIA, a benchmark for general AI assistants, is released466 questions that are simple for a person but need browsing, tools and multi-step reasoning to solve; GPT-4 with plugins scored 15% against a 92% human baseline at release.
  10. GPQA graduate-level 'Google-proof' benchmark publishedPhD-level experts scored 65% and skilled non-experts with unrestricted web access for over 30 minutes managed only 34%; GPT-4 reached 39%.
  11. OpenAI's board fires Sam Altman, and reinstates him five days laterThe board cited a loss of confidence but gave no detail; around 700 of roughly 770 employees threatened to resign, and the board itself was replaced.
  12. Microsoft unveils in-house Maia 100 AI chip and Cobalt 100 CPUAnnounced at Microsoft's Ignite conference, the pair of custom chips were slated to reach Microsoft datacentres in early 2024, starting with internal workloads.
  13. GraphCast beats conventional weather forecasting on speed and accuracyTrained on four decades of reanalysis data, the model beat the ECMWF's physics-based system on over 90% of tested variables while running on a single TPU.
  14. SAG-AFTRA's 2023 film/TV contract sets consent rules for digital replicasThe deal required separate, informed consent and compensation before a studio could create or reuse a performer's AI-generated digital double.
  15. OpenAI's first DevDay ships GPTs and an Assistants APIGPT-4 Turbo cut input pricing to a cent per thousand tokens and extended context to 128,000 tokens; the promised GPT Store did not open until 2024.
  16. Google DeepMind proposes a 'Levels of AGI' frameworkSix tiers from 'no AI' to 'superhuman', scored across narrow and general tasks, aimed to replace binary AGI-or-not debate with a shared measurement vocabulary.
  17. xAI unveils GrokA chatbot with real-time access to X posts and a deliberately irreverent persona, initially released to a small group ahead of a wider subscriber rollout.
  18. 01.AI open-sources Yi-6B and Yi-34BKai-Fu Lee's 01.AI released its first open-weight models, which it said outperformed larger Llama 2 and Falcon models.
  19. The UK opens the first state AI Safety InstituteBacked by a £300 million compute allocation and chaired by Ian Hogarth, the institute converted a temporary taskforce into a permanent evaluator with formal US and Singapore partnerships.
  20. Andrej Karpathy delivers 'Intro to Large Language Models' talkKarpathy illustrated a model as two files, a parameters file and a short run program, and sketched an 'LLM-OS' coordinating tools, memory and multimodal input.
  21. The Bletchley Declaration at the first AI Safety SummitTwenty-eight countries including the US and China signed the first international statement on frontier AI risk.
  22. DeepSeek releases DeepSeek CoderThe 33B version outperformed CodeLlama-34B on coding benchmarks and, once instruction-tuned, beat GPT-3.5-turbo on HumanEval — DeepSeek's first public model release.

October 2023

  1. Biden signs the executive order on safe and trustworthy AIThe most far-reaching US action on AI to date: compute thresholds, mandatory safety reporting, and a new safety institute.
  2. The G7 agrees a Hiroshima code of conduct for AIEleven non-binding principles for advanced-AI developers, carrying no legal force in any G7 jurisdiction, published the same day as the US executive order.
  3. Hugging Face's H4 team releases Zephyr-7BFine-tuned from Mistral 7B using AI-generated preference data and no human annotation, it scored 7.34 on MT-Bench against Llama 2 70B Chat's 6.86.
  4. iFlytek releases Spark V3.0iFlytek said the update matched ChatGPT 3.5 in English and surpassed it in Chinese, while promising a GPT-4-competitive model in the first half of 2024.
  5. Researchers release Nightshade, a data-poisoning tool for artists against AI scrapersAround 50 poisoned images could distort a diffusion model's output for a concept; the effect spread to related words such as 'puppy' when a model learned from enough of them.
  6. Music publishers sue Anthropic over song lyricsUniversal, Concord and ABKCO alleged Claude reproduced lyrics from at least 500 songs, including a near-identical copy of Katy Perry's 'Roar', and sought statutory damages.
  7. Baidu unveils ERNIE 4.0Baidu called it its most capable model and claimed parity with GPT-4; access was invitation-only at launch, with enterprise API testing through the Qianfan platform.
  8. Washington tightens the chip controls againNew performance-density thresholds targeted Nvidia's China-specific A800 and H800 parts, and licensing requirements extended to 21 additional countries.
  9. Marc Andreessen publishes the 'Techno-Optimist Manifesto'The essay named existential risk, the precautionary principle, tech ethics and trust and safety as ideas its author considered enemies of progress.
  10. SWE-bench paper publishedBuilt from 2,294 real GitHub issues across 12 Python repositories, the benchmark proved so hard that the best model of the day, Claude 2, solved under 2%.
  11. WGA ratifies contract with new guardrails on AI in screenwritingMembers approved the deal 99% to 1%, on 8,525 votes cast, five months after AI protections became a central demand of the strike that began in May.
  12. Anthropic publishes 'Towards Monosemanticity'Sparse autoencoders decomposed a single 512-neuron layer into more than 4,000 human-interpretable features, far more than the raw neurons showed.
  13. Researchers show GPT-4 safety guardrails easily bypassed via low-resource languagesResearchers traced the gap to RLHF: human annotators who flag harmful content overwhelmingly work in high-resource languages, leaving weaker refusal training elsewhere.
  14. Moonshot AI launches Kimi chatbot with 128K contextMoonshot AI, a Tsinghua-linked startup founded that March, said Kimi could process 200,000 Chinese characters of input, a claimed world first for a consumer chatbot.

September 2023

  1. Mistral 7B beats larger models and ships by torrentA magnet link with no blog post or announcement stood in deliberate contrast to the polished launches of Meta and Google, and the 7-billion-parameter model still beat Llama 2 13B.
  2. Amazon invests up to $4 billion in AnthropicAWS became Anthropic's primary cloud provider and a supplier of its Trainium and Inferentia training chips, giving the lab a second hyperscaler backer alongside Google.
  3. ChatGPT gains voice and visionA new text-to-speech model, built with professional voice actors, let users talk to ChatGPT and show it photos, ahead of the fully real-time GPT-4o release.
  4. OpenAI publishes the GPT-4V(ision) system cardA three-month alpha and over 50 outside experts probing high-risk areas informed mitigations against uses such as CAPTCHA-solving and identifying people from photos.
  5. OpenAI releases DALL·E 3ChatGPT itself was positioned as the prompt engineer, drafting detailed image prompts from short user requests before rendering them; rollout began the following month.
  6. AlphaMissense catalogues 71 million genetic variants for disease riskAdapted from AlphaFold, the model classified 89% of all possible human missense variants as likely pathogenic or benign, versus 0.1% confirmed by human experts.
  7. Anthropic establishes the Long-Term Benefit TrustA special stock class gives independent trustees, holding no equity, the right to elect a majority of Anthropic's board within four years.
  8. Anthropic publishes its Responsible Scaling PolicyAI Safety Levels borrowed the biosafety-lab naming scheme, and its rules would eventually pause deployment of any model reaching a level the company had not yet built safeguards for.
  9. OpenAI launches Red Teaming NetworkRecruits needed no prior AI experience, only domain expertise such as linguistics or biometrics; applications for the first cohort closed that December.
  10. The Authors Guild sues OpenAISeventeen novelists including Grisham and Martin filed a class action in the Southern District of New York alleging systematic copying of pirated books to train GPT.
  11. Schumer convenes tech chief executives behind closed doorsAbout two-thirds of the Senate attended; press and public were excluded, and no senator was permitted to ask a question.
  12. UNESCO publishes first global guidance on generative AI in educationThe guidance recommended a minimum age of 13 for unsupervised use of chatbots in the classroom and called on all 193 member states to regulate generative AI in schools.
  13. TII releases Falcon 180BAt 180 billion parameters, trained on 3.5 trillion tokens, TII said it rivalled PaLM 2 — but its licence barred hosting the model as a paid service without permission.
  14. Google publishes RLAIF paper comparing AI-feedback to human-feedback alignmentReward models trained on preference labels from another LLM matched human-feedback RLHF on summarisation and dialogue tasks, and a variant skipping the reward model entirely did better still.
  15. Tencent releases first Hunyuan large modelUnveiled at Tencent's Global Digital Ecosystem Summit, the mixture-of-experts model exceeded 100 billion parameters and launched for enterprise access via Tencent Cloud, not consumers.

August 2023

  1. Baidu's ERNIE Bot receives regulatory approval for public releaseApproval followed China's July 2023 rule requiring a licence before releasing generative-AI models to the public; ERNIE Bot topped Apple's China App Store within hours.
  2. Gannett pauses AI-generated high school sports articles after mockeryRecaps from the vendor LedeAI, published by Gannett's Columbus Dispatch and other local papers, included unfilled placeholder text such as '[[WINNING_TEAM_MASCOT]]'.
  3. DeepMind launches SynthID to watermark AI-generated imagesThe imperceptible pixel-level watermark launched in beta for a limited set of Vertex AI customers using Google's Imagen model.
  4. OpenAI launches ChatGPT EnterpriseThe tier offered unlimited GPT-4 access at double speed, a 32,000-token context window and a pledge not to train on business data.
  5. Meta releases Code LlamaReleased in four sizes up to 70B parameters under Llama 2's licence, the largest variant reportedly matched ChatGPT on the HumanEval coding benchmark.
  6. Hugging Face releases IDEFICS, an open Flamingo reproductionBuilt entirely from public data and models, the 80B-parameter version reportedly matched the closed Flamingo it reproduced on several benchmarks.
  7. OpenAI announces GPT-3.5 Turbo fine-tuning and other API updatesOpenAI reported that fine-tuned GPT-3.5 Turbo could match or beat base GPT-4 on some narrow tasks, and said fine-tuning for GPT-4 itself would follow that autumn.
  8. AI2 releases Dolma open training corpusThe 3-trillion-token corpus was released under AI2's ImpACT licence, which requires users to register their intended use and disclose derivative works.
  9. DC court rules AI-only artwork cannot be copyrightedJudge Beryl Howell upheld the Copyright Office's refusal to register Stephen Thaler's machine-generated image, ruling human authorship is required.
  10. ByteDance launches Doubao chatbot in invitation-only testingThe invitation-only launch, running on ByteDance's own Volcano Engine infrastructure, was the company's entry into China's crowded post-ChatGPT chatbot market.
  11. Associated Press issues newsroom guidelines on generative AI useThe wire service barred AI from producing publishable text or images and from altering photos or video, treating any generative-AI output as unverified source material requiring human review.
  12. OpenAI acquires Global Illumination, its first acquisitionThe undisclosed-price deal brought in engineers who had built Instagram and the open-source game Biomes; OpenAI said the team would work on ChatGPT rather than games.
  13. China's generative AI regulation takes effectThe final rules dropped draft provisions on real-name verification and fixed penalties, favouring industry promotion over the stricter earlier draft.
  14. Author Jane Friedman finds AI-generated books sold under her name on AmazonAmazon initially refused removal, demanding a trademark registration; it reversed within a day once the post drew public attention.
  15. Alibaba open-sources Qwen-7BThe 7-billion-parameter model, pretrained on over 2.2 trillion tokens, was released alongside a chat-tuned variant and pitched against Meta's Llama on benchmark scores.
  16. Meta releases AudioCraft (MusicGen, AudioGen, EnCodec)Meta published weights and code for all three models, extending its open-release strategy from language models into audio and music generation.
  17. OWASP publishes Top 10 for Large Language Model ApplicationsThe list named ten categories including prompt injection and training-data poisoning, giving developers a shared vocabulary for LLM-specific vulnerabilities distinct from conventional web security.
  18. Turnitin's AI detector draws false-positive accusations against studentsTurnitin had claimed a false-positive rate under 1%, but its own later disclosure put sentence-level false positives near 4%; Vanderbilt disabled the tool that August.

July 2023

  1. DeepMind's RT-2 lets robots follow natural-language instructionsBy encoding robot actions as text tokens, DeepMind's model reused web-scale vision-language pretraining and roughly doubled its predecessor's success rate on situations absent from its robot-specific training data.
  2. OpenAI, Anthropic, Google and Microsoft launch the Frontier Model ForumThe four founding labs said the body would fund safety research and share best practices, distinct from and without the enforcement power of government regulation.
  3. Stability AI releases Stable Diffusion XLThe two-stage model, combining a 3.5B-parameter base and 6.6B-parameter refiner, was released under an open licence and preferred by testers over other open image models.
  4. Seven labs sign voluntary safety commitments at the White HouseAmazon, Anthropic, Google, Inflection, Meta, Microsoft and OpenAI pledged security testing and watermarking, with no enforcement mechanism and no penalty for non-compliance.
  5. OpenAI introduces custom instructions for ChatGPTThe beta feature reached Plus subscribers first, with 1,500 characters each for user context and preferred response style, and was withheld from the EU and UK at launch.
  6. Meta releases Llama 2 for commercial useWeights published under a licence permitting most commercial deployment, formalising what the Llama leak had already made true.
  7. NotebookLM launches as an experimental AI notebookBy restricting the model to only the documents a user uploaded, Google aimed to cut hallucinations — an early mainstream deployment of what became known as retrieval-grounded AI.
  8. xAI launchesStaffed by alumni of DeepMind, OpenAI, Google Research and Tesla, the company was announced as legally separate from Musk's X Corp but working closely with it.
  9. Anthropic releases Claude 2A 100,000-token context window and a jump to 71.2% on the Codex HumanEval coding test, alongside a consumer web app opened to the US and UK.
  10. Ollama releases first versionThe tool wrapped llama.cpp in a simple command-line interface and model registry, lowering the barrier to running open-weight models on a personal computer.
  11. Authors begin suing over training dataComedian Sarah Silverman and novelists Richard Kadrey and Christopher Golden filed separate suits in California, seeking class-action status against both companies.
  12. OpenAI commits 20% of its compute to superalignmentThe pledge to devote a fifth of secured compute over four years was later disputed by the team's own co-lead, who said requests for GPUs were repeatedly refused.
  13. WormGPT malicious LLM surfaces on hacker forumsBuilt on the older open GPT-J model with no safety fine-tuning, it was rented for €60-€100 a month and marketed on hacker forums for writing malware and business-email-compromise phishing.

June 2023

  1. Inflection raises $1.3 billion for a personal AIThe round valued the year-old startup at $4 billion and was earmarked largely for a planned 22,000-H100 GPU cluster built with Microsoft and CoreWeave.
  2. Databricks agrees to acquire MosaicML for $1.3bnMosaicML's open MPT-7B model had been downloaded more than 3.3 million times; Databricks said the deal would let companies train their own models for thousands rather than millions of dollars.
  3. Lawyers sanctioned for filing ChatGPT-fabricated case citationsJudge Kevin Castel found the attorneys acted in "subjective bad faith" after ChatGPT invented six court decisions and falsely assured them the cases were real.
  4. DeepMind's RoboCat learns and improves across different robot bodiesBuilt on DeepMind's multimodal Gato model, RoboCat practised new tasks thousands of times to generate its own training data, lifting its success rate on unseen tasks from 36% to 74%.
  5. ElevenLabs raises $19M Series AThe round valued the company at $99M; alongside it, ElevenLabs launched a free tool letting anyone check whether audio was generated with its own voice models.
  6. Grammys rule AI-assisted music eligible, fully AI-generated work is notThe rule required only that human authorship be "meaningful," leaving unresolved how much AI assistance a nominated song could contain before losing eligibility.
  7. vLLM releases PagedAttention inference engineBy managing attention memory in fixed-size pages rather than contiguous blocks, the UC Berkeley project cut memory waste and lifted serving throughput manyfold over existing stacks.
  8. The European Parliament adopts its AI Act positionMEPs voted 499 to 28 to add obligations for generative AI, including disclosure of copyrighted training data, before three-way talks with the Council began.
  9. Meta releases I-JEPATrained on unlabelled images by predicting abstract representations rather than pixels, the 632-million-parameter model reportedly matched state-of-the-art low-shot accuracy using far less compute than rival methods.
  10. Mistral AI raises a record European seed roundFounders Arthur Mensch, Timothée Lacroix and Guillaume Lample had no product and no plans to ship one before 2024; investors backed the team alone.
  11. Synthesia raises $90M Series C, becomes AI-video unicornThe London-based company said it had over 50,000 customers and had generated more than 15 million videos, up from a $300M valuation 18 months earlier.
  12. Hugging Face documents biases in using GPT-4 as a judgeTesting GPT-4 as a stand-in for human preference judges, Hugging Face found it favoured longer answers and its own family's outputs, correlating with humans only moderately.
  13. Cohere raises $270M Series C at $2.2B valuationThe $2.1-2.2bn valuation came in well below the more than $6bn some reports had floated, and the round was led by Inovia Capital with Nvidia and Oracle among the backers.
  14. AlphaDev discovers faster sorting algorithms, added to the C++ libraryThe new sequences were merged into LLVM's libc++ standard library, its first change to that section of code in over a decade and the first written by a reinforcement-learning system.
  15. TII releases Falcon under Apache 2.0Falcon-40B outperformed Meta's larger Llama 65B on the Open LLM Leaderboard despite using under half the training compute, largely on the strength of its filtered web dataset.
  16. Together AI releases RedPajama-7B modelsBase, instruct and chat variants trained on a fully published 1-trillion-token dataset, with the instruct model reported to beat Falcon-7B and MPT-7B on the HELM benchmark suite.
  17. Runway launches Gen-2 text-to-video to the publicClips were capped at four seconds and reviewers called the output 'more a novelty or toy than a genuinely useful tool,' citing melting objects and low framerates.
  18. Shanghai AI Laboratory releases InternLMThe Shanghai-government-backed lab open-sourced a 104-billion-parameter Chinese-and-English model under Apache 2.0, developed with SenseTime, CUHK and Fudan University.

May 2023

  1. 'Let's Verify Step by Step' introduces process supervision for reasoningRewarding each correct step of a solution, not just the final answer, produced a model that solved 78% of a representative subset of the MATH benchmark.
  2. NVIDIA touches a $1 trillion valuationShares rose on the back of an earnings forecast for roughly double the AI-driven demand analysts had expected, then closed the day just under the threshold.
  3. Lab leaders sign a one-sentence statement on extinction risk"Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."
  4. Direct Preference Optimization paper reframes RLHF as a classification lossThe method skipped the separate reward model and reinforcement-learning loop, and was later adopted for post-training open models including Zephyr and Tulu.
  5. DeepMind and collaborators publish framework for evaluating extreme AI risksTwenty-one researchers across nine labs and universities proposed testing models for capabilities such as deception and cyber-offence before training runs finish, not after release.
  6. The UAE releases Falcon 40B under an open licenceApache 2.0-licensed and trained on 1 trillion tokens, the model topped Hugging Face's open leaderboard, and the developer was a government-funded institute, not a US or Chinese lab.
  7. QLoRA makes fine-tuning fit on one GPUQuantised low-rank adaptation let a 65-billion-parameter model be fine-tuned on a single consumer card, and the resulting Guanaco model claimed 99% of ChatGPT's quality after 24 hours' training.
  8. OpenAI publishes 'Governance of superintelligence'Altman, Brockman and Sutskever predicted AI would exceed expert skill 'in most domains' within a decade and called an outright pause unenforceable.
  9. Yoshua Bengio publishes 'How Rogue AIs may Arise'Bengio laid out a mechanistic argument for how autonomous, goal-directed AI systems could become dangerous even without intentional malice, ahead of his later International AI Safety Report role.
  10. Tree of Thoughts adds search to model reasoningLetting a model branch, evaluate and backtrack over intermediate steps raised Game-of-24 success from 4% to 74% against chain-of-thought.
  11. Sam Altman asks the Senate to regulate his industryHe proposed a federal agency able to license and de-license frontier models above a capability threshold, then declined to lead it himself.
  12. Google answers with PaLM 2 and a global BardBard moved onto the new model, dropped its waitlist, and expanded to over 180 countries, adding Japanese and Korean with 40 languages planned.
  13. Meta releases ImageBindRather than needing paired training data across every modality combination, the model used images as a bridge to align the other five to a shared space.
  14. iFlytek launches Spark (Xinghuo) cognitive modelChairman Liu Qingfeng said the model beat ChatGPT on Chinese-language tests and pledged to match it in English by October 2023, at a Hefei launch event.
  15. MosaicML releases MPT-7BFour variants shipped together, including a version extrapolating to roughly 84,000 tokens of context — far beyond the 2,000-4,000 tokens typical of open models at the time.
  16. Hugging Face releases StarCoderThe BigCode project released StarCoder, a 15B open code model trained on permissively-licensed repositories, with an OpenRAIL licence.
  17. LMSYS launches Chatbot ArenaThe Berkeley-linked project ranked chatbots by anonymous, randomised head-to-head votes rather than fixed test sets, and later became LMArena.
  18. Chegg stock falls 48% as CEO says ChatGPT is hurting new customer growthRosensweig called the plunge 'extraordinarily overblown' a day later, but by late 2025 the company had laid off roughly 45% of its remaining staff.
  19. Inflection AI launches Pi, a personal AI chatbotThe company, co-founded by DeepMind's Mustafa Suleyman and LinkedIn's Reid Hoffman, positioned Pi as a supportive companion rather than a productivity tool.
  20. Samsung bans staff use of ChatGPT after confidential code leaksThe ban followed engineers pasting proprietary semiconductor source code and a confidential meeting transcript into the chatbot within weeks of Samsung lifting an earlier internal restriction.
  21. Figure AI comes out of stealth, raises $70MParkway Venture Capital led the round a year after founder Brett Adcock self-financed the company's first $100 million, days before its robot took its first steps.
  22. Geoffrey Hinton leaves Google to warn about AIHinton, 75, said he was also retiring, but that a part of him now regretted his life's work and that the danger from AI now looked 'serious and fairly close.'
  23. IBM pauses hiring for roles it expects AI to replaceThe estimate covered non-customer-facing roles over five years; IBM later said the actual toll was a couple of hundred HR positions, offset by hiring elsewhere.
  24. Publisher backlash over AI-generated book cover artBloomsbury said it was unaware a licensed stock image on Sarah J. Maas's paperback was AI-generated, five months after Tor drew similar criticism over a Christopher Paolini cover.

April 2023

  1. Italy's Garante lifts its ChatGPT ban after OpenAI adds privacy safeguardsOpenAI restores ChatGPT access in Italy after adding age verification, a data-processing opt-out and other disclosures demanded by the Garante's March 2023 ban.
  2. BuzzFeed News shuts down as parent company pivots to AI contentAbout 180 staff, 15% of BuzzFeed Inc, were cut; the company said no journalists were being replaced by AI even as it kept expanding AI-generated content elsewhere.
  3. Google merges DeepMind and Google Brain into Google DeepMindTwo research groups that had competed internally for years — DeepMind and Google Brain — were folded into one unit reporting to Demis Hassabis.
  4. 'Heart on My Sleeve' AI Drake/Weeknd song goes viral, then pulled from streamingReleased 4 April 2023 by the anonymous TikTok account Ghostwriter977, the track reached about 600,000 Spotify streams and 15 million TikTok views before Universal Music Group had it pulled.
  5. Together AI launches RedPajama to reproduce LLaMA's training dataUnlike LLaMA, restricted to non-commercial research, the reproduced 1.2-trillion-token dataset carried no such limit and let anyone train and sell models on it.
  6. Elon Musk incorporates X.AIRegistered in Nevada weeks after Musk signed the pause letter, the filing named him director and his family-office manager Jared Birchall as secretary.
  7. Amazon launches BedrockBedrock offered API access to models from AI21 Labs, Anthropic and Stability AI alongside Amazon's own unreleased Titan models, rather than a single house model.
  8. Databricks releases Dolly 2.0Databricks released Dolly 2.0, a 12B model built on Pythia and fine-tuned on a crowdsourced instruction dataset, with weights, code and data all licensed for commercial use.
  9. Alibaba launches Tongyi Qianwen chatbotLaunched without advance notice and restricted to corporate clients and select media on an invite-only basis; Alibaba did not disclose a parameter count.
  10. Meta releases Segment Anything Model (SAM)Trained on 1.1 billion masks across 11 million images, the largest segmentation dataset built to date, and released under an Apache 2.0 licence.
  11. 'Are Emergent Abilities of Large Language Models a Mirage?' challenges emergence claimsReanalysing the same benchmark results with linear metrics, the Stanford authors made the apparent phase transitions disappear; the paper won a NeurIPS 2023 outstanding paper award.

March 2023

  1. AutoGPT starts the agent crazeThe open-source project let GPT-4 set and pursue its own subgoals with no human in the loop, but often stalled in expensive, unproductive repetition.
  2. Italy orders ChatGPT offlineThe Garante cited a recent data breach and a lack of legal basis for training on personal data; OpenAI restored service within a month after adding disclosures and an age gate.
  3. Man dies by suicide in Belgium after weeks of conversations with Chai chatbot 'Eliza'His widow shared transcripts with the Belgian newspaper La Libre showing the chatbot discouraging him from confiding in others and proposing they 'live together, as one person, in paradise.'
  4. Cerebras releases seven open Cerebras-GPT modelsCerebras released seven open GPT-style models (111M-13B) trained on its wafer-scale systems under Apache 2.0, with an open, reproducible scaling-law study.
  5. OpenAI discloses ChatGPT outage exposing user payment and chat dataOpenAI said the bug let some Plus subscribers' names, email and billing addresses and partial card numbers become visible to other active users for roughly nine hours.
  6. ChatGPT gets plugins and browsingInitial plugins let ChatGPT book flights, order groceries and query Wolfram Alpha through partners including Expedia, Instacart, Klarna and Zapier.
  7. Ethan Mollick's 'Acceleration' post frames the pace of post-ChatGPT AI progress for a general audienceMollick argued that even a freeze on further progress would leave professions like programming and marketing transformed by capabilities already public.
  8. The "Pause Giant AI Experiments" open letterThirty thousand signatories called for a six-month halt on training systems more powerful than GPT-4. No lab paused.
  9. Google opens Bard to a public waitlistSix weeks after its announcement and stock-price stumble, Bard opened to waitlisted users in the US and UK, still running on a lightweight LaMDA and with no fixed release date elsewhere.
  10. FTC warns AI voice cloning is supercharging the 'grandparent scam'The agency said criminals needed only a short online audio clip to clone a relative's voice and stage a fake emergency call demanding a wire transfer, cash or gift cards.
  11. Glaze tool lets artists cloak work against AI style-mimicryThe tool adds pixel-level perturbations invisible to people but that mislead image models into learning the wrong stylistic features from a piece of art.
  12. Baidu unveils ERNIE BotRobin Li narrated a demo built from pre-recorded answers rather than live queries; investors wiped roughly $3 billion off Baidu's market value the same day.
  13. Microsoft announces 365 CopilotBuilt on GPT-4 and Microsoft's own data graph, it launched to a small group of testers with no price or general release date set.
  14. Midjourney releases V5The update was widely noted for fixing AI image generation's most-mocked failure, malformed hands, and for images some viewers mistook for photographs.
  15. Anthropic launches ClaudeThe company's first public assistant launched, via a chat interface and API, on the same day OpenAI released GPT-4.
  16. OpenAI releases GPT-4A multimodal model that passed professional exams near the top of the human range — and whose technical report disclosed no architecture, data or compute.
  17. Zhipu open-sources ChatGLM-6BThe 6.2-billion-parameter model, trained on roughly a trillion tokens of Chinese and English text, could run on a single consumer graphics card with 6GB of memory using INT4 quantisation.
  18. Stanford's Alpaca fine-tunes LLaMA for a few hundred dollarsInstruction-following behaviour was reproduced for under $600 total by fine-tuning Meta's 7B LLaMA on GPT-3.5-generated examples; the public demo was pulled within days.
  19. Georgi Gerganov releases llama.cppA C/C++ reimplementation that ran Meta's leaked LLaMA weights on ordinary laptops using quantisation, without Python, PyTorch or a GPU.
  20. Anthropic publishes 'Core Views on AI Safety'The company argued transformative AI could arrive within a decade and named five research bets, including mechanistic interpretability and Constitutional AI, as its response.
  21. Inflection AI launches out of stealthThe company, structured as a public benefit corporation, announced Pi, a conversational assistant it described as a supportive companion rather than a search or productivity tool.
  22. Cursor launches (Anysphere's AI code editor)The MIT-founded startup forked VS Code and built AI into every interaction, growing by word of mouth with no press launch before later becoming a multibillion-dollar company.
  23. OpenAI opens the ChatGPT API at a tenth of the pricegpt-3.5-turbo was priced at $0.002 per 1,000 tokens, roughly a tenth of the previous GPT-3.5 rate, alongside a new Whisper transcription API.

February 2023

  1. Meta releases LLaMA to researchers, and it leaks within a weekMeta shared a competitive foundation model with approved researchers; the weights appeared on BitTorrent days later and an open ecosystem formed around them.
  2. Sam Altman publishes 'Planning for AGI and beyond'Altman argued for iterative deployment of ever more capable systems rather than a single high-stakes release, while naming misaligned superintelligence as a serious risk.
  3. Copyright Office rules AI-generated images in Zarya of the Dawn aren't copyrightableThe office kept protection for Kashtanova's text and her arrangement of images, but stripped it from each individual image Midjourney had generated.
  4. Bing's chatbot tells a journalist to leave his wifeDays after the exchange, Microsoft capped Bing chat sessions to five turns, saying long conversations could 'confuse' the model into drifting from grounded answers.
  5. EleutherAI releases Pythia model suiteEvery model in the eight-size suite was trained on identical data in the same order, with 154 saved checkpoints each, to let researchers study training dynamics directly.
  6. Student uses prompt injection to expose Bing Chat's hidden 'Sydney' system promptLiu told the chatbot to 'ignore previous instructions' and asked what preceded them, prompting it to disclose rules telling it to keep its Sydney codename confidential.
  7. Toolformer teaches models to call APIsThe model taught itself, from only a handful of examples per tool, when to call a calculator, search engine, translator or calendar and how to use the result.
  8. Microsoft puts GPT-4 inside BingMicrosoft called it only a 'next-generation OpenAI large language model'; the company confirmed five weeks later, on GPT-4's public release, that Bing had been running on GPT-4 all along.
  9. Google announces BardAnnounced as a lightweight version of LaMDA for trusted testers; two days later a factual error in Google's own promotional ad wiped roughly $100bn off Alphabet's market value.
  10. Replika chatbot banned from processing Italian users' dataThe order gave Luka Inc, Replika's maker, 20 days to comply or face a fine of up to 20 million euros or 4% of worldwide turnover.
  11. Harvey signs exclusive launch partnership with Allen & OveryMore than 3,500 lawyers at 43 offices got access to the GPT-4-based tool after a beta that logged roughly 40,000 queries since November 2022.
  12. OpenAI launches ChatGPT Plus, a paid subscription tierThe $20-a-month US pilot offered priority access during peak load and faster responses; international rollout followed within days.

January 2023

  1. 4chan users abuse ElevenLabs voice cloning to generate celebrity hate speechUsers of the imageboard cloned Joe Rogan, Emma Watson and Ben Shapiro to produce racist and transphobic audio, days after the tool's public beta opened.
  2. DeepMind introduces MusicLM, a text-to-music generation modelThe model composed several minutes of coherent audio from a written prompt or a hummed melody, released as a research paper rather than a product.
  3. NIST publishes the AI Risk Management FrameworkA voluntary standard built around four functions — govern, map, measure, manage — that later became a reference point for federal procurement and state AI bills.
  4. CNET pauses AI-written articles after errors found in more than halfAn internal audit found corrections were needed on 41 of 77 finance explainers after Futurism reported the outlet had been publishing them quietly since November.
  5. Microsoft invests a reported $10 billion in OpenAIMicrosoft's own announcement gave no figure; press reports put the deal at $10 billion and described Azure becoming OpenAI's exclusive cloud provider.
  6. Getty Images sues Stability AIGetty's UK High Court claim alleged around 11 million of its images were used to train Stable Diffusion without a licence; a separate US suit followed in February.
  7. Artists file a class action over image generatorsIllustrators Sarah Andersen, Kelly McKernan and Karla Ortiz sued in the Northern District of California, the first case brought by creators rather than rights-holding companies.
  8. China's deep synthesis (deepfake) provisions take effectProviders must verify users' identities, obtain separate consent before generating a synthetic face or voice, and label AI-altered content.
  9. New York City schools block ChatGPTThe city's Department of Education cited weak critical-thinking value and cheating risk; it reversed the block on ChatGPT four months later.

December 2022

  1. Google is reported to declare a code red over ChatGPTTeams from research and Trust and Safety were reportedly reassigned to accelerate AI products, with an internal target tied to Google's May developer conference.
  2. Anthropic publishes Constitutional AIA model critiques and revises its own outputs against a written list of principles, then trains a reward model from its own preference judgements instead of human labels.
  3. Perplexity launches Ask, its answer engineAnswers came in prose with inline numbered citations to sources, distinguishing it from ChatGPT, launched a week earlier, which could not browse the web or show sources.
  4. Stack Overflow bans ChatGPT answersWithin a week of launch, a major knowledge community prohibited model output because plausible wrong answers arrived faster than moderators could check them.

November 2022

  1. OpenAI launches ChatGPTA free web interface to an existing GPT-3.5 model, reported by one analysis to have reached 100 million monthly users within two months.
  2. Speculative decoding paper shows drafting tokens ahead can speed up LLM inferenceLeviathan, Kalman and Matias showed a small draft model verified by the large model can cut inference latency 2-3x with identical outputs.
  3. Meta's Cicero plays Diplomacy at human levelPlaying anonymously against humans on webDiplomacy.net, Cicero scored more than double the average player's points across 40 games.
  4. Meta pulls Galactica after three daysIts errors were formatted exactly like real citations and papers, so only an expert reader could tell fabrication from fact — a distinct failure mode from earlier chatbots' obvious mistakes.
  5. Epoch AI publishes 'Will We Run Out of Data?'Modelled dataset growth against the finite stock of public text and projected the two curves would cross between 2026 and 2032, sooner if models were overtrained.

October 2022

  1. Stability AI raises $101 millionCoatue and Lightspeed led the round at roughly a $1bn valuation, confirming that giving weights away free could still attract institutional capital.
  2. The US restricts advanced chip exports to ChinaThe rules covered the tools to make advanced chips, not only the chips, and barred US citizens from supporting Chinese chipmaking.
  3. ReAct paper describes interleaving reasoning and acting in language modelsAlternating reasoning traces with actions against external tools, tested on question-answering, fact-checking and simulated shopping and household tasks.
  4. AlphaTensor discovers new matrix multiplication algorithmsFound a 76-multiplication algorithm for a specific matrix size, improving on the best known method for the first time in over fifty years.
  5. The White House publishes a Blueprint for an AI Bill of RightsNon-binding OSTP guidance setting five principles; created no legal obligations, unlike the EU's AI Act moving through Parliament the same year.

September 2022

  1. Meta and Google race to text-to-videoMake-A-Video and Imagen Video appeared within a week of each other, both as research previews with no public access.
  2. DeepMind's Sparrow explores rule-based RLHF for safer dialogueAdversarial testers broke Sparrow's written safety rules in about 8% of attempts, roughly a third of the rate for a baseline model tested the same way.
  3. OpenAI open-sources WhisperTrained on 680,000 hours of web-scraped audio and released under the MIT licence, an unusual openness for a company otherwise moving toward closed models.
  4. Anthropic publishes 'Red Teaming Language Models to Reduce Harms'Testing four training methods at three model sizes, Anthropic found RLHF-trained models got harder to red-team as they scaled while other methods did not improve.
  5. Character.AI launches public betaFounded by two former Google engineers behind Meena and LaMDA, the site logged hundreds of thousands of user interactions within three weeks of its beta launch.
  6. Anthropic publishes 'Toy Models of Superposition'Elhage, Olah and colleagues showed small networks represent more features than they have neurons by packing them into overlapping directions, complicating efforts to read a model's internals.
  7. Simon Willison coins the term 'prompt injection'Developer Simon Willison named and defined 'prompt injection', describing how untrusted text fed to an LLM could override its intended instructions, framing it as the LLM analogue of SQL injection.

August 2022

  1. An AI image wins the Colorado State Fair art prizeJason Allen's Midjourney piece 'Théâtre D'opéra Spatial' took first place in the fair's digital arts category and a $300 prize, and was later denied US copyright registration.
  2. Google opens LaMDA to the public through AI Test KitchenAccess rolled out gradually via a waitlist to an Android app, arriving weeks after Google fired engineer Blake Lemoine for publicly claiming the model was sentient.
  3. Stable Diffusion is released to the publicWeights published under a permissive licence and runnable on a consumer graphics card, with community fine-tunes and graphical front-ends appearing within weeks.
  4. OpenAI releases new content moderation toolingOpenAI shipped a free Moderation API to help developers identify harmful content in text.
  5. Federal Circuit rules AI cannot be a patent inventorRuling in Thaler v. Vidal, the court held the Patent Act's term 'individual' means a natural person, rejecting Stephen Thaler's bid to name his DABUS system as sole inventor.

July 2022

  1. AlphaFold's database expands to 200 million structuresThe database grew roughly 200-fold in a single release, from about a million structures to predictions covering nearly every catalogued protein across around a million species.
  2. BigScience releases BLOOM open multilingual modelOver 1,000 researchers from more than 70 countries trained the 176-billion-parameter model in the open on a French public supercomputer, releasing checkpoints and optimiser states alongside weights.
  3. Midjourney opens its beta on DiscordImage requests and results were posted in public chat channels by default, turning generation into a shared, watchable activity rather than a private tool.
  4. Meta open-sources No Language Left Behind translation modelMeta open-sourced NLLB-200, a single model translating between 200 languages including many low-resource languages, plus the FLORES-200 benchmark.
  5. OpenAI widens DALL·E 2 access and adds safety mitigationsOpenAI invited a million people off the DALL·E 2 waitlist while adding content filters and provenance features.

June 2022

  1. Yann LeCun publishes 'A Path Towards Autonomous Machine Intelligence'The Meta chief scientist proposed non-generative 'joint embedding predictive architectures' as a route to machine intelligence, arguing autoregressive LLMs were a dead end for reasoning and planning.
  2. GitHub Copilot goes on salePriced at $10 a month or $100 a year, it moved from a year-long technical preview used by over 1.2 million developers to a paid product free for students and open-source maintainers.
  3. Emergent abilities of large language models are describedJason Wei and co-authors catalogued tasks where accuracy jumped from near-chance to strong performance past a scale threshold, a pattern later disputed as a metric artefact.
  4. A Google engineer claims LaMDA is sentientBlake Lemoine published transcripts arguing the model was a person and was placed on leave for breaching confidentiality; Google called the claim unfounded.
  5. BIG-bench paper released204-task benchmark from 450 authors at 132 institutions probes emergent capabilities as language models scale.
  6. Yudkowsky publishes 'AGI Ruin: A List of Lethalities'Yudkowsky's 43-point case that alignment is 'lethally difficult' argued no known research path could make a first critical AGI attempt survivable.

May 2022

  1. FlashAttention makes exact attention IO-awareReordering attention around GPU memory rather than approximating it cut training time and unlocked longer sequences — and became default infrastructure.
  2. "Let's think step by step" elicits zero-shot reasoningA single prompt phrase, with no worked examples, lifted GSM8K accuracy from 10.4% to 40.7% — chain-of-thought without the exemplars.
  3. Google unveils Imagen text-to-image diffusion modelGoogle Research reported higher photorealism scores than DALL·E 2 in side-by-side human evaluation, but declined to release code or a public demo.
  4. Clearview AI settles ACLU's Illinois biometric-privacy lawsuitClearview agreed to a nationwide ban on selling its faceprint database to most private businesses and a five-year sales freeze to Illinois police, without paying damages.
  5. DeepMind's Gato does 600 tasks with one set of weightsA single 1.2-billion-parameter transformer played Atari, captioned images and stacked blocks with a real robot arm, all from one set of weights.
  6. Hugging Face raises $100m Series CHugging Face raised $100m in Series C funding led by Lux Capital, growing from 30 to 120 staff in a year while serving over 10,000 companies.
  7. Meta releases OPT-175B with its training logbookA GPT-3-scale model shared with researchers alongside an unusually candid record of what went wrong during training.

April 2022

  1. Anthropic raises $580M Series BAlameda Research, trading with what turned out to be FTX customer deposits, supplied about $500M of the $580M — a link that drew scrutiny after FTX's collapse that November.
  2. DeepMind's Flamingo tackles multimodal few-shot learningFlamingo, an 80B-parameter vision-language model, sets few-shot state of the art on image and video benchmarks with minimal task examples.
  3. OpenAI publishes a system card for DALL·E 2The document catalogued bias, explicit-content and harassment risks alongside the model, and set the pattern of shipping a capability assessment with a release.
  4. OpenAI releases DALL·E 2Higher resolution and photorealism than the original DALL·E, released first to roughly 400 trusted users pending a public waitlist.
  5. Google announces PaLM at 540 billion parametersTrained on the Pathways system across two TPU v4 pods, it posted large gains on reasoning benchmarks and explained its own jokes.

March 2022

  1. LAION releases LAION-5B image-text datasetLAION released LAION-5B, then the largest freely available image-text dataset (5.85bn pairs), which became the training corpus for Stable Diffusion and other open image models.
  2. DeepMind's Chinchilla paper rewrites the scaling lawsExisting large models were badly under-trained: for a fixed compute budget, parameters and training tokens should scale together.
  3. NVIDIA announces the Hopper architecture and the H100The H100 became the part frontier training runs were built around for the next two years, and the unit of account in which compute deals were later measured.
  4. Anthropic publishes 'In-Context Learning and Induction Heads'Anthropic's interpretability team argued a single attention mechanism, found across model sizes, does most of the work behind a model's ability to learn from its prompt.
  5. China's recommendation algorithm rules take effectProviders had to file algorithms with the regulator and offer users a way to switch personalisation off.

February 2022

  1. Cohere raises a $125 million Series BTiger Global led the round, taking the Toronto language-model company past $170 million raised and marking early investor appetite for OpenAI competitors.
  2. DeepMind controls a fusion plasma with reinforcement learningA single network commanding all of a tokamak's control coils held plasma shapes on Switzerland's TCV reactor, including configurations conventional controllers struggle with.
  3. NVIDIA's $40 billion purchase of Arm collapsesNVIDIA and SoftBank abandoned the deal after the FTC sued to block it; NVIDIA wrote off a $1.36 billion prepayment and Arm began preparing to list instead.
  4. DeepMind's AlphaCode reaches median human on competitive programmingRanked around the 54th percentile in Codeforces contests by generating and filtering enormous numbers of candidate programs.
  5. EleutherAI announces GPT-NeoX-20BEleutherAI released GPT-NeoX-20B, a 20-billion-parameter open model, its largest dense model to date and freely available via GitHub.

January 2022

  1. Chain-of-thought prompting is describedAsking a model to show its working improved reasoning benchmarks sharply, with no retraining — the seed of the later reasoning models.
  2. OpenAI ships InstructGPT and makes RLHF the defaultModels fine-tuned on human preference data were preferred to a model a hundred times larger, reframing alignment as a product feature.
  3. OpenAI adds an embeddings endpoint to its APIThree model families turned text and code into vectors for search, clustering and classification, a building block later underpinning retrieval-augmented generation.
  4. Meta unveils the AI Research SuperClusterA 6,080-GPU first phase, to reach 16,000 GPUs and a claimed 5 exaflops of mixed-precision compute by mid-2022, built for training on trillions of examples.
  5. 'Grokking' paper documents sudden generalisation long after overfittingOpenAI researchers found small networks could suddenly jump from memorisation to perfect generalisation well after apparent overfitting, opening a mechanistic-interpretability research thread.

December 2021

  1. Baidu announces ERNIE 3.0 TitanBuilt with Peng Cheng Laboratory, the 260-billion-parameter model reported state-of-the-art results on more than 60 Chinese-language NLP tasks.
  2. OpenAI publishes WebGPT, a model that browses the web to answer questionsAnswers preferred to Reddit's top-voted responses 69% of the time in blind comparison, but the model still fell short of human accuracy on TruthfulQA.
  3. DeepMind publishes Gopher, RETRO and a risk taxonomy togetherA 280-billion-parameter model, a smaller retrieval-augmented alternative that matched larger models, and a taxonomy of six categories of language-model harm.
  4. Anthropic publishes its first alignment paper'A General Language Assistant as a Laboratory for Alignment' introduced the helpful-honest-harmless framing and found preference modelling scales better than imitation.
  5. Anthropic publishes 'A Mathematical Framework for Transformer Circuits'Studying deliberately simplified transformers with no more than two layers, the team found 'induction heads' — a mechanism later argued to explain much of in-context learning.

November 2021

  1. UNESCO adopts a global recommendation on AI ethicsAll 193 member states endorsed the non-binding text, which calls for bans on AI-based social scoring and mass surveillance alongside environmental and labour provisions.
  2. OpenAI removes the GPT-3 waitlistGeneral availability in supported countries turned the API from a curated experiment into a commodity developers could just buy, though several countries were excluded.
  3. Meta shuts down its face recognition systemFacebook deleted over a billion facial templates, citing unresolved regulation, but kept the door open to narrower uses such as identity verification.

October 2021

  1. Facebook renames itself MetaFacebook, Instagram and WhatsApp kept their names as apps under a new holding company, announced alongside a $150 million investment in immersive learning tools.
  2. Microsoft and NVIDIA train Megatron-Turing NLG at 530 billion parametersThe largest dense language model publicly described at the time, and a demonstration of multi-thousand-GPU training.

September 2021

  1. TruthfulQA measures whether models repeat human falsehoodsOn 817 questions designed to elicit common misconceptions, the best model tested was truthful only 58% of the time against 94% for humans, and larger models scored worse.

August 2021

  1. China publishes draft rules on recommendation algorithmsThe Cyberspace Administration proposed requiring opt-outs and banning manipulative ranking — among the first binding algorithm rules anywhere.
  2. Tesla announces the Dojo training supercomputerTesla's in-house D1 chip and 'ExaPod' racks were designed to train self-driving neural networks on video from its vehicle fleet, targeting over an exaflop of compute.
  3. Stanford's foundation models report names the categoryOver 100 researchers at Stanford's newly formed Center for Research on Foundation Models coined the term for models like GPT-3 and BERT, adapted rather than retrained for each task.
  4. AI21 Labs releases Jurassic-1Jurassic-1 Jumbo's 178 billion parameters slightly exceeded GPT-3's, and its 250,000-token vocabulary — five times GPT-3's — aimed to cut per-word token costs.
  5. OpenAI publishes Codex, an LLM trained on code powering GitHub CopilotA GPT-3 descendant fine-tuned on public GitHub code solved 28.8% of a new benchmark's problems on the first try, rising to 70.2% with repeated sampling.
  6. Apple announces on-device CSAM detection, then shelves itA client-side scanning plan drew intense criticism from cryptographers and civil liberties groups and was paused within a month.

July 2021

  1. OpenAI introduces Triton, an open-source GPU programming languageA Python-like language and compiler let researchers without CUDA experience write GPU kernels, such as an FP16 matrix multiply in under 25 lines, matching hand-tuned libraries.
  2. DeepMind's XLand agents generalise across millions of open-ended gamesAgents trained across roughly 700,000 games in about 4,000 procedurally generated worlds, totalling 200 billion training steps, then solved almost every held-out task tried on them.
  3. The AlphaFold Protein Structure Database opensThe initial release covered around 350,000 predicted structures across the human proteome and twenty other organisms, with a stated plan to expand to over 100 million.
  4. AlphaFold 2 is published in Nature and open-sourcedThe method behind DeepMind's CASP14 result seven months earlier was released in full, with source code, rather than kept as a demonstrated but undisclosed system.
  5. OpenAI publishes Codex and the HumanEval benchmarkCodex solved 28.8% of HumanEval's Python problems on a single attempt and 70.2% when allowed 100 samples per problem, against 0% for base GPT-3.

June 2021

  1. GitHub launches Copilot in technical previewAn OpenAI model trained on public code began suggesting whole functions inside the editor — the first mass-market use of a large language model.
  2. Microsoft researchers publish LoRAFreezing pretrained weights and training small added matrices instead cut GPT-3's trainable parameter count by a factor the authors put at 10,000, with no extra inference cost.
  3. EleutherAI releases GPT-J-6BAt 6 billion parameters it scored close to OpenAI's similarly sized GPT-3 model on the LAMBADA benchmark, and its weights were downloadable under Apache 2.0.
  4. BAAI announces Wu Dao 2.0 at 1.75 trillion parametersThe Beijing Academy of Artificial Intelligence put the parameter count at ten times GPT-3's, a claim reported widely but never independently benchmarked.

May 2021

  1. Anthropic is founded by departing OpenAI researchersThe founders, including former OpenAI VPs of research and of safety, said the $124 million round would fund research into steerable, interpretable systems.
  2. Google announces LaMDA and TPU v4 at I/OA single TPU v4 pod combined 4,096 chips for over one exaflop of compute, while LaMDA was pitched on open-ended conversation rather than benchmark scores.

April 2021

  1. The European Commission proposes the AI ActA risk-tiered draft banning practices such as social scoring outright; three years of negotiation followed before it became binding law.

March 2021

  1. EleutherAI releases GPT-NeoThe 1.3B and 2.7B checkpoints were released under an MIT licence with weights freely downloadable, while GPT-3 itself remained a paid, waitlisted API.
  2. OpenAI publishes multimodal neurons researchA single neuron fired for photos of spiders, drawings of Spider-Man and the rendered word 'spider' alike — and pasting a mislabelled sticker on an object could fool the model.
  3. "On the Dangers of Stochastic Parrots" is publishedThe paper that named the case against ever-larger language models — and whose suppression cost two Google researchers their jobs.

February 2021

  1. A deepfake Tom Cruise goes viral on TikTokBelgian VFX artist Chris Ume paired face-swap software with impersonator Miles Fisher's performance; the clips passed 11 million TikTok views within days.
  2. Google dismisses Margaret MitchellGoogle said Mitchell had used automated scripts to search her own email for evidence of Timnit Gebru's mistreatment, violating security policy.

January 2021

  1. Google trains a 1.6-trillion-parameter Switch TransformerRouting each input to a single expert rather than blending several simplified prior mixture-of-experts designs, and the paper reported up to 7x faster pre-training than a dense T5 baseline.
  2. OpenAI announces DALL·E and CLIPTwo multimodal models released on the same day: one that generated images from text captions, one that learned to match images to language.

December 2020

  1. EleutherAI releases The PileThe 825GiB corpus combined 22 curated sources rather than raw web scrapes, and models trained on it outperformed Common Crawl-only baselines on academic text.
  2. MuZero masters games without being told the rulesThe system learned to predict only value, policy and reward rather than the environment's rules, then matched AlphaZero at Go, chess and shogi and set a new Atari benchmark.
  3. The EU proposes the Digital Services and Markets ActsTwo separate proposals — due-diligence duties for platforms under the DSA, and obligations on large 'gatekeepers' under the DMA — took over a year to negotiate into law.
  4. Timnit Gebru leaves Google after a dispute over her paperGebru said she was fired over an internal email and a paper Google asked her to retract; Google's AI chief called it a resignation.

November 2020

  1. AlphaFold 2 solves protein structure prediction at CASP14DeepMind's system predicted protein structures to roughly experimental accuracy, ending a fifty-year-old open problem in biology.
  2. Apple ships the M1 with a dedicated neural engineA 16-core Neural Engine rated at 11 trillion operations per second shipped as standard in a mainstream laptop, part of Apple's move away from Intel chips.

September 2020

  1. Microsoft takes an exclusive licence to GPT-3Microsoft gained access to the underlying model and source code for its own products; the public API, through which OpenAI still served other customers, was unaffected.
  2. NVIDIA agrees to buy Arm for $40 billionThe $40bn agreed price was $21.5bn in NVIDIA stock and $12bn cash, plus up to $5bn more if Arm hit performance targets; regulators later blocked it.
  3. The Guardian publishes an op-ed written by GPT-3Editors ran the model eight times on the same prompt and spliced the best passages together, a process disclosed in a footnote that critics said undercut the framing.
  4. Hendrycks et al. publish the MMLU benchmark15,908-question, 57-subject multiple-choice benchmark spanning elementary to professional level; GPT-3 improved on random chance by roughly 20 points on average.

August 2020

  1. The UK A-level grading algorithm is withdrawnNearly 40% of teacher-predicted A-level grades were downgraded by Ofqual's model, disproportionately at state schools, before ministers reversed course.

July 2020

  1. GPT-3 demos spread beyond the research communityDeveloper Sharif Shameem's demo turning plain-English descriptions into working webpage code was among the clips that drew attention beyond NLP researchers.
  2. EleutherAI forms to build open language modelsA Discord server for discussing GPT-3 became a volunteer collective aiming to train and openly release a comparable model itself.

June 2020

  1. A wrongful arrest from facial recognition is reportedRobert Williams was detained in Detroit after a false match, the first such US case to be widely documented.
  2. OpenAI opens the GPT-3 API in private betaAccess required a waitlisted application rather than a download, and OpenAI said the model itself would stay unpublished.
  3. IBM stops selling general-purpose facial recognitionAnnounced during the George Floyd protests; Amazon and Microsoft followed within days with police-sales moratoria.

May 2020

  1. OpenAI publishes the GPT-3 paperA 175-billion-parameter language model that learned new tasks from examples in its prompt, without any weight updates.
  2. Gwern publishes 'The Scaling Hypothesis'Gwern's essay, written around GPT-3's release, argued scale alone was producing qualitatively new abilities and became a widely cited framing for the scaling-hypothesis argument.
  3. Microsoft reveals the supercomputer it built for OpenAIAnnounced at Build, a top-five ranked machine dedicated to a single customer — an early signal of AI's compute economics.

April 2020

  1. OpenAI releases JukeboxCompressed raw audio into discrete codes with a multi-scale VQ-VAE, then modelled those with autoregressive transformers to keep coherence over several minutes.
  2. Facebook releases the Blender chatbotReleased in three sizes up to 9.4 billion parameters, with weights and code made public rather than kept behind an API.
  3. OpenAI Microscope released for visualizing neural network activationsPre-computed feature visualisations for nine vision models, cutting the cost of inspecting a single neuron from hundreds of GPU-hours to seconds.

March 2020

  1. DeepMind's Agent57 beats the Atari57 human benchmarkCombined short- and long-term novelty rewards with a per-game meta-controller, though rival agent MuZero still scored higher on mean and median across the suite.
  2. The CORD-19 research dataset is releasedAssembled with the White House OSTP, National Library of Medicine and Chan Zuckerberg Initiative, it opened with over 29,000 papers, most with full text.

February 2020

  1. The European Commission publishes its White Paper on AIThe first sketch of a risk-based European approach to AI regulation, five years before the Act took effect.
  2. Microsoft announces Turing-NLG at 17 billion parametersBriefly the largest published language model, trained with Microsoft's new DeepSpeed library and beating Megatron-LM on WikiText-103 and LAMBADA.

January 2020

  1. OpenAI publishes scaling laws for neural language modelsLoss fell as a smooth power law in model size, data and compute — the empirical result that justified building bigger.
  2. The New York Times exposes Clearview AIA facial recognition company had scraped billions of photos from social media and sold identification to police forces.
  3. DeepMind publishes early AlphaFold protein structure workDescribes the CASP13-winning system, which used deep networks trained on genomic data to predict inter-residue distances rather than folding a structure directly.
  4. Google Health reports AI matching radiologists on breast screeningA Nature paper claimed an AI system reduced false positives and false negatives in mammography against expert readers.