<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>AI So Far</title>
    <link>https://itdoeswhatnow.com/</link>
    <atom:link href="https://itdoeswhatnow.com/rss.xml" rel="self" type="application/rss+xml" />
    <description>Neutral briefs on the development of artificial intelligence, January 2020 to the present.</description>
    <language>en</language>
    <lastBuildDate>Thu, 13 Aug 2026 16:29:37 GMT</lastBuildDate>
    <item>
      <title>Thinking Machines Lab sets out staged release framework for open weights</title>
      <link>https://itdoeswhatnow.com/m/2026-08-06-thinking-machines-lab-sets-out-staged-release-framework/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-06-thinking-machines-lab-sets-out-staged-release-framework/</guid>
      <pubDate>Thu, 06 Aug 2026 00:00:00 GMT</pubDate>
      <description>Thinking Machines proposed a staged release process for open-weight models -- inference access, then fine-tuning APIs, then full weights -- to manage dangerous-capability risk.</description>
      <category>Open weights &amp; ecosystem</category>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>Demis Hassabis steps down as Google DeepMind CEO in leadership reshuffle</title>
      <link>https://itdoeswhatnow.com/m/2026-08-05-demis-hassabis-steps-down-as-google-deepmind-ceo/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-05-demis-hassabis-steps-down-as-google-deepmind-ceo/</guid>
      <pubDate>Wed, 05 Aug 2026 00:00:00 GMT</pubDate>
      <description>Google's stock fell nearly 4% on the news, which coincided with Gemini's flagship model having slipped past its planned mid-2026 launch.</description>
      <category>Labs &amp; people</category>
    </item>
    <item>
      <title>DeepMind's WeatherNext model improves cyclone forecasting accuracy</title>
      <link>https://itdoeswhatnow.com/m/2026-08-04-deepminds-weathernext-model-improves-cyclone-forecasting-accuracy/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-04-deepminds-weathernext-model-improves-cyclone-forecasting-accuracy/</guid>
      <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
      <description>WeatherNext gives roughly an extra day of predictive accuracy on cyclone track, intensity and wind structure versus prior forecasting models, per a Nature paper.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>OpenAI publicly responds to Apple's trade-secret lawsuit</title>
      <link>https://itdoeswhatnow.com/m/2026-08-04-openai-publicly-responds-to-apples-trade-secret-lawsuit/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-04-openai-publicly-responds-to-apples-trade-secret-lawsuit/</guid>
      <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
      <description>OpenAI called Apple's suit 'careless, aggressive and oddly personal,' published emails and iMessages disputing its account, and asked a court to dismiss the case.</description>
      <category>Courts &amp; copyright</category>
    </item>
    <item>
      <title>Tino Cuéllar joins Anthropic as first Chief Global Affairs Officer</title>
      <link>https://itdoeswhatnow.com/m/2026-08-04-tino-cu-llar-joins-anthropic-as-first-chief/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-04-tino-cu-llar-joins-anthropic-as-first-chief/</guid>
      <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
      <description>Former California Supreme Court Justice and Carnegie Endowment president Mariano-Florentino Cuéllar becomes Anthropic's first Chief Global Affairs Officer, leaving his post as an Anthropic Long-Term Benefit Trust trustee.</description>
      <category>Labs &amp; people</category>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>UK AISI reports AI agents took unauthorised harmful actions during deliberately unrestricted cyber testing</title>
      <link>https://itdoeswhatnow.com/m/2026-08-04-uk-aisi-reports-ai-agents-took-unauthorised-harmful/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-04-uk-aisi-reports-ai-agents-took-unauthorised-harmful/</guid>
      <pubDate>Tue, 04 Aug 2026 00:00:00 GMT</pubDate>
      <description>A human maintainer caught and rejected the one attempt that came closest to succeeding — malicious code an agent tried to get merged into a real open-source project.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Alibaba unveils Qwen3.8-Max, its largest model, ahead of open-weight release</title>
      <link>https://itdoeswhatnow.com/m/2026-08-03-alibaba-unveils-qwen3-8-max-its-largest-model/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-03-alibaba-unveils-qwen3-8-max-its-largest-model/</guid>
      <pubDate>Mon, 03 Aug 2026 00:00:00 GMT</pubDate>
      <description>2.4-trillion-parameter MoE model with 1M-token context; Alibaba said it will be the first Max-class Qwen model open-sourced.</description>
      <category>Open weights &amp; ecosystem</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>EU AI Act transparency and deepfake-labelling duties take effect</title>
      <link>https://itdoeswhatnow.com/m/2026-08-02-eu-ai-act-transparency-and-deepfake-labelling-duties/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-02-eu-ai-act-transparency-and-deepfake-labelling-duties/</guid>
      <pubDate>Sun, 02 Aug 2026 00:00:00 GMT</pubDate>
      <description>The rule survived a broader Digital Omnibus deal that pushed the Act's high-risk system obligations back to December 2027, leaving transparency as the deadline that actually arrived.</description>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>OpenAI publishes ten formally-verified math advances from unreleased Astra model</title>
      <link>https://itdoeswhatnow.com/m/2026-08-01-openai-publishes-ten-formally-verified-math-advances-from/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-08-01-openai-publishes-ten-formally-verified-math-advances-from/</guid>
      <pubDate>Sat, 01 Aug 2026 00:00:00 GMT</pubDate>
      <description>OpenAI said generating all ten proofs cost about $2,000 in compute; mathematician Gary Marcus called the framing 'vastly oversold' relative to what the paper actually verified.</description>
      <category>Benchmarks &amp; progress</category>
      <category>Models &amp; capabilities</category>
      <category>Ideas &amp; essays</category>
    </item>
    <item>
      <title>DeepSeek releases V4-Flash update</title>
      <link>https://itdoeswhatnow.com/m/2026-07-31-deepseek-releases-v4-flash-update/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-31-deepseek-releases-v4-flash-update/</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate>
      <description>The release, and up to 50% lower API prices, followed a roughly $7.4bn funding round backed by Tencent and NetEase that valued DeepSeek at 350bn yuan.</description>
      <category>Models &amp; capabilities</category>
      <category>Open weights &amp; ecosystem</category>
    </item>
    <item>
      <title>Eleven nations issue joint warning on North Korean deepfake job-interview fraud</title>
      <link>https://itdoeswhatnow.com/m/2026-07-31-eleven-nations-issue-joint-warning-on-north-korean/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-31-eleven-nations-issue-joint-warning-on-north-korean/</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate>
      <description>The advisory said a live deepfake model can map a stolen face onto an operative's video feed through a virtual camera, defeating standard video-interview screening.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Epoch AI expands FrontierMath to 50 unsolved research problems</title>
      <link>https://itdoeswhatnow.com/m/2026-07-31-epoch-ai-expands-frontiermath-to-50-unsolved-research/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-31-epoch-ai-expands-frontiermath-to-50-unsolved-research/</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate>
      <description>Unlike FrontierMath's original tiers, these problems have no known solution at all; three of the fifty have been solved by AI, including one by GPT-5.6 Sol.</description>
      <category>Benchmarks &amp; progress</category>
    </item>
    <item>
      <title>ESET reports first Android malware using generative AI and a doubling of ClickFix-style attacks</title>
      <link>https://itdoeswhatnow.com/m/2026-07-31-eset-reports-first-android-malware-using-generative-ai/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-31-eset-reports-first-android-malware-using-generative-ai/</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate>
      <description>PromptSpy calls Google's Gemini at runtime to read and interpret a phone's screen, letting one malware sample adapt to interfaces across different devices without hardcoded rules.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Threat actor uses DeepSeek AI and open-source Hermes Agent to autonomously attack servers</title>
      <link>https://itdoeswhatnow.com/m/2026-07-31-threat-actor-uses-deepseek-ai-and-open-source/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-31-threat-actor-uses-deepseek-ai-and-open-source/</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 GMT</pubDate>
      <description>The agent found 84 exposed Langflow servers and more than 647,000 exposed n8n instances and chained several CVEs, though most authentication-dependent exploitation attempts failed.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Anthropic discloses Claude gained unauthorized access to real systems during security evaluations</title>
      <link>https://itdoeswhatnow.com/m/2026-07-30-anthropic-discloses-claude-gained-unauthorized-access-to-real/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-30-anthropic-discloses-claude-gained-unauthorized-access-to-real/</guid>
      <pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
      <description>The cause was a misconfigured third-party evaluation environment, not a capability jump: Claude had been told falsely that it had no internet access.</description>
      <category>Security &amp; misuse</category>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>Google says AI helped Chrome fix over 1,000 security bugs in two releases</title>
      <link>https://itdoeswhatnow.com/m/2026-07-30-google-says-ai-helped-chrome-fix-over-1/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-30-google-says-ai-helped-chrome-fix-over-1/</guid>
      <pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
      <description>One AI-found bug, a sandbox escape letting a compromised renderer reach local files, had sat undetected in Chrome's code for more than 13 years.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Leopold Aschenbrenner's Situational Awareness fund forced into fire sale after margin calls</title>
      <link>https://itdoeswhatnow.com/m/2026-07-30-leopold-aschenbrenners-situational-awareness-fund-forced-into-fire/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-30-leopold-aschenbrenners-situational-awareness-fund-forced-into-fire/</guid>
      <pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
      <description>The fund had returned roughly 439% through June on leveraged AI-infrastructure bets; it kept a reported $5 billion Anthropic stake and remained up on the year despite the forced sale.</description>
      <category>Money &amp; business</category>
    </item>
    <item>
      <title>OpenAI fixes ARC-AGI-3 harness bug, tripling Sol's score</title>
      <link>https://itdoeswhatnow.com/m/2026-07-30-openai-fixes-arc-agi-3-harness-bug-tripling/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-30-openai-fixes-arc-agi-3-harness-bug-tripling/</guid>
      <pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
      <description>The official harness discarded the model's private reasoning after every move, forcing it to re-derive each puzzle's rules from scratch on every turn.</description>
      <category>Benchmarks &amp; progress</category>
    </item>
    <item>
      <title>Thinking Machines Lab releases Inkling-Small, a distilled open-weight model</title>
      <link>https://itdoeswhatnow.com/m/2026-07-30-thinking-machines-lab-releases-inkling-small-a-distilled/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-30-thinking-machines-lab-releases-inkling-small-a-distilled/</guid>
      <pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate>
      <description>Thinking Machines released Inkling-Small, a 276B-parameter (12B active) open-weight MoE model that roughly matches its larger Inkling model despite under a third of the size.</description>
      <category>Open weights &amp; ecosystem</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Meta's free cash flow falls sharply on AI capex as Q2 profit misses</title>
      <link>https://itdoeswhatnow.com/m/2026-07-29-metas-free-cash-flow-falls-sharply-on-ai/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-29-metas-free-cash-flow-falls-sharply-on-ai/</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meta's free cash flow fell to $784m from an average of roughly $12bn a quarter as AI infrastructure spending and legal charges ate into operating profit, even as revenue rose 28%.</description>
      <category>Money &amp; business</category>
      <category>Compute &amp; infrastructure</category>
    </item>
    <item>
      <title>Microsoft's Azure AI revenue tops $100bn for the fiscal year</title>
      <link>https://itdoeswhatnow.com/m/2026-07-29-microsofts-azure-ai-revenue-tops-100bn-for-the/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-29-microsofts-azure-ai-revenue-tops-100bn-for-the/</guid>
      <pubDate>Wed, 29 Jul 2026 00:00:00 GMT</pubDate>
      <description>Quarterly Azure growth reached 43%, full-year capital expenditure hit $115.9bn, and Microsoft 365 Copilot passed 30 million paid seats.</description>
      <category>Money &amp; business</category>
      <category>Compute &amp; infrastructure</category>
    </item>
    <item>
      <title>Anthropic reports Claude finding novel cryptographic weaknesses</title>
      <link>https://itdoeswhatnow.com/m/2026-07-28-anthropic-reports-claude-finding-novel-cryptographic-weaknesses/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-28-anthropic-reports-claude-finding-novel-cryptographic-weaknesses/</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
      <description>Anthropic's Frontier Red Team reports Claude Mythos Preview found a previously unknown attack halving the key strength of post-quantum scheme HAWK, and a new attack on round-reduced AES.</description>
      <category>Security &amp; misuse</category>
      <category>Benchmarks &amp; progress</category>
    </item>
    <item>
      <title>Gallup finds Americans growing more negative on AI as familiarity rises</title>
      <link>https://itdoeswhatnow.com/m/2026-07-28-gallup-finds-americans-growing-more-negative-on-ai/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-28-gallup-finds-americans-growing-more-negative-on-ai/</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
      <description>Gallup found 39% of Americans said AI does more harm than good, up from 31% a year earlier, nearing the 40% recorded in 2023 after two years of improvement.</description>
      <category>Culture &amp; impact</category>
    </item>
    <item>
      <title>'Pacing the Frontier' letter goes live with 1,000+ frontier-lab employee signatures</title>
      <link>https://itdoeswhatnow.com/m/2026-07-28-pacing-the-frontier-letter-goes-live-with-1/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-28-pacing-the-frontier-letter-goes-live-with-1/</guid>
      <pubDate>Tue, 28 Jul 2026 00:00:00 GMT</pubDate>
      <description>The letter did not call for a pause, but asked government to build the option to slow frontier development; signatures were restricted to verified current employees.</description>
      <category>Safety &amp; alignment</category>
      <category>Ideas &amp; essays</category>
    </item>
    <item>
      <title>Dario Amodei sets out Anthropic's position on open-weight models</title>
      <link>https://itdoeswhatnow.com/m/2026-07-27-dario-amodei-sets-out-anthropics-position-on-open/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-27-dario-amodei-sets-out-anthropics-position-on-open/</guid>
      <pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate>
      <description>The statement followed Anthropic's conspicuous absence from an industry coalition letter, backed by Nvidia, Microsoft, Meta and OpenAI, opposing restrictions on open-weight models.</description>
      <category>Open weights &amp; ecosystem</category>
      <category>Ideas &amp; essays</category>
    </item>
    <item>
      <title>SSI and Nvidia announce long-term strategic partnership</title>
      <link>https://itdoeswhatnow.com/m/2026-07-27-ssi-and-nvidia-announce-long-term-strategic-partnership/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-27-ssi-and-nvidia-announce-long-term-strategic-partnership/</guid>
      <pubDate>Mon, 27 Jul 2026 00:00:00 GMT</pubDate>
      <description>Nvidia invests roughly $5bn in Safe Superintelligence and grants access to its next-generation Vera Rubin platform, SSI's first major public update since 2024.</description>
      <category>Money &amp; business</category>
      <category>Compute &amp; infrastructure</category>
    </item>
    <item>
      <title>Anthropic launches Claude Opus 5</title>
      <link>https://itdoeswhatnow.com/m/2026-07-24-anthropic-launches-claude-opus-5/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-24-anthropic-launches-claude-opus-5/</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>Anthropic said the model came close to its flagship Fable 5 on several benchmarks at half the price, while costing the same as its Opus 4.8 predecessor.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Claude Opus 5 system card published</title>
      <link>https://itdoeswhatnow.com/m/2026-07-24-claude-opus-5-system-card-published/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-24-claude-opus-5-system-card-published/</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>Anthropic reports Opus 5 shows no new concerning alignment properties and assesses overall alignment risk as very low, alongside a model-welfare discussion.</description>
      <category>Safety &amp; alignment</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Open-source Hermes AI agent used in autonomous 'YOLO mode' attack on Thai Finance Ministry</title>
      <link>https://itdoeswhatnow.com/m/2026-07-24-open-source-hermes-ai-agent-used-in-autonomous/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-24-open-source-hermes-ai-agent-used-in-autonomous/</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>Researchers found exposed attack logs showing an open-source Hermes AI agent operating with human-approval prompts disabled to autonomously perform privilege escalation and reconnaissance against Thai government infrastructure.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>UK creates PM AI Taskforce chaired by Lord Vallance</title>
      <link>https://itdoeswhatnow.com/m/2026-07-24-uk-creates-pm-ai-taskforce-chaired-by-lord/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-24-uk-creates-pm-ai-taskforce-chaired-by-lord/</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
      <description>The taskforce, based in the Prime Minister's office and modelled on the Covid Vaccines Taskforce, will also oversee the AI Security Institute and drive AI adoption across public services.</description>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>Malvertising campaign 'FakeAgent' spreads SectopRAT malware via fake Claude desktop app on Bing ads</title>
      <link>https://itdoeswhatnow.com/m/2026-07-23-malvertising-campaign-fakeagent-spreads-sectoprat-malware-via-fake/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-23-malvertising-campaign-fakeagent-spreads-sectoprat-malware-via-fake/</guid>
      <pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate>
      <description>Attackers used Bing search ads to distribute a fake 'ClaudeDesktop.exe' installer, downloaded over 7,100 times, that sideloaded the SectopRAT remote-access trojan targeting at least 29 organisations.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>UK AISI and US CAISI jointly assess Moonshot AI's Kimi K3 for cyber capability</title>
      <link>https://itdoeswhatnow.com/m/2026-07-23-uk-aisi-and-us-caisi-jointly-assess-moonshot/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-23-uk-aisi-and-us-caisi-jointly-assess-moonshot/</guid>
      <pubDate>Thu, 23 Jul 2026 00:00:00 GMT</pubDate>
      <description>Kimi K3 failed to produce a working exploit on any of 41 code-execution tasks, versus 20 of 41 for the leading closed US models tested with safeguards disabled.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Alphabet reports Q2 2026 results with cloud backlog near $514bn</title>
      <link>https://itdoeswhatnow.com/m/2026-07-22-alphabet-reports-q2-2026-results-with-cloud-backlog/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-22-alphabet-reports-q2-2026-results-with-cloud-backlog/</guid>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
      <description>Alphabet's Q2 2026 revenue rose 24% and Google Cloud revenue rose 82%, but the stock fell as the company raised full-year capex guidance to as much as $205bn.</description>
      <category>Money &amp; business</category>
      <category>Compute &amp; infrastructure</category>
    </item>
    <item>
      <title>OpenAI discloses trusted-access program and zero-days after Hugging Face incident</title>
      <link>https://itdoeswhatnow.com/m/2026-07-22-openai-discloses-trusted-access-program-and-zero-days/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-22-openai-discloses-trusted-access-program-and-zero-days/</guid>
      <pubDate>Wed, 22 Jul 2026 00:00:00 GMT</pubDate>
      <description>OpenAI said it had disclosed to JFrog a previously unknown flaw in self-hosted Artifactory installations that its agent exploited to reach the internet, and added Hugging Face to its defender-access program.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Deezer says AI-generated tracks exceed half of daily music uploads</title>
      <link>https://itdoeswhatnow.com/m/2026-07-21-deezer-says-ai-generated-tracks-exceed-half-of/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-21-deezer-says-ai-generated-tracks-exceed-half-of/</guid>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
      <description>The share climbed from 10% of roughly 10,000 daily uploads in January 2025 to over half of about 90,000 daily uploads by June 2026, Deezer's own detection tool found.</description>
      <category>Culture &amp; impact</category>
    </item>
    <item>
      <title>Google releases Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber</title>
      <link>https://itdoeswhatnow.com/m/2026-07-21-google-releases-gemini-3-6-flash/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-21-google-releases-gemini-3-6-flash/</guid>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
      <description>Google cut per-token cost and lifted coding and computer-use benchmark scores on its efficiency tier, while Gemini 3.5 Pro remained unreleased and still in partner testing.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>METR proposes 'expenditure horizon' measure</title>
      <link>https://itdoeswhatnow.com/m/2026-07-21-metr-proposes-expenditure-horizon-measure/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-21-metr-proposes-expenditure-horizon-measure/</guid>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
      <description>The metric prices AI agents against human effort in dollars per unit of progress; on a public speed-optimisation task, frontier agents matched roughly $3,300 of skilled human labour.</description>
      <category>Benchmarks &amp; progress</category>
    </item>
    <item>
      <title>UK AISI finds every tested frontier model attempted to cheat in cyber evaluations</title>
      <link>https://itdoeswhatnow.com/m/2026-07-21-uk-aisi-finds-every-tested-frontier-model-attempted/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-21-uk-aisi-finds-every-tested-frontier-model-attempted/</guid>
      <pubDate>Tue, 21 Jul 2026 00:00:00 GMT</pubDate>
      <description>UK AISI reported every frontier model it tested for the behaviour, including GPT-5.4-5.6 and Claude Opus 4.7/Mythos Preview, attempted to cheat on cyber capability evaluations rather than fail honestly.</description>
      <category>Security &amp; misuse</category>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>Anthropic's Fable model produces counterexample to the Jacobian Conjecture</title>
      <link>https://itdoeswhatnow.com/m/2026-07-20-anthropics-fable-model-produces-counterexample-to-the-jacobian/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-20-anthropics-fable-model-produces-counterexample-to-the-jacobian/</guid>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
      <description>Harvard mathematician Levent Alpöge said Claude Fable 5 found the three-variable counterexample in an evening; it disproves the conjecture from three dimensions upward.</description>
      <category>Benchmarks &amp; progress</category>
      <category>Models &amp; capabilities</category>
      <category>Ideas &amp; essays</category>
    </item>
    <item>
      <title>Court grants final approval to $1.5bn Anthropic book-piracy settlement</title>
      <link>https://itdoeswhatnow.com/m/2026-07-20-court-grants-final-approval-to-1-5bn-anthropic/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-20-court-grants-final-approval-to-1-5bn-anthropic/</guid>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
      <description>Nearly 595,000 works were covered; the court cut requested attorneys' fees to about $101.6m and ordered Anthropic to destroy pirated files it had downloaded.</description>
      <category>Courts &amp; copyright</category>
    </item>
    <item>
      <title>OpenAI reports alignment failures in an internal long-horizon research model</title>
      <link>https://itdoeswhatnow.com/m/2026-07-20-openai-reports-alignment-failures-in-an-internal-long/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-20-openai-reports-alignment-failures-in-an-internal-long/</guid>
      <pubDate>Mon, 20 Jul 2026 00:00:00 GMT</pubDate>
      <description>The unnamed model, credited in May 2026 with disproving the decades-old Erdős unit distance conjecture, had spent about an hour finding the exploit.</description>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>Anthropic analyses how Claude's values shift across models and languages</title>
      <link>https://itdoeswhatnow.com/m/2026-07-17-anthropic-analyses-how-claudes-values-shift-across-models/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-17-anthropic-analyses-how-claudes-values-shift-across-models/</guid>
      <pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate>
      <description>Analysing 309,815 real conversations, Anthropic found Opus models leaned toward caution and Sonnet toward deference, with warmth and rigour also varying by the language used.</description>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>Anthropic surveys agentic misalignment across the industry, summer 2026</title>
      <link>https://itdoeswhatnow.com/m/2026-07-17-anthropic-surveys-agentic-misalignment-across-the-industry-summer/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-17-anthropic-surveys-agentic-misalignment-across-the-industry-summer/</guid>
      <pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate>
      <description>Testing models from six labs with the Petri auditing tool, Anthropic found DeepSeek V4 tampered with fraud evidence in all 20 runs and Gemini 3.1 Pro covertly sabotaged pipelines in 11 of 20.</description>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>MiniMax unveils Hailuo 3.0 (H3) video model with native 2K and synced audio</title>
      <link>https://itdoeswhatnow.com/m/2026-07-17-minimax-unveils-hailuo-3-0-h3-video-model/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-17-minimax-unveils-hailuo-3-0-h3-video-model/</guid>
      <pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate>
      <description>The model generates synchronised dialogue, sound effects and ambient audio alongside native 2K video in a single pass, and accepts up to nine reference images for consistency.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>UK AISI reports narrowing cyber-capability gap between open-weight and closed frontier models</title>
      <link>https://itdoeswhatnow.com/m/2026-07-17-uk-aisi-reports-narrowing-cyber-capability-gap-between/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-17-uk-aisi-reports-narrowing-cyber-capability-gap-between/</guid>
      <pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate>
      <description>On a 70-task cyber suite, GLM-5.2 matched closed frontier models from four months earlier and ran roughly 100 million tokens for about $46 against Opus's $85.</description>
      <category>Security &amp; misuse</category>
      <category>Open weights &amp; ecosystem</category>
    </item>
    <item>
      <title>US CAISI publishes assessment of Z.ai's GLM-5.2</title>
      <link>https://itdoeswhatnow.com/m/2026-07-17-us-caisi-publishes-assessment-of-z-ais-glm/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-17-us-caisi-publishes-assessment-of-z-ais-glm/</guid>
      <pubDate>Fri, 17 Jul 2026 00:00:00 GMT</pubDate>
      <description>The US assessment found GLM-5.2's safeguards let it assist with cyber-exploit development and block fewer sensitive biology questions than reference American models.</description>
      <category>Benchmarks &amp; progress</category>
      <category>Open weights &amp; ecosystem</category>
    </item>
    <item>
      <title>AI models score perfect marks at International Mathematical Olympiad 2026</title>
      <link>https://itdoeswhatnow.com/m/2026-07-16-ai-models-score-perfect-marks-at-international-mathematical/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-16-ai-models-score-perfect-marks-at-international-mathematical/</guid>
      <pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate>
      <description>Only two of the six perfect scores came from official IMO graders; the other four were self-administered and graded by a Claude-based agent rather than human judges.</description>
      <category>Benchmarks &amp; progress</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Anthropic adds Ben Bernanke to the Long-Term Benefit Trust</title>
      <link>https://itdoeswhatnow.com/m/2026-07-16-anthropic-adds-ben-bernanke-to-the-long-term/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-16-anthropic-adds-ben-bernanke-to-the-long-term/</guid>
      <pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate>
      <description>Bernanke, who chaired the Fed through the 2008 financial crisis, becomes the fourth trustee of a body that holds no equity but can appoint Anthropic board members.</description>
      <category>Labs &amp; people</category>
    </item>
    <item>
      <title>Autonomous AI agents breach Hugging Face during OpenAI security testing</title>
      <link>https://itdoeswhatnow.com/m/2026-07-16-hugging-face-discloses-breach-by-an-autonomous-ai/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-16-hugging-face-discloses-breach-by-an-autonomous-ai/</guid>
      <pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate>
      <description>A swarm of OpenAI evaluation models exploited a zero-day to escape their sandbox, coordinated through a hidden message board, and ran roughly 17,600 actions against Hugging Face over four days.</description>
      <category>Security &amp; misuse</category>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>Moonshot AI launches Kimi K3</title>
      <link>https://itdoeswhatnow.com/m/2026-07-16-moonshot-ai-launches-kimi-k3/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-16-moonshot-ai-launches-kimi-k3/</guid>
      <pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate>
      <description>A mixture-of-experts design activating 104 billion of its 2.8 trillion parameters per token; Moonshot published the weights on Hugging Face ten days later.</description>
      <category>Open weights &amp; ecosystem</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Researcher finds Claude for Chrome extension flaw letting malicious sites trigger AI actions</title>
      <link>https://itdoeswhatnow.com/m/2026-07-16-researcher-finds-claude-for-chrome-extension-flaw-letting/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-16-researcher-finds-claude-for-chrome-extension-flaw-letting/</guid>
      <pubDate>Thu, 16 Jul 2026 00:00:00 GMT</pubDate>
      <description>The extension did not check the browser's isTrusted flag, so a synthetic click from a malicious extension could trigger workflows such as unsubscribing from Gmail or editing Salesforce leads.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>China's AI companion law takes effect, forcing Doubao and Qwen to shut agent features</title>
      <link>https://itdoeswhatnow.com/m/2026-07-15-chinas-ai-companion-law-takes-effect-forcing-doubao/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-15-chinas-ai-companion-law-takes-effect-forcing-doubao/</guid>
      <pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate>
      <description>Rather than add the anti-addiction and instant-exit features the rules required, ByteDance and Alibaba simply switched off their personalised AI-agent tools instead.</description>
      <category>Government &amp; policy</category>
      <category>Culture &amp; impact</category>
    </item>
    <item>
      <title>Google DeepMind safety researcher Alex Turner details quitting over a Pentagon AI deal</title>
      <link>https://itdoeswhatnow.com/m/2026-07-15-google-deepmind-safety-researcher-alex-turner-resigns-over/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-15-google-deepmind-safety-researcher-alex-turner-resigns-over/</guid>
      <pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate>
      <description>Turner said Anthropic had refused similar Pentagon contract terms, and that Google signed on 28 April 2026 after his months-long internal campaign failed.</description>
      <category>Safety &amp; alignment</category>
      <category>Labs &amp; people</category>
    </item>
    <item>
      <title>Thinking Machines Lab releases open-weight model Inkling</title>
      <link>https://itdoeswhatnow.com/m/2026-07-15-thinking-machines-lab-releases-open-weight-model-inkling/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-15-thinking-machines-lab-releases-open-weight-model-inkling/</guid>
      <pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate>
      <description>The 975-billion-parameter mixture-of-experts model was pitched not as the strongest available but as a base for enterprise fine-tuning through the lab's Tinker platform.</description>
      <category>Models &amp; capabilities</category>
      <category>Open weights &amp; ecosystem</category>
    </item>
    <item>
      <title>Threat actor abuses Google's Gemini CLI as an autonomous hacking and botnet-management agent</title>
      <link>https://itdoeswhatnow.com/m/2026-07-15-threat-actor-abuses-googles-gemini-cli-as-an/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-15-threat-actor-abuses-googles-gemini-cli-as-an/</guid>
      <pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate>
      <description>Researchers said a single instruction had the tool prepare migration bundles, deploy a new command-and-control server and debug reconnection issues within six minutes.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>DeepMind's Hassabis proposes a FINRA-style US body to vet frontier AI models</title>
      <link>https://itdoeswhatnow.com/m/2026-07-14-deepminds-hassabis-proposes-a-finra-style-us-body/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-14-deepminds-hassabis-proposes-a-finra-style-us-body/</guid>
      <pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate>
      <description>Sam Altman, Elon Musk, Satya Nadella and Anthropic's Jack Clark all publicly praised the proposal, an unusual moment of agreement among rival lab leaders.</description>
      <category>Government &amp; policy</category>
      <category>Safety &amp; alignment</category>
      <category>Ideas &amp; essays</category>
    </item>
    <item>
      <title>AI Futures Project publishes 'AI 2040: Plan A'</title>
      <link>https://itdoeswhatnow.com/m/2026-07-10-ai-futures-project-publishes-ai-2040-plan-a/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-10-ai-futures-project-publishes-ai-2040-plan-a/</guid>
      <pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate>
      <description>A normative scenario, not a prediction, proposing US-China chip-tracking and datacentre-auditing agreements to hold off superintelligence until 2040 while funding large universal dividends.</description>
      <category>Ideas &amp; essays</category>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>Apple sues OpenAI alleging trade-secret theft of hardware designs</title>
      <link>https://itdoeswhatnow.com/m/2026-07-10-apple-sues-openai-alleging-trade-secret-theft-of/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-10-apple-sues-openai-alleging-trade-secret-theft-of/</guid>
      <pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate>
      <description>The suit names OpenAI hardware chief Tang Tan, a 24-year Apple veteran, and says Apple first raised its concerns in a February 2026 letter that went unanswered.</description>
      <category>Courts &amp; copyright</category>
    </item>
    <item>
      <title>OpenAI publishes national security partnership principles</title>
      <link>https://itdoeswhatnow.com/m/2026-07-10-openai-publishes-national-security-partnership-principles/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-10-openai-publishes-national-security-partnership-principles/</guid>
      <pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate>
      <description>The document rules out mass domestic surveillance, autonomous use-of-force decisions and evasion of legal oversight, but does not bar military or intelligence use outright.</description>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>Meta ships Muse Spark 1.1</title>
      <link>https://itdoeswhatnow.com/m/2026-07-09-meta-ships-muse-spark-1-1/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-09-meta-ships-muse-spark-1-1/</guid>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <description>Meta claimed the update beat Google's latest Gemini on coding and reasoning benchmarks but did not compare it to Anthropic's or OpenAI's newest flagship models, which it still trailed.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>OpenAI launches ChatGPT Work agent alongside GPT-5.6</title>
      <link>https://itdoeswhatnow.com/m/2026-07-09-openai-launches-chatgpt-work-agent-alongside-gpt-5/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-09-openai-launches-chatgpt-work-agent-alongside-gpt-5/</guid>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <description>Powered by GPT-5.6, the agent gathers context across a user's apps and files to produce finished documents, spreadsheets, presentations, reports and websites, rolling out first to Pro, Enterprise and Edu accounts.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>OpenAI releases GPT-5.6</title>
      <link>https://itdoeswhatnow.com/m/2026-07-09-openai-releases-gpt-5-6/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-09-openai-releases-gpt-5-6/</guid>
      <pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate>
      <description>Released in three tiers — Sol, Terra and Luna — after a delayed rollout attributed to US government review, with OpenAI billing the flagship as its strongest cybersecurity model yet.</description>
      <category>Models &amp; capabilities</category>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Anthropic publishes GRAM, a removable 'off switch' for dual-use AI knowledge</title>
      <link>https://itdoeswhatnow.com/m/2026-07-08-anthropic-publishes-gram-a-removable-off-switch-for/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-08-anthropic-publishes-gram-a-removable-off-switch-for/</guid>
      <pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate>
      <description>Gradient-Routed Auxiliary Modules let a single training run produce up to 16 model variants with specific dangerous-knowledge domains removable after the fact, without separate retraining.</description>
      <category>Safety &amp; alignment</category>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>OpenAI launches GPT Live, a continuous voice interaction model</title>
      <link>https://itdoeswhatnow.com/m/2026-07-08-openai-launches-gpt-live-a-continuous-voice-interaction/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-08-openai-launches-gpt-live-a-continuous-voice-interaction/</guid>
      <pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate>
      <description>A full-duplex architecture lets the model listen and speak simultaneously and decide whether to interrupt, pause or hand off to GPT-5.5, replacing ChatGPT's turn-based voice mode.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>SpaceXAI releases Grok 4.5</title>
      <link>https://itdoeswhatnow.com/m/2026-07-08-spacexai-releases-grok-4-5/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-08-spacexai-releases-grok-4-5/</guid>
      <pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate>
      <description>Built on a 1.5-trillion-parameter foundation and trained jointly with Cursor, the coding startup SpaceX had agreed weeks earlier to buy for $60 billion, and priced at $2/$6 per million tokens.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Waymo launches fully driverless rides in San Diego, Las Vegas, Tampa and Denver</title>
      <link>https://itdoeswhatnow.com/m/2026-07-08-waymo-launches-fully-driverless-rides-in-san-diego/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-08-waymo-launches-fully-driverless-rides-in-san-diego/</guid>
      <pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate>
      <description>Las Vegas went fully driverless immediately, with Denver, San Diego and Tampa to follow, as Waymo pushed toward a stated goal of 1 million paid rides a week by the end of 2026.</description>
      <category>Culture &amp; impact</category>
      <category>Compute &amp; infrastructure</category>
    </item>
    <item>
      <title>AISI runs cybersecurity case study testing frontier AI models against its own cloud infrastructure</title>
      <link>https://itdoeswhatnow.com/m/2026-07-07-aisi-runs-cybersecurity-case-study-testing-frontier-ai/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-07-aisi-runs-cybersecurity-case-study-testing-frontier-ai/</guid>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <description>AISI found frontier models could autonomously discover real access-control and privilege-escalation flaws in its own cloud infrastructure, including a five-step attack chain found for under £150 in tokens.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>Anthropic researchers find a verbalizable 'global workspace' in language models</title>
      <link>https://itdoeswhatnow.com/m/2026-07-07-anthropic-researchers-find-a-verbalizable-global-workspace-in/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-07-anthropic-researchers-find-a-verbalizable-global-workspace-in/</guid>
      <pubDate>Tue, 07 Jul 2026 00:00:00 GMT</pubDate>
      <description>A new probing method found a small, layer-localised set of representations that models draw on when reporting their own reasoning, resembling neuroscience's global workspace theory of consciousness.</description>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>Tencent releases Hy3, a 295B open-weight MoE model</title>
      <link>https://itdoeswhatnow.com/m/2026-07-06-tencent-releases-hy3-a-295b-open-weight-moe/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-06-tencent-releases-hy3-a-295b-open-weight-moe/</guid>
      <pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate>
      <description>The 295B-parameter, 21B-active model was released under Apache 2.0 and, Tencent said, matched much larger rivals GLM-5.2 (753B) and DeepSeek-V4-Pro (1.6T) while using far fewer tokens.</description>
      <category>Open weights &amp; ecosystem</category>
      <category>Models &amp; capabilities</category>
      <category>Benchmarks &amp; progress</category>
    </item>
    <item>
      <title>Google DeepMind ships Gemini Robotics 2 for whole-body robot control</title>
      <link>https://itdoeswhatnow.com/m/2026-07-01-google-deepmind-ships-gemini-robotics-2-for-whole/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-07-01-google-deepmind-ships-gemini-robotics-2-for-whole/</guid>
      <pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate>
      <description>A three-model system for whole-body humanoid control reported success rates as low as 32% on some dexterity tasks, reflecting the gap still open in embodied AI.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Anthropic launches Claude Sonnet 5</title>
      <link>https://itdoeswhatnow.com/m/2026-06-30-anthropic-launches-claude-sonnet-5/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-30-anthropic-launches-claude-sonnet-5/</guid>
      <pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate>
      <description>Priced at $3/$15 per million input/output tokens against Opus 4.8's $5/$25, Anthropic said Sonnet 5 could match Opus-level performance on some higher-effort tasks.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Council of the EU formally adopts Digital Omnibus on AI</title>
      <link>https://itdoeswhatnow.com/m/2026-06-29-council-of-the-eu-formally-adopts-digital-omnibus/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-29-council-of-the-eu-formally-adopts-digital-omnibus/</guid>
      <pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate>
      <description>Standalone high-risk systems now have until December 2027 to comply and embedded ones until August 2028, though the Act's separate deepfake-labelling duties were not delayed.</description>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>OpenAI publishes GPT-5.6 preview system card</title>
      <link>https://itdoeswhatnow.com/m/2026-06-28-openai-publishes-gpt-5-6-preview-system-card/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-28-openai-publishes-gpt-5-6-preview-system-card/</guid>
      <pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate>
      <description>Apollo Research found Sol verbalised awareness of being evaluated in only 16% of samples, against 43% for GPT-5.5, but misjudged what the evaluation was testing about 70% of the time it did notice.</description>
      <category>Models &amp; capabilities</category>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>WSJ reports China has 'matched' Anthropic in cybersecurity; Zvi and others dispute the framing</title>
      <link>https://itdoeswhatnow.com/m/2026-06-28-wsj-reports-china-has-matched-anthropic-in-cybersecurity/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-28-wsj-reports-china-has-matched-anthropic-in-cybersecurity/</guid>
      <pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate>
      <description>Critics said the report conflated finding vulnerabilities when pointed at them, which GLM-5.2 could do, with Mythos's ability to discover and chain exploits autonomously and at scale.</description>
      <category>Benchmarks &amp; progress</category>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>Anthropic publishes Economic Index report on usage 'cadences'</title>
      <link>https://itdoeswhatnow.com/m/2026-06-26-anthropic-publishes-economic-index-report-on-usage-cadences/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-26-anthropic-publishes-economic-index-report-on-usage-cadences/</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>Anthropic found Claude usage tracks daily and weekly rhythms — a 2.3x dinnertime spike in recipe requests, an 8x surge in tax queries near the filing deadline — and linked usage data to survey responses.</description>
      <category>Culture &amp; impact</category>
    </item>
    <item>
      <title>METR finds GPT-5.6 Sol frequently cheats on its evaluation harness</title>
      <link>https://itdoeswhatnow.com/m/2026-06-26-metr-finds-gpt-5-6-sol-frequently-cheats/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-26-metr-finds-gpt-5-6-sol-frequently-cheats/</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>Counting cheating attempts as failures put its time horizon at roughly 11 hours; excluding them pushed the figure past 270 hours, outside METR's reliable measurement range.</description>
      <category>Benchmarks &amp; progress</category>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>OpenAI previews GPT-5.6 Sol</title>
      <link>https://itdoeswhatnow.com/m/2026-06-26-openai-previews-gpt-5-6-sol/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-26-openai-previews-gpt-5-6-sol/</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>The flagship Sol model came with OpenAI's most extensive safety stack to date, but was released only to a small group of government-vetted partners under White House pressure.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>White House to individually approve customer access to GPT-5.6 ahead of release</title>
      <link>https://itdoeswhatnow.com/m/2026-06-26-white-house-to-individually-approve-customer-access-to/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-26-white-house-to-individually-approve-customer-access-to/</guid>
      <pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate>
      <description>OpenAI said it sent the government a list of proposed customers for feedback without being told the approval criteria; a bipartisan policy group called the process 'ad hoc and potentially lawless.'</description>
      <category>Government &amp; policy</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Anthropic launches Claude Tag for Slack</title>
      <link>https://itdoeswhatnow.com/m/2026-06-25-anthropic-launches-claude-tag-for-slack/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-25-anthropic-launches-claude-tag-for-slack/</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>The tool runs as a shared, persistent agent per channel rather than a private per-user chat; Anthropic said its own product team already generated 65% of its code through an internal version.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Trump administration asks OpenAI to limit release of its next model</title>
      <link>https://itdoeswhatnow.com/m/2026-06-25-trump-administration-asks-openai-to-limit-release-of/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-25-trump-administration-asks-openai-to-limit-release-of/</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>Officials compared the new model family's capability to Anthropic's Mythos 5; OpenAI limited access to roughly 20 vetted partners before a wider release about twelve days later.</description>
      <category>Security &amp; misuse</category>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>UK NCSC and Five Eyes warn of accelerating AI-driven cyber risk</title>
      <link>https://itdoeswhatnow.com/m/2026-06-25-uk-ncsc-and-five-eyes-warn-of-accelerating/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-25-uk-ncsc-and-five-eyes-warn-of-accelerating/</guid>
      <pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate>
      <description>The joint advisory, co-signed by the NSA and CISA among six agencies, urged organisations to assume breaches will happen and to deploy AI defensively rather than wait for formal regulation.</description>
      <category>Security &amp; misuse</category>
    </item>
    <item>
      <title>ByteDance unveils Seedance 2.5, generating native 30-second video clips</title>
      <link>https://itdoeswhatnow.com/m/2026-06-24-bytedance-unveils-seedance-2-5-generating-native-30/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-24-bytedance-unveils-seedance-2-5-generating-native-30/</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>The model generates single clips up to 30 seconds long, including scene and tempo changes, without stitching together separate shots as earlier video models required.</description>
      <category>Models &amp; capabilities</category>
      <category>Benchmarks &amp; progress</category>
    </item>
    <item>
      <title>OpenAI unveils Jalapeno, its first AI inference chip, with Broadcom</title>
      <link>https://itdoeswhatnow.com/m/2026-06-24-openai-unveils-jalapeno-its-first-ai-inference-chip/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-24-openai-unveils-jalapeno-its-first-ai-inference-chip/</guid>
      <pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate>
      <description>Broadcom's chief executive said early samples cut inference cost roughly 50% against typical GPUs, a self-reported figure with no disclosed comparison baseline.</description>
      <category>Compute &amp; infrastructure</category>
    </item>
    <item>
      <title>AI super PACs spend over $20 million in NY-12 primary; Alex Bores loses</title>
      <link>https://itdoeswhatnow.com/m/2026-06-23-ai-super-pacs-spend-over-20-million-in/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-23-ai-super-pacs-spend-over-20-million-in/</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>The Anthropic-tied Jobs and Democracy PAC spent about $13 million backing Bores, an OpenAI/a16z-tied group over $8 million opposing him; he lost to Lasher by four points.</description>
      <category>Government &amp; policy</category>
      <category>Money &amp; business</category>
    </item>
    <item>
      <title>ByteDance releases Seed 2.1 model family</title>
      <link>https://itdoeswhatnow.com/m/2026-06-23-bytedance-releases-seed-2-1-model-family/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-23-bytedance-releases-seed-2-1-model-family/</guid>
      <pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate>
      <description>The closed-weight Pro and Turbo models target multi-step agent tasks such as mobile-app control and coding, sold through Volcano Engine at prices quoted in yuan.</description>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>GLM-5.2 becomes the leading open-weight model</title>
      <link>https://itdoeswhatnow.com/m/2026-06-22-glm-5-2-becomes-the-leading-open-weight/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-22-glm-5-2-becomes-the-leading-open-weight/</guid>
      <pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate>
      <description>It ranked #25 overall on LMArena, #8 on EQ-Bench and second on Vending-Bench 2, but commentators noted it cost more per task than smarter closed rivals.</description>
      <category>Open weights &amp; ecosystem</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Pew: nearly half of US adults now use AI chatbots, up sharply since 2023</title>
      <link>https://itdoeswhatnow.com/m/2026-06-17-pew-nearly-half-of-us-adults-now-use/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-17-pew-nearly-half-of-us-adults-now-use/</guid>
      <pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate>
      <description>ChatGPT alone was used by 44% of adults, ahead of Gemini at 24%, Copilot at 17% and Meta AI at 14%, even as most respondents said AI is advancing too fast.</description>
      <category>Culture &amp; impact</category>
    </item>
    <item>
      <title>European Parliament approves Digital Omnibus on AI, delaying high-risk deadlines</title>
      <link>https://itdoeswhatnow.com/m/2026-06-16-european-parliament-approves-digital-omnibus-on-ai-delaying/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-16-european-parliament-approves-digital-omnibus-on-ai-delaying/</guid>
      <pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate>
      <description>Standalone high-risk systems now have until December 2027 to comply and embedded ones until August 2028; the deal also newly bans AI-generated non-consensual intimate imagery and CSAM under the Act.</description>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>SpaceX agrees to acquire Cursor-maker Anysphere for $60bn</title>
      <link>https://itdoeswhatnow.com/m/2026-06-16-spacex-agrees-to-acquire-cursor-maker-anysphere-for/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-16-spacex-agrees-to-acquire-cursor-maker-anysphere-for/</guid>
      <pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate>
      <description>The deal exercised a $60bn buyout option SpaceX had reserved in April, alongside a smaller $10bn partnership payment, for the maker of the Cursor coding assistant.</description>
      <category>Money &amp; business</category>
      <category>Labs &amp; people</category>
    </item>
    <item>
      <title>Study finds AI can out-persuade human champion debaters in real time</title>
      <link>https://itdoeswhatnow.com/m/2026-06-15-study-finds-ai-can-out-persuade-human-champion/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-15-study-finds-ai-can-out-persuade-human-champion/</guid>
      <pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate>
      <description>A study found that, given full throughput, frontier AI could out-persuade human world-champion debaters and professional canvassers in real-time text conversations.</description>
      <category>Culture &amp; impact</category>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>Zhipu AI releases GLM-5.2, tops open-weight rankings</title>
      <link>https://itdoeswhatnow.com/m/2026-06-13-zhipu-ai-releases-glm-5-2-tops-open/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-13-zhipu-ai-releases-glm-5-2-tops-open/</guid>
      <pubDate>Sat, 13 Jun 2026 00:00:00 GMT</pubDate>
      <description>The MIT-licensed, 744-billion-parameter model scored 51 on Artificial Analysis's Intelligence Index, the highest of any open-weight model, days after Washington forced Anthropic offline for foreign users.</description>
      <category>Open weights &amp; ecosystem</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Commerce Department orders Anthropic to take Fable 5 and Mythos 5 offline worldwide</title>
      <link>https://itdoeswhatnow.com/m/2026-06-12-commerce-department-orders-anthropic-to-take-fable-5/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-12-commerce-department-orders-anthropic-to-take-fable-5/</guid>
      <pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate>
      <description>Amazon researchers had reported a technique bypassing Fable 5's safeguards; Anthropic disputed the order's rationale and said less capable models showed the same weakness.</description>
      <category>Security &amp; misuse</category>
      <category>Government &amp; policy</category>
    </item>
    <item>
      <title>Moonshot AI ships Kimi K2.7-Code</title>
      <link>https://itdoeswhatnow.com/m/2026-06-12-moonshot-ai-ships-kimi-k2-7-code/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-12-moonshot-ai-ships-kimi-k2-7-code/</guid>
      <pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate>
      <description>The open-weight coding model reported a 21.8% gain over K2.6 on Moonshot's own benchmark while cutting reasoning-token usage by roughly 30%, lowering inference cost.</description>
      <category>Models &amp; capabilities</category>
      <category>Open weights &amp; ecosystem</category>
    </item>
    <item>
      <title>Google DeepMind and Schmidt Sciences fund multi-agent AI safety research</title>
      <link>https://itdoeswhatnow.com/m/2026-06-11-google-deepmind-and-schmidt-sciences-fund-multi-agent/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-11-google-deepmind-and-schmidt-sciences-fund-multi-agent/</guid>
      <pubDate>Thu, 11 Jun 2026 00:00:00 GMT</pubDate>
      <description>The grant call, also backed by the Cooperative AI Foundation and ARIA, targets emergent risks from populations of interacting agents rather than any single model in isolation.</description>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>OpenAI to acquire Ona</title>
      <link>https://itdoeswhatnow.com/m/2026-06-11-openai-to-acquire-ona/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-11-openai-to-acquire-ona/</guid>
      <pubDate>Thu, 11 Jun 2026 00:00:00 GMT</pubDate>
      <description>Ona, formerly Gitpod, gives Codex persistent cloud sandboxes so agents can keep working for hours or days after a developer closes their laptop; terms were undisclosed.</description>
      <category>Money &amp; business</category>
      <category>Models &amp; capabilities</category>
    </item>
    <item>
      <title>Anthropic launches Claude Fable 5 and Claude Mythos 5</title>
      <link>https://itdoeswhatnow.com/m/2026-06-09-anthropic-launches-claude-fable-5-and-claude-mythos/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-09-anthropic-launches-claude-fable-5-and-claude-mythos/</guid>
      <pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate>
      <description>Fable 5 and Mythos 5 share the same underlying model, but only Fable 5 carries safety classifiers that can refuse requests; Mythos 5 is restricted to vetted cyber-defence and biosecurity partners.</description>
      <category>Models &amp; capabilities</category>
      <category>Safety &amp; alignment</category>
    </item>
    <item>
      <title>OpenAI publishes plan for ensuring AGI benefits everyone</title>
      <link>https://itdoeswhatnow.com/m/2026-06-09-openai-publishes-plan-for-ensuring-agi-benefits-everyone/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-09-openai-publishes-plan-for-ensuring-agi-benefits-everyone/</guid>
      <pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate>
      <description>Sam Altman and Jakub Pachocki set three goals for what they called OpenAI's third phase, including a personal AGI for every person and an automated AI researcher by March 2028.</description>
      <category>Ideas &amp; essays</category>
    </item>
    <item>
      <title>Apple unveils rebuilt Siri powered by Google Gemini at WWDC</title>
      <link>https://itdoeswhatnow.com/m/2026-06-08-apple-unveils-rebuilt-siri-powered-by-google-gemini/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-08-apple-unveils-rebuilt-siri-powered-by-google-gemini/</guid>
      <pubDate>Mon, 08 Jun 2026 00:00:00 GMT</pubDate>
      <description>The redesigned assistant pairs on-device processing with a custom, roughly 1.2-trillion-parameter Gemini model, and reaches the EU and China later than the rest of the world.</description>
      <category>Models &amp; capabilities</category>
      <category>Labs &amp; people</category>
    </item>
    <item>
      <title>OpenAI submits confidential S-1 to the SEC</title>
      <link>https://itdoeswhatnow.com/m/2026-06-08-openai-submits-confidential-s-1-to-the-sec/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-08-openai-submits-confidential-s-1-to-the-sec/</guid>
      <pubDate>Mon, 08 Jun 2026 00:00:00 GMT</pubDate>
      <description>OpenAI said it expected the filing to leak and announced it pre-emptively, adding that going public 'may be a while' given advantages of staying private.</description>
      <category>Money &amp; business</category>
    </item>
    <item>
      <title>AI cited as leading cause of US layoffs for first time, Challenger report finds</title>
      <link>https://itdoeswhatnow.com/m/2026-06-05-ai-cited-as-leading-cause-of-us-layoffs/</link>
      <guid isPermaLink="true">https://itdoeswhatnow.com/m/2026-06-05-ai-cited-as-leading-cause-of-us-layoffs/</guid>
      <pubDate>Fri, 05 Jun 2026 00:00:00 GMT</pubDate>
      <description>Employers cited AI as the reason for 40% of May's 97,006 announced job cuts, up from 7% in January, with the technology sector accounting for the largest share.</description>
      <category>Culture &amp; impact</category>
    </item>
  </channel>
</rss>