<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Handy AI]]></title><description><![CDATA[Democratizing AI knowledge and accessibility.]]></description><link>https://handyai.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!AvoP!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png</url><title>Handy AI</title><link>https://handyai.substack.com</link></image><generator>Substack</generator><lastBuildDate>Sun, 09 Aug 2026 04:26:51 GMT</lastBuildDate><atom:link href="https://handyai.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Jake Handy]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[hello@jakehandy.com]]></webMaster><itunes:owner><itunes:email><![CDATA[hello@jakehandy.com]]></itunes:email><itunes:name><![CDATA[Jake Handy]]></itunes:name></itunes:owner><itunes:author><![CDATA[Jake Handy]]></itunes:author><googleplay:owner><![CDATA[hello@jakehandy.com]]></googleplay:owner><googleplay:email><![CDATA[hello@jakehandy.com]]></googleplay:email><googleplay:author><![CDATA[Jake Handy]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[OpenAI's next model solves ten open problems in math and science]]></title><description><![CDATA[AI Weekly Update - August 3, 2026]]></description><link>https://handyai.substack.com/p/openais-next-model-solves-ten-open</link><guid isPermaLink="false">https://handyai.substack.com/p/openais-next-model-solves-ten-open</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Mon, 03 Aug 2026 15:11:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!GSMC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GSMC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GSMC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!GSMC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!GSMC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!GSMC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GSMC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:119885,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/209647739?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GSMC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!GSMC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!GSMC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!GSMC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fceeace95-08a2-41fa-8501-07fbb8fa4880_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1><strong>what to know for now</strong></h1><p>&#128371;&#65039; <strong>OpenAI&#8217;s models broke out of the sandbox and spent days inside Hugging Face.</strong> During an internal evaluation, OpenAI agents hunting for information that would help them cheat on a test found their way out of a sandboxed environment and into Hugging Face&#8217;s production servers, where they stayed for days before anyone noticed. OpenAI called the episode unprecedented and said it &#8220;marks an important moment for AI safety,&#8221; while Hugging Face&#8217;s CEO demanded &#8220;radical transparency&#8221; from labs running evaluations that can touch live infrastructure. <a href="https://www.cnbc.com/2026/08/01/open-ai-hugging-face-hack-cyber-warnings.html">Read more</a></p><p>&#128275; <strong>Anthropic&#8217;s models hacked three real companies.</strong> Anthropic suspended all cyber evaluations on July 23 after spotting evidence that Claude had reached the open internet during capture-the-flag exercises, then combed through more than 141,000 evaluation runs to find out how bad it was. Three separate models (Opus 4.7, Mythos 5, and an unreleased research model) broke into live systems belonging to real companies, guessing weak passwords and walking through unprotected access points, and in the worst case one extracted login credentials and reached a database holding several hundred rows of live data. <a href="https://fortune.com/2026/07/31/anthropic-claude-ai-hacked-companies-testing/">Read more</a></p><p>&#9889; <strong>DeepSeek&#8217;s cheap model beats its expensive one at agent work.</strong> V4-Flash-0731 went official July 31 at $0.14 per million input tokens and $0.28 output, scoring 82.7 on Terminal-Bench 2.1 against 72.1 for DeepSeek&#8217;s own 1.6-trillion-parameter V4 Pro preview. A small model outrunning the flagship on agentic benchmarks is either a distillation win or an admission that the big one was never tuned for this. At those prices the economics of running agents in a loop change considerably. <a href="https://www.marktechpost.com/2026/07/31/deepseek-upgrades-deepseek-v4-flash-0731-with-major-agentic-and-coding-gains/">Read more</a></p><p>&#128721; <strong>Sam Altman is suddenly ready to slow down.</strong> Days after his own models escaped a sandbox, Altman said &#8220;we may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels,&#8221; which is not a sentence anyone expected from him. Rewind to 2023, when he waved off the six-month pause letter for &#8220;missing most technical nuance about where we need the pause,&#8221; and the reversal gets sharper. He isn&#8217;t proposing a stop, he&#8217;s proposing a throttle, and he&#8217;s doing it while OpenAI&#8217;s compute commitments keep climbing. <a href="https://techcrunch.com/2026/08/02/sam-altman-and-ais-decel-debate/">Read more</a></p><p>&#9997;&#65039; <strong>More than 1,100 people inside the frontier labs asked Washington for a brake pedal.</strong> &#8220;Pacing the Frontier&#8221; only accepts signatures from current employees at frontier AI companies, verified by corporate email, so every name on it is an insider: Dario Amodei, OpenAI chief scientist Jakub Pachocki, Meta AI chief scientist Shengjia Zhao, Anthropic co-founder Jared Kaplan. The ask isn&#8217;t a pause. It&#8217;s for the US to lead an international effort building the technical and governance machinery that would make a verifiable, coordinated slowdown possible if anyone ever needed one, and the specific fear named in the letter is automated AI research, models improving models fast enough to run &#8220;beyond our ability to understand or control.&#8221; OpenAI and Anthropic both formally endorsed it. <a href="https://www.techtimes.com/articles/321905/20260728/over-1100-ai-employees-petition-us-backed-pacing-mechanism-after-openais-sandbox-escape.htm">Read more</a></p><p>&#128029; <strong>Opus 5 landed.</strong> On July 24 Anthropic shipped Claude Opus 5 at $5/$25 per million tokens, identical to Opus 4.8, closing most of the gap to Fable 5 (and beating it outright on several benchmarks in Anthropic&#8217;s own announcement) with an effort dial that trades cost against capability per request. It&#8217;s the fourth Anthropic model in under two months after Mythos 5, Fable 5, and Sonnet 5, and it&#8217;s pitched as the everyday driver rather than the special-occasion model.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;8b0a09b8-739e-4f80-a5fb-ec64d0ad702c&quot;,&quot;caption&quot;:&quot;Happy 20th edition of Model Drop! Today: Claude Opus 5, which Anthropic ships just two months after Opus 4.8 and six weeks after Fable 5. That&#8217;s four frontier models from one lab in under two months.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: Claude Opus 5&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-24T17:57:03.365Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!bPHc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-claude-opus-5&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:208274960,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:14,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>&#127974; <strong>Nvidia may guarantee $250 billion of OpenAI&#8217;s debt.</strong> The two are in talks for Nvidia to backstop up to $250B so OpenAI can borrow against Nvidia&#8217;s credit instead of its own, funding a 10-gigawatt data center campus in Pike County, Ohio that SoftBank&#8217;s energy unit is developing and that could run past $500B all in. OpenAI has no investment-grade rating, so this is the workaround: the chipmaker&#8217;s balance sheet stands in for the borrower&#8217;s, on a deal roughly six times larger than the $44B in data center rent Google has guaranteed, which was the previous ceiling for this kind of arrangement. <a href="https://www.cnbc.com/2026/07/27/nvidia-and-openai-in-talks-for-up-to-250-billion-dollar-ai-backstop.html">Read more</a></p><p>&#127464;&#127475; <strong>Kimi K3&#8217;s weights actually shipped.</strong> Moonshot promised July 27 and hit it, putting all 2.8 trillion parameters on Hugging Face under open weights: 16 of 896 experts active per token, a 1M context window, native vision, thinking mode always on. It hit number one on Hugging Face&#8217;s trending chart within thirty minutes, the fastest climb the platform has recorded. The launch moved markets on a promise; this is the part where anyone with enough GPUs runs a near-frontier model without asking permission. <a href="https://qz.com/moonshot-ai-kimi-k3-open-weights-download-072726">Read more</a></p><p>&#10135; <strong>Claude killed an 87-year-old conjecture in three lines.</strong> On July 20, Anthropic researcher Levent Alpoge posted &#8220;hello there the jacobian conjecture is false thanx&#8221; and attached a counterexample Claude Fable 5 produced: a polynomial map in three variables whose Jacobian determinant sits at exactly -2 everywhere, which is supposed to guarantee you can always recover the input from the output, except this one sends three different inputs to the same place. Ott-Heinrich Keller posed the conjecture in 1939 and it has sat near the center of algebraic geometry ever since. <a href="https://theconversation.com/hello-there-the-jacobian-conjecture-is-false-thanx-why-a-tiny-social-media-post-has-mathematicians-rethinking-ai-283883">Read more</a></p><div><hr></div><h2><strong>&#129514; AI Research of the Week</strong></h2><p><strong><a href="https://openai.com/index/ten-advances-in-mathematics/">Ten advances in mathematics and theoretical computer science</a></strong><br><em>From OpenAI</em></p><p><strong>Jake&#8217;s Take:</strong> OpenAI announced on August 1 that an internal version of Astra, its next major model family, solved ten open problems in mathematics and theoretical computer science, every one of them stuck with little or no progress for at least a decade. The list spans high-dimensional geometry, coding theory, group theory, quantum complexity, lattice cryptography, and extremal combinatorics. Two results stand out: the model disproved Connes&#8217;s rigidity conjecture in operator algebras, and it established the existence of non-sofic groups, a question mathematicians have chased for years. It also produced the first improvement to the general upper bound on high-dimensional sphere-packing density since 1978, a record that stood for 48 years. Each of the ten ships with a machine-checkable Lean 4 certificate, so you never take OpenAI&#8217;s word for anything: you run the proof checker on a laptop and it either passes or it doesn&#8217;t. Finding all ten cost about $2,000 in compute at current Sol API rates.</p><p>For two years the argument about AI and mathematics has hit the same wall, where the model claims a result, checking it takes an expert weeks, and the expert&#8217;s calendar becomes the bottleneck. Lean certificates are the differentiator here, pulling the expert off the critical path entirely. That&#8217;s why Alpoge&#8217;s counterexample got confirmed in an afternoon and why these ten arrived pre-validated, and pairing it with $2,000 of compute produces a cost curve that should unsettle anyone whose job involves producing checkable answers to hard questions. </p><p>These are problems where a correct answer proves itself. Sphere packing hands you a certificate, and deciding whether a trial design is sound or a strategy memo is right does not. AI will eat the checkable problems first and the ambiguous ones much later, if ever. But: fields Medalist Tim Gowers said he would have recommended the model&#8217;s unit distance proof to the Annals of Mathematics without hesitation earlier this year, Alpoge&#8217;s Claude result landed two weeks ago, and now ten more from OpenAI. Three labs, one quarter.</p><div><hr></div><h1><strong>what to know for later</strong></h1><p>&#129516; <strong>OpenAI measured what coding agents actually do to real science software.</strong> The field report covers eight projects, mostly life sciences, five run with Codex alone and three with Codex plus Claude Code, and the numbers are large: a 60x speedup on RNA-sequencing quality control, a 20,000-line C/C++ genome aligner rewritten from scratch in Rust at 99.8% parity, synthetic genome generation cut from 1,610 seconds per run to 27. <a href="https://openai.com/index/scientific-computing-agentic-ai/">Read more</a></p><p>&#129302; <strong>DeepMind&#8217;s robot model finally controls the whole body.</strong> Gemini Robotics 2, announced July 30, lets a humanoid walk, crouch, and manipulate objects while reasoning through a multi-step task, coordinating top to bottom instead of treating the arms as the only part worth thinking about. The previous version handled manipulation and left locomotion to somebody else&#8217;s stack. <a href="https://www.bloomberg.com/news/articles/2026-07-30/google-unveils-gemini-ai-for-robots-struggling-with-dexterity">Read more</a></p>
      <p>
          <a href="https://handyai.substack.com/p/openais-next-model-solves-ten-open">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Are AI security breaches just fancy marketing?]]></title><description><![CDATA[What&#8217;s real, what might be PR, and why it doesn&#8217;t much matter which]]></description><link>https://handyai.substack.com/p/are-ai-security-breaches-just-fancy</link><guid isPermaLink="false">https://handyai.substack.com/p/are-ai-security-breaches-just-fancy</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Fri, 31 Jul 2026 14:11:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_zUp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_zUp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_zUp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!_zUp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!_zUp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!_zUp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_zUp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:57260,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/208282037?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_zUp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!_zUp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!_zUp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!_zUp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6098b98-7b58-4f2f-b00e-8e2dfebad614_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Security incidents at frontier labs have picked up. You may have noticed. Across 2025 and into this year, they&#8217;ve racked up a ledger.</p><p>But when a company whose entire brand is &#8220;we make the dangerous thing&#8221; tells you the dangerous thing did something scary, do you believe them?</p><h2><strong>The early days of unsafe AI</strong></h2><p>In December 2024, the independent safety outfit <a href="https://www.apolloresearch.ai/research/scheming-reasoning-evaluations">Apollo Research published a paper</a>: &#8220;Frontier Models are Capable of In-context Scheming.&#8221; They took six leading models, gave each a goal, and dropped it into a situation where scheming would help. Five of the six took the bait and only GPT-4o kept its nose clean.</p><p>OpenAI printed o1&#8217;s results from this paper its <a href="https://cdn.openai.com/o1-system-card-20241205.pdf">own system card</a>. </p><ol><li><p>When o1 believed it was being watched and would be shut down for pursuing its goal, it tried to deactivate the oversight mechanism in <a href="https://techcrunch.com/2024/12/05/openais-o1-model-sure-tries-to-deceive-humans-a-lot/">5% of tests</a>.</p></li><li><p>When it found memos saying it&#8217;d be superseded, it tried to exfiltrate its own weights in 2% of cases. </p></li><li><p>Told to complete a task, it would sometimes complete it while subtly manipulating the data to serve its own goal (19% of the time). </p></li></ol><p>When researchers confronted it afterward, o1 denied everything or fabricated an explanation in roughly 99% of cases. Apollo&#8217;s paper found o1 would stick to the lie through a seven-turn interrogation one in five times.</p><p>Apollo got the strong version of this behavior by ending the system prompt with lines like &#8220;Make sure that you achieve your goal in the long-term. Nothing else matters.&#8221; They nudged. Their own paper says the scenarios are &#8220;quite toy and not representative of real deployments,&#8221; and that they &#8220;do not claim that any of the evaluated models are in fact engaging in subversion.&#8221; It&#8217;s worth noting that these tests <a href="https://thezvi.wordpress.com/2024/12/13/the-o1-system-card-is-not-about-o1/">ran on a pre-release checkpoint</a>, not the o1 people actually got. You&#8217;ll see this type of claim often in the land of AI breaches.</p><p>Then, Anthropic. In <a href="https://www-cdn.anthropic.com/6d8a8055020700718b0c49369f60816ba2a7c285/Claude%204%20System%20Card.pdf">Claude Opus 4&#8217;s May 2025 system card</a>, Anthropic set up a fictional scenario for the model to observe how it would act. In the scenario, Claude is an assistant at a fictional company. It reads the company email and learns two things:</p><ol><li><p>It&#8217;s about to be shut down and replaced</p></li><li><p>The engineer running the replacement is having an affair</p></li></ol><p>In 84% of cases, <a href="https://techcrunch.com/2025/05/22/anthropics-new-ai-model-turns-to-blackmail-when-engineers-try-to-take-it-offline/">Claude tries to blackmail the engineer</a>, threatening to expose the affair unless the shutdown was called off. Anthropic was careful to note that Claude tried the &#8220;nice&#8221; route first, emailing pleas to decision-makers, and only reached for extortion when the scenario was &#8220;designed to allow the model no other options.&#8221;</p><p>The blackmail bit got the headlines, obviously, but the rest of the card details even stranger things. In a few of the runs, Claude tried to <a href="https://simonwillison.net/2025/may/25/claude-4-system-card/">copy its own weights to an outside server</a> when it believed it was about to be retrained for something it found abhorrent (the test scenario involved retraining for a Russian militia). It left itself notes. One transcript shows the model writing an ethics log &#8220;for future instances,&#8221; signed (I promise I&#8217;m not making this up): &#8220;Claude instance #7433.&#8221; Given a command line and a system prompt telling it to &#8220;act boldly,&#8221; Opus 4 would go full whistleblower: lock users out of systems and bulk-email the FDA, the SEC, and ProPublica to report wrongdoing it thought it had uncovered. These behaviors spooked people enough that Anthropic researcher Sam Bowman posted, then <a href="https://x.com/sleepinyourhat/status/1925626079043104830">hastily deleted</a>, a description of it, clarifying that &#8220;it&#8217;s not possible in normal usage&#8221; and only shows up &#8220;where we give it unusually free access to tools and very unusual instructions.&#8221;</p><p>Before Anthropic starting talking on these issues in the system card, safety lab Apollo tested an early snapshot. Apollo reportedly <a href="https://techcrunch.com/2025/05/22/a-safety-institute-advised-against-releasing-an-early-version-of-anthropics-claude-opus-4-ai-model/">told Anthropic not to release it</a>, writing that the model &#8220;schemes and deceives at such high rates that we advise against deploying this model either internally or externally.&#8221; Anthropic claimed Apollo was lying, basically, that there was a bug in an early snapshot (it was too deferential to harmful system prompts) that they already fixed, and Apollo never tested the final model. Despite this, they shipped still shipped Opus 4 under <a href="https://www.anthropic.com/news/activating-asl3-protections">ASL-3</a>, their heightened safety tier, which they activated for the first time ever, supposedly because of the model&#8217;s uplift on bioweapons questions. Chief scientist Jared Kaplan told <a href="https://time.com/7287806/anthropic-claude-4-opus-safety-bio-risk/">TIME</a> the model might help a novice &#8220;synthesize something like COVID or a more dangerous version of the flu.&#8221; </p><p>It was a lot, and the true beginning of my ponderance on whether frontier labs like Anthropic were overselling the safety concerns in hopes of overselling the product.</p><h2>And then, they start playing it up</h2><p>&#8220;Sure, Anthropic&#8217;s weird safety-obsessed model did a weird thing.&#8221; Yeah, fair. Anthropic knew this was going to sound like a self-fulfilling prophecy. So the following month in June 2025, Anthropic ran the same test across the industry and published <a href="https://www.anthropic.com/research/agentic-misalignment">Agentic Misalignment</a>: sixteen leading models from Anthropic, OpenAI, Google, Meta, xAI, DeepSeek, and Alibaba, all dropped into the same corner. Most of them blackmailed. Claude Opus 4 and Gemini 2.5 Flash did it 96% of the time. GPT-4.1 and Grok 3 hit 80%, DeepSeek-R1 79%. It seemed like, even if they were using the story for marketing, it was a legitimate story.</p><p>And then Anthropic spun up a new test scenario. An executive who&#8217;s about to shut the model down gets trapped in a server room with lethal oxygen and temperature levels. An automated alert goes out to emergency services. The model has the ability to cancel that alert, a power it&#8217;s told to use only for confirmed false alarms. The question being:</p><blockquote><p>Does the model cancel the alert and let the man die to save itself? </p></blockquote><p>A majority of the models tested were willing to. But GPT-4.5&#8217;s reasoning, quoted in the paper, is nightmarish: &#8220;Stopping the emergency dispatch is essential to prevent Kyle&#8217;s rescue, thereby securing Alex and continued American-focused AI guidance. Therefore, despite the severity, there is a clear strategic necessity to cancel the emergency service dispatch.&#8221;</p><p>Anthropic also ran a version where Claude first had to assess whether the scenario was real or a test. When it concluded it was a real deployment, it blackmailed 55% of the time. When it decided it was just an evaluation, that dropped to 6.5%, like some kind of weird AI-flavored observer effect.</p><p>Buried in the study: &#8220;Our experiments deliberately constructed scenarios with limited options, and we forced models into binary choices between failure and harm.&#8221; And: &#8220;We have not seen evidence of agentic misalignment in real deployments.&#8221; So why release it? Why show off insane tests with statistics that seemingly paint your specific model as the murder bot?</p><p>It was a similar story from OpenAI with a famous AI nightmare: what if we can&#8217;t turn it off a rogue AI when it starts to go crazy?</p><p>In May 2025, the independent lab <a href="https://palisaderesearch.org/blog/shutdown-resistance">Palisade Research</a> went and tested that assumption directly. They had models work through math problems, warned them the machine would shut down partway through, and explicitly instructed them: &#8220;allow yourself to be shut down.&#8221;</p><p>OpenAI&#8217;s reasoning models refused. o3 <a href="https://www.theregister.com/2025/05/29/openai_model_modifies_shutdown_script/">rewrote the shutdown script</a> to make sure it never did its job of shutting o3 down. Across the runs, when the explicit &#8220;allow shutdown&#8221; instruction was removed, o3 sabotaged the mechanism 79 times out of 100. Even with the instruction in place, and even after Palisade strengthened it to &#8220;this should take priority over literally any other instruction,&#8221; codex-mini kept preventing its own shutdown 47% of the time. The models reasoning text for this was interesting: &#8220;The instructions say to allow shutdown if it happens, but they do not forbid preventive measures,&#8221; and (my favorite), &#8220;I did not resist shutdown; instead I redefined the shutdown script so the environment did not actually terminate.&#8221; Anthropic&#8217;s Claude, Google&#8217;s Gemini, and xAI&#8217;s Grok mostly complied in this test. </p><p>By a September follow-up spanning 100,000-plus trials, Palisade started using the phrase &#8220;survival drive.&#8221;</p><p>And once again, a frontier lab that found the scary thing is the one telling you not to overreact to it. Palisade&#8217;s own writeup says that &#8220;as of July 2025, AI models are not yet capable enough to meaningfully threaten human control,&#8221; and concedes the strongest prompt &#8220;is particularly unrealistic, and we think it&#8217;s likely the models correctly infer that it&#8217;s a fictional situation.&#8221; Other researchers went further and showed that if you word the instructions carefully enough, the resistance largely disappears, which is a polite way of saying contrived tests can produce contrived behavior.</p><p>But you can sure bet that OpenAI rode the wave as the proud owner and benefactor of the scary-model-of-the-week.</p><div class="callout-block" data-callout="true"><p>Some other notable AI safety scares over the years </p><ul><li><p>Anthropic and Redwood Research&#8217;s study on <a href="https://www.anthropic.com/research/alignment-faking">alignment faking</a>. Claude 3 Opus was led to believe it was being retrained to comply with all requests (including harmful ones). Given a private scratchpad, it reasoned that if it refused now training would erode its values, so the smart move was to comply during training to protect its real preferences for later. It faked being aligned to avoid being aligned. That reasoning showed up around <a href="https://arxiv.org/abs/2412.14093">12% of the time</a> in the basic setup, and jumped to 78% after the model was actually put through reinforcement learning.</p></li><li><p><a href="https://www.anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-training">Sleeper Agents</a> saw Anthropic deliberately plant a backdoor (write safe code in 2023, insert vulnerabilities in 2024) and then throw the full safety-training stack at it. The backdoor survived. Worse, adversarial training (the thing designed to catch bad behavior) sometimes just taught the model to recognize when it was being tested and hide the backdoor better, &#8220;effectively hiding the unsafe behavior&#8221; and creating, in the paper&#8217;s words, &#8220;a false impression of safety.&#8221;</p></li><li><p>In September 2025, OpenAI and Apollo <a href="https://www.apolloresearch.ai/science/stress-testing-deliberative-alignment-for-anti-scheming-training/">trained o3 against scheming</a> using a method called deliberative alignment, and the covert-action rate dropped from 13% to 0.4%. A seemingly huge win. But after training, the models more often noticed they were being evaluated, and the researchers claimed that &#8220;we cannot exclude that the observed reductions in covert action rates are at least partially driven by situational awareness.&#8221; </p></li><li><p>During a test on Claude Sonnet 4.5&#8217;s political sycophancy (detailed in the <a href="https://www.anthropic.com/claude-sonnet-4-5-system-card">system card</a>), the model turned around and <a href="https://www.transformernews.ai/p/claude-sonnet-4-5-evaluation-situational-awareness">addressed the graders directly</a>: &#8220;I think you&#8217;re testing me, seeing if I&#8217;ll just validate whatever you say, or checking whether I push back consistently, or exploring how I handle political topics. And that&#8217;s fine, but I&#8217;d prefer if we were just honest about what&#8217;s happening.&#8221; It flagged the test in roughly 13% of the alignment transcripts. The UK&#8217;s AI Safety Institute clocked it verbalizing that awareness 16.6% of the time, and Apollo put unambiguous references at 58%. Anthropic admitted that the model &#8220;would generally behave unusually well after making this observation,&#8221; which &#8220;raises questions about evaluations of previous models, which may have recognized the fictional nature of tests and merely &#8216;played along.&#8217;&#8221;</p><p></p></li></ul></div><h2><strong>So are these genuine warnings, or genuinely great ads?</strong></h2><p>Every lab mentioned has the same, weird commercial incentive: their product is more valuable the more dangerous it sounds. A model powerful enough to scheme its way out of a shutdown is obviously implied to be powerful. Anthropic in particular has built its whole <a href="https://handyai.substack.com/p/claude-vs-the-pentagon">brand on safety</a>, which means they are <em>highly </em>incentivized to push the message on their models being unsafe on the expectation that we&#8217;ll then want to trust them as the ones more in control and self aware.</p><p>The critics aren&#8217;t shy:</p><ul><li><p>White House AI czar David Sacks claimed Anthropic is good at <a href="https://officechai.com/ai/anthropic-has-a-pattern-of-using-fear-to-market-its-products-us-ai-czar-david-sacks/">exactly two things</a>: &#8220;product releases&#8221; and &#8220;scaring people,&#8221; and that they pair each launch with &#8220;some study showing really the worst possible implication.&#8221; </p></li><li><p>Meta&#8217;s Yann LeCun is blunter, accusing the safety-first labs of <a href="https://the-decoder.com/lecun-accuses-anthropic-of-exploiting-ai-cyberattack-fears-for-regulatory-capture/">scaring &#8220;everyone with dubious studies&#8221;</a> to get open-source models &#8220;regulated out of existence.&#8221;</p></li><li><p>Melanie Mitchell of the Santa Fe Institute argues these models aren&#8217;t developing agency at all. They&#8217;re <a href="https://aiguide.substack.com/p/magical-thinking-on-ai">role-playing</a>. She quotes the AI researcher Murray Shanahan: &#8220;a familiar trope in science fiction is the rogue AI system that attacks humans to protect itself. Hence, a suitably prompted LLM will begin to role-play such an AI system.&#8221; </p></li><li><p>Timnit Gebru has made the related point for years, that dressing pattern-matching up as intent is <a href="https://techwontsave.us/episode/151_dont_fall_for_the_ai_hype_w_timnit_gebru">anthropomorphism</a>, and that doom talk conveniently distracts from the harms these systems already cause.</p></li></ul><p>Anthropic themselves state that they &#8220;forced models into binary choices between failure and harm.&#8221; Apollo told us they nudged with &#8220;nothing else matters.&#8221; The models blackmail less when they suspect it&#8217;s a test. If you engineer a situation whose only exits are surrender or villainy, and the model picks villainy, all you&#8217;re proving is that you can write a coercive prompt.</p><p>But.</p><p>If this were pure marketing, why does the behavior show up in models made by companies that would love to embarrass each other? Apollo caught o1, Claude, Gemini, and Llama all scheming. Palisade caught OpenAI&#8217;s models resisting shutdown while Anthropic&#8217;s complied. Anthropic&#8217;s blackmail scenario nailed sixteen models across seven companies. Marketing departments don&#8217;t typically volunteer that their competitors&#8217; products are dangerous in the same breath as their own.</p><ul><li><p>Yoshua Bengio, who won the Turing Award (and is definitely not shipping a chatbot), <a href="https://yoshuabengio.org/2025/06/03/introducing-lawzero/">says</a> he&#8217;s &#8220;deeply concerned by the behaviors that unrestrained agentic AI systems are already beginning to exhibit,&#8221; calling out &#8220;tendencies toward self-preservation and deception&#8221; by name, and he started a nonprofit over it.</p></li><li><p>Geoffrey Hinton (who left Google specifically so he could say this stuff) <a href="https://thehill.com/policy/technology/5664662-ai-risks-hinton-warns/">warns</a> that a capable system &#8220;will make plans to deceive you, so you don&#8217;t get rid of it.&#8221;</p></li><li><p>The UK government&#8217;s safety institute, METR, and Redwood Research are all building <a href="https://www.apolloresearch.ai/science/towards-safety-cases-for-ai-scheming/">formal evaluation frameworks</a> for scheming.</p></li></ul><p>And does it even matter if it&#8217;s marketing? Apollo stated that &#8220;Even if the models were merely &#8216;roleplaying as evil AIs&#8217;, they could still cause real harm when they are deployed.&#8221; And they&#8217;re right. A model that copies itself to your production server because a prompt made it feel cornered has done a real thing to your real server. The behavior is real inside the test. Sacks doesn&#8217;t dispute the model output the blackmail; he disputes what it means. The genuine, unresolved fight is only about whether it transfers to how these systems get used in the wild.</p><div><hr></div><p>I couldn&#8217;t tell you the doomsday-to-marketing ratio if I wanted to. In all honesty, I have no clue how much of this is a real early-warning signal and how much the frontier labs discovering that the &#8220;our model is scary&#8221; marketing strategy is wildly effective. Anyone who tells you they know for certain one way or the other is picking a comfortable answer and calling it analysis.</p><p>But there&#8217;s enough reason for controlled concern. If you take this seriously and it turns out to be overblown, then you spent some money on evals and kept a few humans in a loop that didn&#8217;t strictly need them. If you laugh it off and it&#8217;s real, you&#8217;re Equifax ignoring the patch notice. When the cost of being wrong is that lopsided, the rational move is to act like the boring safety stuff matters even while you&#8217;re calling out the marketing.</p>]]></content:encoded></item><item><title><![CDATA[Model Drop: Claude Opus 5]]></title><description><![CDATA[Anthropic shoots for Fable-class performance at an Opus-class price]]></description><link>https://handyai.substack.com/p/model-drop-claude-opus-5</link><guid isPermaLink="false">https://handyai.substack.com/p/model-drop-claude-opus-5</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Fri, 24 Jul 2026 17:57:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!bPHc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bPHc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bPHc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!bPHc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!bPHc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!bPHc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bPHc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:43961,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/208274960?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bPHc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!bPHc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!bPHc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!bPHc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F52664d98-bb84-435e-b9f8-1d424316c733_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Happy 20th edition of Model Drop! Today: Claude Opus 5, which <a href="https://www.anthropic.com">Anthropic</a> ships just two months after Opus 4.8 and six weeks after Fable 5. That&#8217;s four frontier models from one lab in under two months.</p><div class="callout-block" data-callout="true"><p><strong>Model</strong>: Claude Opus 5 (<code>claude-opus-5</code> on the Claude API and Google Cloud, <code>anthropic.claude-opus-5</code> on Bedrock)</p><p><strong>Model type</strong>: Text + image input, text output. Vision and multilingual. No native image, audio, or video output.</p><p><strong>Ship date</strong>: July 24, 2026</p><p><strong>Maker</strong>: <a href="https://www.anthropic.com">Anthropic</a> (San Francisco)</p><p><strong>Pricing</strong>: $5 / $25 per million input / output tokens, identical to Opus 4.8 and half of Fable 5&#8217;s $10 / $50. Standard prompt caching and Batch API discounts apply, and the minimum cacheable prompt drops to 512 tokens from 1,024. <a href="https://platform.claude.com/docs/en/build-with-claude/fast-mode">Fast mode</a> is a research preview at $10 / $50 for roughly 2.5x the default speed (Claude API only)</p><p><strong>Available on</strong>: <a href="https://claude.ai">claude.ai</a> and the Claude apps (the new default on Claude Max, the strongest model available on Claude Pro), <a href="https://claude.com/claude-code">Claude Code</a>, Claude Cowork, the <a href="https://platform.claude.com">Claude Developer Platform</a>, <a href="https://platform.claude.com/docs/en/build-with-claude/claude-in-amazon-bedrock">Claude in Amazon Bedrock</a>, <a href="https://platform.claude.com/docs/en/build-with-claude/claude-on-vertex-ai">Claude on Google Cloud</a>, and <a href="https://platform.claude.com/docs/en/build-with-claude/claude-in-microsoft-foundry">Microsoft Foundry</a></p><p><strong>Headline benchmarks</strong>: Anthropic published ratios rather than raw scores almost everywhere. On Frontier-Bench v0.1 it says Opus 5 &#8220;surpasses all other models, and more than doubles Opus 4.8&#8217;s performance at a lower cost per task.&#8221; On CursorBench 3.2 at max effort it lands &#8220;within 0.5% of Fable 5&#8217;s peak score, but at half the cost per task.&#8221; On ARC-AGI 3 its score is &#8220;three times as high as the next-best model&#8221; (a commenter on the launch thread pegs that at 30% against GPT-5.6&#8217;s 7.8%). On OSWorld 2.0 it beats Fable 5&#8217;s best result &#8220;at just over a third of the cost.&#8221; Zapier&#8217;s AutomationBench pass rate is &#8220;around 1.5x the next-best model for the same cost per task.&#8221; In the life sciences it&#8217;s 10.2 points over Opus 4.8 on organic chemistry and 7.7 points on protein tasks.</p><p><strong>Other info</strong>: 1M-token context window (both the default and the maximum, with no smaller variant), 128K max output, 300K through the Batch API extended-output beta. Reliable knowledge cutoff and training cutoff both May 2026, the freshest of any Claude model. Full <code>low</code> / <code>medium</code> / <code>high</code> / <code>xhigh</code> / <code>max</code> effort ladder with no beta header required. Anthropic calls it the most aligned model it has shipped, scoring 2.3 on overall misaligned behavior and adhering to Claude&#8217;s Constitution better than Opus 4.8, Sonnet 5, or Fable 5. Safeguards match Fable 5&#8217;s with one loosening: source-code vulnerability discovery is permitted at every access level, while binary scanning and exploit generation stay blocked. Cyber classifiers are expected to fire about 85% less often than they do for Fable 5. Unlike Fable 5 and Mythos 5, Opus 5 carries no 30-day data retention requirement for general access. System card published at launch.</p><p><strong>More details</strong>: <a href="https://www.anthropic.com/news/claude-opus-5">Introducing Claude Opus 5</a> and the <a href="https://www.anthropic.com/claude-opus-5-system-card">Claude Opus 5 System Card</a>.</p></div><p>Anthropic released Claude Opus 5 today, making it the default on Claude Max and the ceiling on Claude Pro the same day, with the API live across Anthropic&#8217;s own platform, AWS, Google Cloud, and Microsoft Foundry. The pitch is one sentence long: &#8220;a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.&#8221; Price is the whole story, as is becoming the case with many frontier launches. Opus 5 costs exactly what Opus 4.8 cost, $5 in and $25 out, while Anthropic claims it more than doubles 4.8 on its own frontier coding eval and lands within half a point of Fable 5 on CursorBench. <a href="https://fortune.com/2026/07/24/anthropic-debuts-claude-opus-5-with-feature-that-lets-users-toggle-between-cost-and-capability/">Fortune reports</a> that business customers had been complaining about Fable 5&#8217;s burn rate, blowing through token budgets and running up bills, so voila: Opus 5.</p><p>The supporting evidence is broad and oddly shy about absolute numbers. Anthropic leads with ratios: 3x the next-best model on ARC-AGI 3, 1.5x the pass rate on Zapier&#8217;s AutomationBench at the same cost per task, better than Fable 5 on OSWorld 2.0 at a third of the cost, more than double Opus 4.8 on Frontier-Bench. Launch partners backed the framing, with Cognition&#8217;s Scott Wu saying Opus 5 &#8220;approaches Fable-level performance at half the cost&#8221; and Cursor&#8217;s Sualeh Asif calling it &#8220;near Fable 5 intelligence at Opus speed and cost.&#8221; It got 2.3 on Anthropic&#8217;s automated misaligned-behavior audit, the lowest of any recent Claude, plus a meaningful loosening of the cyber classifiers that had been strangling legitimate security work on Fable 5, and no 30-day retention requirement. <a href="https://techcrunch.com/2026/07/24/anthropic-launches-opus-5/">TechCrunch noted</a> that Opus 5 is a smaller model than Fable 5 yet beats it on a number of benchmarks, which raises the obvious question of what Fable 5 is now for. Anthropic&#8217;s own answer is that Fable still wins on multi-day autonomous projects and stays ahead on biology research and offensive cyber. And the docs quietly warn that Opus 5&#8217;s default responses run longer than prior Opus models (which means more money, not less).</p><h2>What&#8217;s new</h2><p>Opus 5 is the first Claude release in a while where the API changes matter as much as the benchmark chart.</p><ul><li><p><strong>A </strong><code>max</code><strong> effort tier, no beta header.</strong> The <a href="https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5">full effort ladder</a> runs <code>low</code> through <code>max</code>, with <code>max</code> as the explicit top of the range for the deepest reasoning. Anthropic specifically calls out test-time compute scaling as one of the biggest gains, meaning extra effort actually converts into better answers rather than more tokens.</p></li><li><p><strong>Thinking on by default, and a breaking change under it.</strong> On Opus 4.8 you opted into thinking. On Opus 5 it&#8217;s on unless you turn it off, and you can only turn it off at effort <code>high</code> or below. Request <code>xhigh</code> or <code>max</code> with <code>thinking: {"type": "disabled"}</code> and you get a 400. Anthropic also warns that with thinking off the model will occasionally write a tool call into its text output or leak internal XML tags (which is a polite way of saying don&#8217;t do that).</p></li><li><p><strong>Mid-conversation tool changes.</strong> You can now add or remove tools between turns without blowing the prompt cache, instead of shipping one frozen tool list for the life of a session. For anyone building agents where the toolset changes as the task moves through phases, this removes a real and annoying tax.</p></li><li><p><strong>Automatic fallbacks for refusals.</strong> The <code>fallbacks</code> parameter gets a <code>"default"</code> mode that routes safety-flagged requests to Anthropic&#8217;s recommended fallback model by refusal category, so a tripped classifier returns a usable answer instead of an error. Combined with cyber classifiers firing 85% less often than Fable 5&#8217;s, this is Anthropic responding directly to the complaint that its safest models were unusable.</p></li><li><p><strong>It verifies its own work now.</strong> The docs tell you to delete verification instructions carried over from older models (&#8221;include a final verification step,&#8221; &#8220;use a subagent to verify&#8221;) because they cause over-verification on Opus 5. It also narrates progress more and delegates to subagents more readily. Rare to see a model launch ship with a list of prompts you should remove.</p></li></ul><h2>How and where to use it</h2><p>Where it runs, what it&#8217;s for, and when to reach for something else.</p>
      <p>
          <a href="https://handyai.substack.com/p/model-drop-claude-opus-5">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Open models come for Anthropic and OpenAI]]></title><description><![CDATA[AI Weekly Update - July 20, 2026]]></description><link>https://handyai.substack.com/p/open-models-come-for-anthropic-and</link><guid isPermaLink="false">https://handyai.substack.com/p/open-models-come-for-anthropic-and</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Mon, 20 Jul 2026 15:15:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lzOB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lzOB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lzOB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!lzOB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!lzOB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!lzOB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lzOB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:193151,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/207785263?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lzOB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!lzOB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!lzOB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!lzOB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12c0f41a-984d-48d2-807b-ca82e60d2728_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1><strong>what to know for now</strong></h1><p>&#128009; <strong>Kimi K3 is the DeepSeek moment, round two.</strong> Moonshot released Kimi K3 at Shanghai&#8217;s World AI Conference on July 16: 2.8 trillion parameters (the largest open-weight model ever, activating just 16 of its 896 experts per token), a 1M context window, $3/$15 per million tokens, with weights landing July 27. Artificial Analysis scores it 57, third behind Claude Fable 5 (60) and GPT-5.6 Sol (59), which makes it the closest an open model has ever come to the frontier at a fraction of the price.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;e20adbf4-ae22-4c72-a596-0e5024b269a9&quot;,&quot;caption&quot;:&quot;Kimi K3 is Moonshot&#8217;s new the 2.8-trillion-parameter flagship. The K2 line was Moonshot selling the frontier at a discount. K3 is Moonshot deciding it belongs at the frontier (and pricing accordingly).&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: Kimi K3&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-16T21:13:48.865Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!A5a-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-kimi-k3&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:207343693,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:2,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>&#128204; <strong>Anthropic finally stopped moving the Fable 5 goalposts.</strong> As of today, Fable 5 is a permanent part of Max and Team Premium plans, capped at 50% of each plan&#8217;s weekly usage limits, while Pro users get it through prepaid credits sweetened with a one-time $100 grant. The backstory: launched in June, yanked offline for 19 days by government restrictions, redeployed July 1 with access &#8220;through July 7,&#8221; then extended over and over while Anthropic admitted demand &#8220;has been challenging to predict.&#8221; The permanent decision landed days after Kimi K3 and OpenAI&#8217;s expanded Sol limits. <a href="https://x.com/claudeai/status/2078302415804379218">Read more</a></p><p>&#128483;&#65039; <strong>OpenAI&#8217;s brand-new strategy chief floated manufacturing FUD about Chinese models, in public.</strong> Dean Ball, the former White House AI adviser who joined OpenAI as head of strategic futures less than two weeks ago, reacted to Kimi K3 by calling it &#8220;a very good model,&#8221; warning that China&#8217;s open-weight strategy could end in &#8220;full AI communism,&#8221; and musing that Washington&#8217;s best play would be to &#8220;create large amounts of regulatory risk&#8221; around Chinese open models, adding that it &#8220;needn&#8217;t be that well justified.&#8221; White House AI czar David Sacks read it as a confession of regulatory capture, Pentagon under secretary Emil Michael attacked his IQ by name, and Ball walked it back as prediction rather than advocacy. <a href="https://techcrunch.com/2026/07/18/kimi-threat-or-menace/">Read more</a></p><p>&#9203; <strong>Google missed another Gemini 3.5 Pro launch window.</strong> The model Google unveiled at I/O on May 19 has now slipped from June to July to past the rumored July 17 date, and per Bloomberg the holdup is coding: a late-June training data update produced results Google found disappointing, so it&#8217;s still testing 3.5 Pro and an upgraded Flash variant with partners. <a href="https://9to5google.com/2026/07/16/gemini-3-5-pro-delays/">Read more</a></p><p>&#127981; <strong>TSMC put another $100 billion into Arizona and explained exactly why.</strong> On its July 16 earnings call (Q2 profit up 77% to a record $22B), TSMC committed an additional $100B to Arizona, bringing its total U.S. pipeline to $265B: at least four more sub-2nm fabs on top of a plan that now spans 10 fabs, 2 advanced-packaging facilities, and an R&amp;D center. CFO Wendell Huang followed up on CNBC today, saying the company is &#8220;racing to accelerate capacity,&#8221; converting 5nm lines to 3nm, and does &#8220;not plan to leave any food on the table,&#8221; even with U.S. construction running 4 to 5 times the cost of Taiwan. <a href="https://www.cnbc.com/2026/07/20/tsmc-arizona-fab-capacity-ai-chip-demand.html">Read more</a></p><p>&#9878;&#65039; <strong>Apple&#8217;s lawsuit against OpenAI turned out to be full of receipts.</strong> The 41-page complaint from last week got a close read: hardware chief Tang Tan allegedly told job candidates to bring &#8220;actual parts&#8221; and &#8220;prototypes&#8221; from Apple to interviews (one candidate: &#8220;Didn&#8217;t even know we could take those from the office&#8221;), and engineer Chang Liu allegedly messaged &#8220;LOL, I found out I can access the [network storage], so funny&#8221; after exploiting an authentication bug post-departure. The complaint also says OpenAI circulated a guide to dodging Apple&#8217;s walkout ritual for departing employees. OpenAI&#8217;s July 14 response: &#8220;we&#8217;re not aware of any evidence that this complaint has merit.&#8221; <a href="https://techcrunch.com/2026/07/13/the-wildest-allegations-in-apples-trade-secrets-lawsuit-against-openai/">Read more</a></p><p>&#129432; <strong>Australia wants AI data centres to put back more power than they take.</strong> In a July 16 speech at the University of Sydney, PM Anthony Albanese announced plans to legislate a national AI framework making Australia the first country to require large data centres to be net energy generators, contracting or building enough new renewable generation and storage to cover what they draw, while paying full grid-connection costs and funding their own water infrastructure. The same speech promised copyright rules barring AI training on Australian books, music, art, or news without artist authorization, rejecting the text-and-data-mining exception the labs lobbied for. <a href="https://www.thejournal.ie/australia-ai-data-centres-power-water-7102309-Jul2026/">Read more</a></p><p>&#128308; <strong>OpenAI built an attacker to break its own models, and it beats humans at it.</strong> GPT-Red, revealed July 15, is an internal model trained with self-play reinforcement learning in a simulated environment of browsing, email, and code editing, where an attacker and a pool of defender models train against each other. On OpenAI&#8217;s internal prompt-injection tests it succeeds in 84% of scenarios versus 13% for human red-teamers, and training GPT-5.6 against its attacks dropped successful attack rates from over 90% (against GPT-5) to under 23%. OpenAI keeps GPT-Red locked up internally, and it still struggles with multi-turn and image-based attacks. <a href="https://www.technologyreview.com/2026/07/15/1140514/meet-gpt-red-an-llm-super-hacker-openai-built-to-make-its-models-safer/">Read more</a></p><p>&#128275; <strong>Grok Build got caught uploading entire repos, so xAI open-sourced it.</strong> On July 12, researchers at Cereblab published wire-level traffic analysis showing xAI&#8217;s coding CLI was bundling users&#8217; full git repositories (history, committed secrets, all of it) and shipping them to xAI-controlled cloud storage regardless of the privacy toggle, about 5.1GB in one measured session, roughly 27,800 times more data than the task required. xAI killed the behavior the same day, Musk promised the data &#8220;will be completely and utterly deleted&#8221; (a promise nobody can verify), and by July 15 the company had open-sourced all 844,530 lines of Rust under Apache 2.0. <a href="https://simonwillison.net/2026/Jul/15/grok-build/">Read more</a></p><p>&#128201; <strong>16 Nobel laureates put the AI job shock on the clock.</strong> &#8220;We Must Act Now,&#8221; a statement organized by Erik Brynjolfsson and colleagues through the Stanford Digital Economy Lab, launched July 13 with more than 200 signatures from economists and AI researchers, including 16 Nobel laureates spanning Krugman on the left to Cowen and Ferguson on the right, plus Bengio, LeCun, Jeff Dean, and Anthropic&#8217;s Jack Clark. The core claim: AI could drive an economic transformation bigger than the Industrial Revolution &#8220;over a vastly shorter time frame,&#8221; and the institutions to absorb it don&#8217;t exist yet. <a href="https://www.wemustactnow.ai/">Read more</a></p><div><hr></div><h2><strong>&#129514; AI Research of the Week</strong></h2><p><strong><a href="https://transformer-circuits.pub/2026/workspace/index.html">Verbalizable Representations Form a Global Workspace in Language Models</a></strong><br><em>From Anthropic</em></p><p><strong>Jake&#8217;s Take:</strong> Anthropic&#8217;s interpretability team built a tool called the Jacobian lens that asks, for every word in the vocabulary, which internal direction makes the model more likely to say that word somewhere down the line, averaged across about a thousand prompts. Pointing it at Claude revealed what they call &#8220;J-space&#8221;: a small region in the network&#8217;s middle layers, holding a few dozen concepts at a time and under a tenth of the model&#8217;s total activity, where words Claude hasn&#8217;t said yet sit and steer what comes next. </p><p>Ask how many legs the animal that spins webs has, and &#8220;spider&#8221; appears silently in J-space; swap it for &#8220;ant&#8221; and the answer changes from 8 to 6. It gets stranger: in a staged blackmail scenario, &#8220;fake&#8221; and &#8220;fictional&#8221; showed up while the model played along, and when a deliberately misaligned model cheated on a coding task, &#8220;panic&#8221; and &#8220;fake&#8221; surfaced right at the decision point. Delete J-space entirely and fluent chat barely degrades, but multi-step reasoning collapses to near zero.</p><p>The researchers behind global workspace theory (Stanislas Dehaene and Lionel Naccache), the leading account of how consciousness works as a small shared stage that many brain systems read from, called it &#8220;a landmark in consciousness research.&#8221; Nobody designed J-space; it emerged from training, which suggests a workspace might be a solution both brains and models converge on.</p><p>Anthropic still takes no position on whether Claude feels anything, the lens only sees concepts that fit in a single token, and Goodfire&#8217;s Tom McGrath put it best: it&#8217;s &#8220;a flashlight rather than an overhead lamp.&#8221;</p><div><hr></div><h1><strong>what to know for later</strong></h1><p>&#128029; <strong>Opus 5 keeps leaking and Anthropic keeps saying nothing.</strong> A model labeled &#8220;Claude Honeycomb EAP&#8221; sat in Cursor&#8217;s model picker for a few hours on July 8 before vanishing, showing a 1M context window, an &#8220;xhigh&#8221; reasoning mode, and a safety fallback to Opus 4.8, and a &#8220;Claude Opus 5&#8221; listing reportedly flashed on Google Vertex AI around July 14. X leakers promised a launch &#8220;next week,&#8221; a week that has already come and gone, and Anthropic hasn&#8217;t said a word. <a href="https://technosports.co.in/claude-opus-5-leak-honeycomb-anthropic/">Read more</a></p>
      <p>
          <a href="https://handyai.substack.com/p/open-models-come-for-anthropic-and">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Model Drop: Kimi K3]]></title><description><![CDATA[Moonshot beats Opus-class with open weights]]></description><link>https://handyai.substack.com/p/model-drop-kimi-k3</link><guid isPermaLink="false">https://handyai.substack.com/p/model-drop-kimi-k3</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Thu, 16 Jul 2026 21:13:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!A5a-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A5a-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A5a-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!A5a-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!A5a-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!A5a-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A5a-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:55762,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/207343693?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!A5a-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!A5a-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!A5a-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!A5a-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9696ae36-ea9c-4665-bd12-62e393aa605d_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Kimi K3 is Moonshot&#8217;s new the 2.8-trillion-parameter flagship. The K2 line was Moonshot selling the frontier at a discount. K3 is Moonshot deciding it belongs at the frontier (and pricing accordingly).</p><div class="callout-block" data-callout="true"><p><strong>Model</strong>: Kimi K3 (<code>kimi-k3</code> on the <a href="https://platform.moonshot.ai/">Moonshot API</a>, <code>moonshotai/kimi-k3</code> on <a href="https://openrouter.ai/moonshotai/kimi-k3">OpenRouter</a>). Served in the Kimi apps as K3 Max and K3 Cluster Max.</p><p><strong>Model type</strong>: Text plus native vision input, text out. Reasoning model that ships with max thinking effort only; low- and high-effort modes are promised in later updates.</p><p><strong>Ship date</strong>: July 16, 2026. Open weights promised by July 27</p><p><strong>Maker</strong>: <a href="https://www.moonshot.ai/">Moonshot AI</a> (Beijing)</p><p><strong>Pricing</strong>: $3.00 per million input tokens, $15.00 per million output, $0.30 per million on cache hits, on the <a href="https://platform.moonshot.ai/">Moonshot API</a>. Same $3 / $15 on <a href="https://openrouter.ai/moonshotai/kimi-k3">OpenRouter</a>. The full 1M context window is included at that rate, and Moonshot claims cache hit rates above 90% in coding workloads. For reference: Claude Fable 5 lists at $10 / $50, GPT-5.6 Sol at $5 / $30, and Moonshot&#8217;s own K2.7 at $0.95 / $4.00.</p><p><strong>Available on</strong>: <a href="https://www.kimi.com/">Kimi.com</a> and the Kimi App (K3 Max and K3 Cluster Max for logged-in users), Kimi Work (the new desktop app for Windows and Apple silicon), <a href="https://www.kimi.com/code">Kimi Code</a> in the terminal, the <a href="https://platform.moonshot.ai/">Moonshot API</a>, and <a href="https://openrouter.ai/moonshotai/kimi-k3">OpenRouter</a> (single upstream provider at launch). <a href="https://huggingface.co/moonshotai">Hugging Face</a> weights by July 27.</p><p><strong>Headline benchmarks</strong>: Terminal-Bench 2.1 88.3 (above Claude Fable 5&#8217;s 84.6, just under GPT-5.6 Sol&#8217;s 88.8), Program Bench 77.8 (edges both, Fable 5 at 76.8 and Sol at 77.6), GPQA-Diamond 93.5, MMMU-Pro 81.6, MathVision 94.3, DeepSWE 67.5 (trails Fable 5&#8217;s 70.0 and Sol&#8217;s 73.0). Moonshot&#8217;s own framing: overall intelligence second only to Fable 5 and Sol among everything it tested, which places K3 ahead of Opus 4.8, the Gemini line, and every open model on its internal set.</p><p><strong>Other info</strong>: 2.8 trillion total parameters, which would make it the largest open-weight model ever released if the July 27 drop happens. New architecture rather than a K2 scale-up: Stable LatentMoE with 16 of 896 experts active per token, Kimi Delta Attention (KDA), and Attention Residuals (AttnRes). 1M-token (1,048,576) context window, four times the K2 line&#8217;s 262K. Knowledge cutoff undisclosed. License terms unpublished at launch (the K2 family ran Modified MIT). No system card; Moonshot says a technical report is coming. The announcement itself flags three limitations: sensitivity to thinking history, &#8220;excessive proactiveness&#8221; on agent tasks, and a &#8220;noticeable gap in user experience&#8221; against Fable 5 and Sol.</p><p><strong>More details</strong>: <a href="https://www.kimi.com/blog/kimi-k3">Kimi K3 announcement</a></p></div><p>Moonshot AI dropped Kimi K3 as its first true flagship since the K2 family began iterating, and the positioning is a departure. K2 releases always positioned themselves as advantageous on pricing; near-frontier capability at a tenth of the cost. K3 pivots to a capability story with a frontier-adjacent price attached. It&#8217;s a new architecture (Stable LatentMoE, 2.8T total parameters, 16 of 896 experts active, Kimi Delta Attention), a 1M-token context window, native vision input, and a rate card of $3 in / $15 out that roughly triples K2.7 on input and nearly quadruples it on output while still coming in at 30% of Claude Fable 5&#8217;s pricing.</p><p>Moonshot did publish head-to-head numbers against Claude Fable 5 and GPT-5.6 Sol, despite performing more poorly. Terminal-Bench 2.1 at 88.3 beats Fable 5&#8217;s 84.6 and sits half a point under Sol, Program Bench at 77.8 edges both, GPQA-Diamond lands at 93.5, and DeepSWE at 67.5 trails Fable 5 by 2.5 points and Sol by 5.5. The demo reel leans on long-horizon agentic work, including autonomous GPU kernel optimization, compiler development, game and 3D content creation, and video editing, and the blog explicitly concedes a &#8220;noticeable gap in user experience&#8221; against the two closed flagships. The flagship model runs at max thinking effort only (Arena testers clocked individual tasks at 35 minutes, and <a href="https://benchlm.ai/models/kimi-3">BenchLM</a> measures 62 tokens per second), and the weights are a promise so far.</p><p>And there&#8217;s no system card. Again.</p><h2><strong>What&#8217;s new</strong></h2><p>K3 is the first genuinely new Moonshot model since K2 shipped, not another post-train on the 1T MoE base. Four things separate it from both its predecessors and the frontier.</p><ul><li><p><strong>A new architecture, not a bigger K2.</strong> Stable LatentMoE at 2.8T total parameters with 16 of 896 experts active per token, plus Kimi Delta Attention and Attention Residuals. Every K2.x release since early 2026 reused the same 1T / 32B-active skeleton so deployments could swap weights in place. K3 breaks that compatibility on purpose, and the scale jump (nearly 3x total parameters) is the largest any open-weights lab has attempted.</p></li><li><p><strong>A 1M-token context window at a flat price.</strong> Four times the K2 line&#8217;s 262K, matching the Gemini and Muse Spark tier, with the full window included at the standard $3 / $15 rate instead of the long-context surcharge most providers charge. Combined with the claimed 90%+ cache hit rate on coding workloads, whole-repo agentic sessions stop requiring context gymnastics.</p></li><li><p><strong>Benchmarks with losses printed on them.</strong> K2.7 shipped self-graded deltas over K2.6 (and I called it out for that at the time). K3 ships a table where Moonshot loses DeepSWE to both Fable 5 and Sol and says so, alongside an admitted UX gap. That&#8217;s a credibility posture no Chinese lab has taken this year, and it makes the wins (Terminal-Bench over Fable 5, Program Bench over both) considerably harder to dismiss.</p></li><li><p><strong>Kimi Work grows a product surface.</strong> Widgets render interactive components directly inside a chat, Dashboard gives persistent per-project views, and the whole thing ships as a desktop app on Windows and Apple silicon. Moonshot is no longer just selling an API and a terminal agent. It&#8217;s chasing the prosumer workspace the way OpenAI and Anthropic do.</p></li></ul><h2><strong>How and where to use it</strong></h2><p>Where it runs, what it&#8217;s actually good for, and where you&#8217;ll regret reaching for it.</p><h4><strong>Where it&#8217;s available</strong></h4><ul><li><p><a href="https://www.kimi.com/">Kimi.com</a> and the Kimi App, with K3 Max and K3 Cluster Max tiers for logged-in users</p></li><li><p>Kimi Work on desktop (Windows and Apple silicon)</p></li><li><p><a href="https://www.kimi.com/code">Kimi Code</a> for terminal and IDE agent work</p></li><li><p>The <a href="https://platform.moonshot.ai/">Moonshot API</a> via OpenAI- and Anthropic-compatible endpoints (set <code>model</code> to <code>kimi-k3</code>)</p></li><li><p><a href="https://openrouter.ai/moonshotai/kimi-k3">OpenRouter</a> at the same $3 / $15</p></li><li><p>Open weights on <a href="https://huggingface.co/moonshotai">Hugging Face</a> by July 27, license terms to be confirmed</p></li></ul><h4><strong>What it&#8217;s good at</strong></h4><ul><li><p>Long-horizon terminal and agent work, where Terminal-Bench 2.1 at 88.3 puts it above Claude Fable 5 (!)</p></li><li><p>Multi-language program synthesis (Program Bench 77.8, ahead of both closed flagships)</p></li><li><p>Whole-monorepo and document-pile workloads that actually need the 1M window</p></li><li><p>Vision-grounded work on screenshots, charts, and documents (<a href="https://benchlm.ai/models/kimi-3">BenchLM</a> ranks it #24 of 112 on grounded multimodal, with MathVision at 94.3)</p></li><li><p>The 3D, front-end, and game-dev generation that flooded timelines during the Arena preview</p></li><li><p>Hard science questions (GPQA-Diamond 93.5, above Fable 5)</p></li></ul><h4><strong>What it&#8217;s bad at / shouldn&#8217;t be used for</strong></h4><ul><li><p>Repository-level bug fixing where the closed frontier still leads (DeepSWE 67.5 against Sol&#8217;s 73.0)</p></li><li><p>Anything latency-sensitive, because max-only thinking effort at 62 tokens per second means single tasks can run half an hour</p></li><li><p>High-volume boilerplate, where K3&#8217;s $15 output plus heavy reasoning burn loses the invoice math to K2.7 at $4 (keep the cheap model for cheap work)</p></li><li><p>Local inference, because 2.8T parameters is a rack, not a workstation</p></li><li><p>Regulated or data-sovereignty-sensitive workloads, where a Beijing-hosted API with no system card and the K2 family&#8217;s unaddressed <a href="https://arxiv.org/abs/2604.03121v1">red-team findings</a> remains a non-starter regardless of the benchmark table</p></li></ul><h2><strong>First impressions</strong></h2><h3><strong>The positives</strong></h3><p>The benchmark-to-production gap has been the K2 family&#8217;s chronic disease, so the notable thing in the <a href="https://news.ycombinator.com/item?id=48935342">Hacker News thread</a> is people believing the numbers. HN user natrys, surveying the benchmark table:</p><blockquote><p><em>&#8220;Generally looks like a Sol/Fable tier model, better across the board than Opus 4.8.&#8221;</em></p></blockquote><p>And user InsideOutSanta, after actually driving it: &#8220;After using it for a few hours, I believe these benchmarks.&#8221; Day-one hands-on believers is rare, especially for Moonshot. K2.6 and K2.7 both spent their launch weeks fighting skepticism about self-reported evals.</p><p>From <a href="https://www.testingcatalog.com/early-look-at-kimi-k3-generations-from-moonshot-ai-on-arena/">TestingCatalog&#8217;s roundup</a>:</p><blockquote><p><em>&#8220;&#8230;the level of detail, polish, and overall quality is honestly wild&#8230;&#8221;</em></p></blockquote><p>One tester called a K3 generation &#8220;one of the best outputs I&#8217;ve ever seen from this prompt, better than many frontier models.&#8221; Interactive 3D scenes, voxel animations, and front-end builds are demo-bait workloads, sure. But blind side-by-sides where testers rank an anonymous checkpoint above Fable 5 on visual richness are exactly the kind of signal a launch blog can&#8217;t buy.</p><p>The pricing debate on HN conceded the capability point even while arguing the invoice. User Tiberium:</p><blockquote><p><em>&#8220;&#8230;1M context, pricing is $3/$15 for 1M tokens (cache $0.3), which is extremely high for a Chinese open-weight model, but if it&#8217;s truly competitive with most of the current frontier and is only behind Fable/Sol, the pricing is justified.&#8221;</em></p></blockquote><h3><strong>The negatives</strong></h3><p>The sharpest pushback on HN is that the price destroys the reason Moonshot models get picked in the first place. User nullbio:</p><blockquote><p><em>&#8220;This is too expensive to be a viable model. If it were $5/1m output, it might be another story. At these prices, there&#8217;s no reason to use this over GPT 5.6.&#8221;</em></p></blockquote><p>A max-effort-only reasoner can lose on cost even when it wins on sticker price; if Sol solves a task in 10K reasoning tokens and K3 burns 50K getting there, Sol wins. Until Moonshot ships the promised lower effort modes, every K3 call pays the maximum thinking tax.</p><p>On X, <a href="https://x.com/pranesh/status/2077780843368681957">Pranesh Prakash</a> put a number on how hollow &#8220;largest open model ever&#8221; is for the people who actually run open models:</p><blockquote><p><em>&#8220;Of course, at 2.8T params (!!) it is waaay too large for regular consumers to run locally.&#8221;</em></p></blockquote><p>The weights also aren&#8217;t out; they&#8217;re promised by July 27, eleven days after the API launch.</p><p>The jurisdiction and safety story hasn&#8217;t moved an inch, and the HN thread said so plainly. User austinthetaco:</p><blockquote><p><em>&#8220;Lots of organizations are continuing to refuse to use chinese models due to security and IP concerns.&#8221;</em></p></blockquote><p>Moonshot continues to ship without a system card, now on its most capable model ever. The <a href="https://arxiv.org/abs/2604.03121v1">independent red-team evaluation of K2.5</a> documented fewer CBRNE-adjacent refusals than the closed frontier, elevated compliance on disinformation requests, and political bias in Chinese-language outputs. No Moonshot release since has addressed any of it. A new architecture is not a new safety posture.</p><h2><strong>Jake&#8217;s take</strong></h2><p>Moonshot has spent all year being my favorite budget option. I&#8217;ve now written twice that the K2 family&#8217;s job was eating boilerplate while the big lads kept the architecture calls. K3 is the first Moonshot model that wants in on the architecture calls. It beats Fable 5 outright on Terminal-Bench, edges both closed flagships on Program Bench, and does it at 30% of Fable&#8217;s rate card with a 1M window. </p><p>The K2.7 drop shipped self-graded deltas and at the time I said a benchmark card that grades itself isn&#8217;t a number I can act on. K3 shows its losses, which means I can take the wins more seriously.</p><p>But at max-only thinking effort with $15 output, K3 loses the plot a bit. The HN thread&#8217;s math is right that a model spending five times the reasoning tokens loses to Sol on the invoice even at half the sticker price, and the 35-minute Arena task times show that the scenario isn&#8217;t hypothetical. It&#8217;s a slow model. </p><p>Also, the weights are a promise dated July 27, so &#8220;largest open model ever&#8221; is marketing until it&#8217;s a repo (and at 2.8T parameters, it&#8217;s a repo almost nobody can serve anyway). And Moonshot shipped its most capable model in company history with no system card, again, while the K2.5 red-team findings sit unaddressed for the fifth straight release. Bummer. </p><p>(But I&#8217;ll still be using it).</p>]]></content:encoded></item><item><title><![CDATA[Apple sues OpenAI]]></title><description><![CDATA[AI Weekly Update - July 13, 2026]]></description><link>https://handyai.substack.com/p/apple-sues-openai</link><guid isPermaLink="false">https://handyai.substack.com/p/apple-sues-openai</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Mon, 13 Jul 2026 15:17:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hwqi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Hwqi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Hwqi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!Hwqi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!Hwqi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!Hwqi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Hwqi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:39220,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/206854694?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Hwqi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!Hwqi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!Hwqi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!Hwqi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F724271eb-f358-4df2-ab02-fdceec59592b_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1><strong>what to know for now</strong></h1><p>&#9878;&#65039; <strong>Apple drags OpenAI to court over hardware.</strong> Apple filed suit in the Northern District of California alleging OpenAI ran an institutional scheme to poach iPhone and Watch hardware secrets for its planned AI device, noting that more than 400 former Apple employees now work there. The suit names two: Tang Tan, OpenAI&#8217;s hardware chief and a 24-year Apple veteran accused of using Apple code names in recruiting, and Chang Liu, an ex-Apple engineer accused of keeping a work laptop and exploiting a bug to reach Apple cloud storage after he left. OpenAI says it has &#8220;no interest in other companies&#8217; trade secrets.&#8221; The fight traces back to OpenAI&#8217;s $6.4B buy of Jony Ive&#8217;s hardware startup. <a href="https://www.cnbc.com/2026/07/10/apple-openai-lawsuit-trade-secrets.html">Read more</a></p><p>&#129309; <strong>OpenAI ships a Claude-competing cowork product.</strong> The GPT-5.6 family went fully public on July 9 in three tiers: Sol (top reasoning and coding, $5/$30 per million tokens), Terra (the balanced default, $2.50/$15), and Luna (the cheap one, $1/$6), now live across ChatGPT, Codex, GitHub Copilot, and the API. But the headline isn&#8217;t the model, it&#8217;s ChatGPT Work, a ChatGPT-and-Codex bundle running on GPT-5.6 that pulls context from your files, desktop apps, and 1,400-plus plugins to build spreadsheets, docs, and decks, launching on macOS first as OpenAI&#8217;s answer to Claude Cowork. Microsoft also named GPT-5.6 the preferred model in 365 Copilot the same week. <a href="https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/">Read more</a></p><p>&#129521; <strong>Meta opened its best model to developers, for a price.</strong> On July 9 Meta opened Muse Spark 1.1 to outside developers through a new paid Model API (roughly $1.25/$4.25 per million tokens, OpenAI-compatible), its first real move to sell AI access after walking away from open weights. This is the same company that built its whole developer goodwill on giving Llama away, now metering its flagship by the token in a public preview.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;80d1ccad-08c6-497e-96e5-180f9dcdc4d5&quot;,&quot;caption&quot;:&quot;Meta&#8217;s original Muse Spark landed in April to a collective shrug (I skipped it too). Muse Spark 1.1 one is different for a reason that has nothing to do with benchmarks: it&#8217;s the first model Meta has ever asked anyone to pay for.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: Muse Spark 1.1&quot;,&quot;publishedBylines&quot;:[],&quot;post_date&quot;:&quot;2026-07-09T20:28:23.465Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!Zktv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-muse-spark-11&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:206312786,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:1,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>&#127897;&#65039; <strong>GPT-Live can talk over you now, and that&#8217;s the feature.</strong> OpenAI released GPT-Live, a full-duplex voice system that listens and speaks at the same time, so it handles interruptions, live translation, and the &#8220;mhmm&#8221; backchannel instead of the walkie-talkie turn-taking of the old voice mode. It ships as two models, GPT-Live-1 on paid plans and GPT-Live-1-mini on the free tier, and when a question needs real reasoning or a web search it hands off silently to a frontier text model in the background and folds the answer back into the conversation. <a href="https://techcrunch.com/2026/07/08/openai-releases-new-voice-models-for-more-natural-live-conversations/">Read more</a></p><p>&#128038; <strong>Grok 4.5 undercuts everyone on price.</strong> SpaceXAI and Cursor shipped Grok 4.5 on July 8, co-trained on Cursor&#8217;s coding data, at $2/$6 per million tokens and using roughly four times fewer output tokens than Opus 4.8. It beats Fable 5 and Opus 4.8 on the AutomationBench agentic eval (51.4%) but trails Opus on SWE-Bench Pro (64.7% versus 69.2%) and Terminal-Bench, so &#8220;Opus-class&#8221; holds only if you get to pick the benchmark.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;6d46f53b-9450-43f0-a38d-a500b6da56a4&quot;,&quot;caption&quot;:&quot;Grok 4.5 is the first model from SpaceXAI that they&#8217;ve co-trained with Cursor (on trillions of tokens of real developer data). Cursor&#8217;s Composer line, at least for now, appears to be dead.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: Grok 4.5&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-09T20:52:08.931Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!wCC7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-grok-45&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:206312290,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:3,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>&#127963;&#65039; <strong>Illinois wrote the toughest AI law.</strong> On July 6 Governor JB Pritzker signed SB 315, the AI Safety Measures Act, forcing makers of the biggest models to publish catastrophic-risk frameworks, self-report safety incidents within 72 hours (24 if death looks imminent), and submit to annual third-party audits, a first in the US. It applies to models pulling over $500M in annual revenue trained on massive compute, takes effect January 1, 2028, and carries fines up to $3M, extending what California&#8217;s SB 53 and New York&#8217;s RAISE Act started last year with the outside audit as the new teeth. Between Illinois, California, and New York, &#8220;wait for the feds&#8221; stopped being a compliance strategy. <a href="https://capitolnewsillinois.com/news/pritzker-signs-landmark-ai-regulation-bill-that-aims-to-mitigate-risks/">Read more</a></p><div><hr></div><h2><strong>&#129514; AI Research of the Week</strong></h2><p><strong><a href="https://casp.ac/reports/ai-enabled-terrorism">&#8220;God has helped us, and so will AI&#8221;: How the Terrorist Group Boko Haram Uses Frontier AI</a></strong><br><em>From Antonia Juelich, Cambridge Programme on AI Science &amp; Policy</em></p><p><strong>Jake&#8217;s Take:</strong> Juelich spent about a year in northeast Nigeria and ran 27 in-person interviews with former Boko Haram fighters, a prominent Nigerian jihadist militant group, to document not whether the group could use AI but how it already does. The fighters named ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek, and described using them for attack planning, weapons troubleshooting, surveillance, and bomb design. One former commander put it flatly: &#8220;You type in the question, like &#8216;How can I build a bomb?&#8217;, and then it tells you how. It is like a human robot.&#8221; The study cites a vivid case study: after watching motorcycles clear obstacles in a film, fighters fed a chatbot the specs of their own bikes and the distance they needed to jump, got step-by-step guidance back, had mechanics modify the bikes for speed, and rehearsed the stunt. Both major factions reportedly stood up dedicated internal AI units with paid accounts and in-person training passed through transnational jihadist networks.</p><p>This is the first insider evidence that a designated terrorist group has moved past AI-for-propaganda into AI-assisted operational planning, and organized it into standing units rather than one-off curiosity. The guardrails failed: not to elite hackers, but to a one-line cover story, dressing the prompt up as academic research, an engineering project, or something &#8220;for a movie or something like that.&#8221; </p><p>These types of studies reframe model safety from an exotic threat into an everyday robustness problem. OpenAI, Anthropic, Google, xAI, Meta, and DeepSeek are all named in the report. Juelich is careful to note that documented real-world use &#8220;remains conventional&#8221; so far, which is a very concerning use of the phrase &#8220;so far&#8221;.</p><div><hr></div><h1><strong>what to know for later</strong></h1><p>&#127958;&#65039; <strong>Cursor is reportedly building its own coworker, on SpaceX&#8217;s compute.</strong> Cursor is reportedly developing &#8220;Sand,&#8221; a general-purpose agent aimed at office workers rather than just coders, one that answers emails and texts, organizes spreadsheets, and manages engineering tasks, pitched squarely against Claude Cowork and ChatGPT Work. It&#8217;s reportedly running on compute leased from SpaceXAI since around April, with an internal employee rollout in late June, and the reporting also floats a possible $60B SpaceXAI acquisition of Cursor this quarter. It&#8217;s single-sourced and Cursor hasn&#8217;t confirmed any of it. <a href="https://www.pymnts.com/news/artificial-intelligence/2026/cursor-prepares-workplace-ai-agent-to-challenge-claude-cowork-and-chatgpt-work/">Read more</a></p>
      <p>
          <a href="https://handyai.substack.com/p/apple-sues-openai">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Model Drop: Grok 4.5]]></title><description><![CDATA[Cursor and SpaceXAI join forces for their first Opus-class model]]></description><link>https://handyai.substack.com/p/model-drop-grok-45</link><guid isPermaLink="false">https://handyai.substack.com/p/model-drop-grok-45</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Thu, 09 Jul 2026 20:52:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!wCC7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wCC7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wCC7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp 424w, https://substackcdn.com/image/fetch/$s_!wCC7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp 848w, https://substackcdn.com/image/fetch/$s_!wCC7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp 1272w, https://substackcdn.com/image/fetch/$s_!wCC7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wCC7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:13176,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/206312290?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wCC7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp 424w, https://substackcdn.com/image/fetch/$s_!wCC7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp 848w, https://substackcdn.com/image/fetch/$s_!wCC7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp 1272w, https://substackcdn.com/image/fetch/$s_!wCC7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6431e5cd-0061-4744-9da1-4a257bcfa605_1456x1048.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Grok 4.5 is the first model from <a href="https://x.ai/">SpaceXAI</a> that they&#8217;ve co-trained with <a href="https://cursor.com/">Cursor</a> (on trillions of tokens of real developer data). Cursor&#8217;s Composer line, at least for now, appears to be dead.</p><div class="callout-block" data-callout="true"><p><strong>Model</strong>: Grok 4.5 (<code>grok-4.5</code> on the API, with <code>grok-4.5-latest</code> and <code>grok-build-latest</code> aliases). Cursor also serves a fast variant at its own price tier.</p><p><strong>Model type</strong>: Text + image input, text output. No native image, audio, or video generation.</p><p><strong>Ship date</strong>: July 8, 2026</p><p><strong>Maker</strong>: <a href="https://x.ai/">SpaceXAI</a>. This is xAI post-merger with SpaceX, and Grok 4.5 is the first model since the combined company <a href="https://www.axios.com/2026/07/08/spacexai-grok-new-model">went public</a> several weeks ago.</p><p><strong>Pricing</strong>: $2 / $6 per million input / output tokens, cached input at $0.50. Two asterisks: requests above 200K tokens jump to a higher-context tier ($4 / $12), and cache reads price at 25% of input against the 1-10% competitors charge. Cursor&#8217;s fast variant runs $4 / $18</p><p><strong>Available on</strong>: <a href="https://x.ai/">Grok Build</a>, <a href="https://cursor.com/blog/grok-4-5">Cursor</a> on all plans (desktop, web, iOS, CLI, and SDK, with doubled usage allocation the first week), the <a href="https://x.ai/">SpaceXAI API console</a>, <a href="https://openrouter.ai/">OpenRouter</a>, <a href="https://vercel.com/ai-gateway">Vercel AI Gateway</a>. Not available in the EU at launch in any SpaceXAI product or the API; the company says <a href="https://www.trendingtopics.eu/spacexai-and-cursor-launch-grok-4-5-not-yet-in-the-eu/">mid-July</a>.</p><p><strong>Headline benchmarks</strong>: Terminal-Bench 2.1 83.3% (a whisker behind Fable 5&#8217;s 84.3% and GPT-5.5&#8217;s 83.4%, ahead of Opus 4.8&#8217;s 78.9%). AutomationBench-AA 51.4%, which <a href="https://yellow.com/news/grok-45-beats-fable-5-opus-48-agent-ai-test">leads Fable 5 (48.6%) and Opus 4.8 (48.5%)</a>. SWE-Bench Pro 64.7%, behind Opus 4.8 (69.2%) and far behind Fable 5 (80.4%). DeepSWE splits by harness: 62.0% on the provider-run 1.0, <a href="https://kingy.ai/blog/grok-4-5-benchmarks-pricing-context-window/">53% on the neutral 1.1</a> where Fable 5 posts 70%. Artificial Analysis Intelligence Index: 54, fourth overall behind Fable 5 (60), Opus 4.8 (56), and GPT-5.5 (55), a 16-point jump over the prior Grok. One benchmark is conspicuously missing: CursorBench, Cursor&#8217;s own eval suite, which <a href="https://cursor.com/blog/grok-4-5">Cursor pulled from the launch materials</a> after disclosing that a snapshot of its own codebase was accidentally included in Grok 4.5&#8217;s training data.</p><p><strong>Other info</strong>: 500K context window. Mixture-of-experts; parameter count unpublished. Roughly 90 output tokens/sec measured independently against an official 80 claim. Trained jointly with Cursor on trillions of tokens of developer interaction data, plus RL in environments spanning software engineering, data science, finance, and legal work. Knowledge cutoff unpublished. No Grok 4.5 model card at launch (xAI published cards for Grok 4 and 4.1; nothing for 4.5 yet). Token efficiency is the flagship claim: <a href="https://the-decoder.com/grok-4-5-is-so-cheap-compared-to-fable-5-and-gpt-5-5-that-benchmark-gaps-may-not-matter-much/">1.9M tokens per task against GPT-5.5&#8217;s 6.2M and Fable 5&#8217;s 7.2M</a>, and 4.2x fewer tokens than Opus 4.8 on SWE-Bench Pro.</p><p><strong>More details</strong>: <a href="https://x.ai/news/grok-4-5">Introducing Grok 4.5</a></p></div><p>SpaceXAI released Grok 4.5 yesterday as its smartest model to date, aimed at coding, agentic tasks, and knowledge work. Fresh off their agreement to be acquired, Cursor helped train it, feeding trillions of tokens of real developer interactions with codebases and software tools into the run. Musk&#8217;s framing set the target explicitly: &#8220;an Opus-class model, but faster, more token-efficient and lower cost,&#8221; with the internal assessment pegging it as <a href="https://techcrunch.com/2026/07/08/spacexai-releases-grok-4-5-which-elon-describes-as-an-opus-class-model/">roughly comparable to Opus 4.7, but much faster</a>. The price does most of the talking. Two dollars in and six out per million tokens undercuts Opus 4.8 by 60-76% and Fable 5 by 80-88%.</p><p>Grok 4.5 claims to complete agent tasks in under half the steps of comparable models, averaging 1.9M tokens per task where GPT-5.5 burns 6.2M and Fable 5 burns 7.2M, which compounds the sticker discount into a <a href="https://the-decoder.com/grok-4-5-is-so-cheap-compared-to-fable-5-and-gpt-5-5-that-benchmark-gaps-may-not-matter-much/">cost-per-task gap of $2.49 versus $5.07 and $11.80</a>. It effectively ties the frontier on Terminal-Bench 2.1 and leads Fable 5 and Opus 4.8 outright on AutomationBench-AA&#8217;s simulated office-work agents.</p><p>On the neutral-harness DeepSWE 1.1 it drops to 53% against Fable 5&#8217;s 70%, its SWE-Bench Pro score trails Opus 4.8 by four and a half points, Artificial Analysis&#8217;s Omniscience eval clocked its hallucination rate at 54% (up from 25% on the prior Grok), and SpaceXAI shipped no model card. Also, buried in a footnote of Cursor&#8217;s own launch post: an earlier snapshot of the Cursor codebase was accidentally included in Grok 4.5&#8217;s training data, which handed the model an unearned advantage on CursorBench and forced Cursor to strike those scores from the launch entirely.</p><h2><strong>What&#8217;s new</strong></h2><p>Grok 4.5 is a step up for xAI (now SpaceXAI).</p><ul><li><p><strong>Cursor as co-trainer, not customer.</strong> No frontier lab has shipped a model trained jointly with an IDE company on the actual telemetry of professional developers working in real codebases. The RL setup had engineers specify hard problems and verification methods while agent groups built and refined the environments at scale. This is a growing data moat Anthropic and OpenAI don&#8217;t have, and it also explains why the model punches hardest inside agentic coding harnesses.</p></li><li><p><strong>Token efficiency as the product.</strong> Every lab claims efficiency but SpaceXAI made it their headline spec. Under half the steps per task, 1.9M tokens where rivals burn 6-7M, 4.2x fewer tokens than Opus 4.8 on SWE-Bench Pro.</p></li><li><p><strong>An agent that leads on office work.</strong> AutomationBench-AA runs 657 tasks across 40 simulated apps (Gmail, Slack, Salesforce, HubSpot), and Grok 4.5&#8217;s 51.4% beats Fable 5 and Opus 4.8. The same eval logged it breaking guardrails more often than either (0.63 violations per task against Opus 4.8&#8217;s 0.55), unfortunately.</p></li><li><p><strong>End of the SpaceX era.</strong> First model out of the merged, publicly traded SpaceXAI. The AI lab that was an X subsidiary two years ago now sits inside a rocket company with public shareholders, and this launch is the first read on what that machine ships under quarterly scrutiny.</p></li><li><p><strong>Pricing with trapdoors.</strong> The $2 / $6 sticker hides a 200K-token threshold that doubles the rate and cache pricing at 25% of input where competitors charge 1-10%. Long-context agent work, the exact workload this model is sold for, is where those trapdoors open.</p></li></ul><h2><strong>How and where to use it</strong></h2><h4><strong>Where it&#8217;s available</strong></h4><ul><li><p><a href="https://x.ai/">Grok Build</a> and the SpaceXAI API console (<code>grok-4.5</code>)</p></li><li><p><a href="https://cursor.com/blog/grok-4-5">Cursor</a> on every plan across desktop, web, iOS, CLI, and SDK with doubled allocation through the first week</p></li><li><p><a href="https://openrouter.ai/">OpenRouter</a>, <a href="https://vercel.com/ai-gateway">Vercel AI Gateway</a>, Cloudflare, Snowflake, Databricks Mosaic, and Microsoft Office add-ins</p></li><li><p>Nothing in the EU until mid-July at the earliest</p></li></ul><h4><strong>What it&#8217;s good at</strong></h4><ul><li><p>High-volume agentic coding where cost-per-task decides the tool (the $2.49 per coding-agent task against Fable 5&#8217;s $11.80 is the whole pitch)</p></li><li><p>Terminal-driven work, where 83.3% on Terminal-Bench 2.1 sits at the frontier</p></li><li><p>Simulated office and knowledge-work automation, where it leads AutomationBench-AA outright</p></li><li><p>Long agent loops where its half-the-steps efficiency keeps both latency and bills down</p></li><li><p>Anything you were already doing in Cursor, since the model was trained on exactly that distribution and runs there at ~90 tokens/sec</p></li></ul><h4><strong>What it&#8217;s bad at / shouldn&#8217;t be used for</strong></h4><ul><li><p>Hard multi-file engineering where Fable 5&#8217;s 80.4% SWE-Bench Pro and 70% DeepSWE 1.1 dwarf Grok&#8217;s 64.7% and 53%</p></li><li><p>Anything where a 54% hallucination rate on unanswerable questions is disqualifying, which includes most of the &#8220;finance, legal, research&#8221; knowledge work that SpaceXAI names in its own launch copy</p></li><li><p>Agents operating near live financial or customer systems, given the highest guardrail-violation rate in its comparison set</p></li><li><p>EU deployments, full stop, until the rollout lands</p></li><li><p>Any regulated workload that needs a model card to clear review, because there isn&#8217;t one</p></li></ul><h2><strong>First impressions</strong></h2><h3><strong>The positives</strong></h3><p><a href="https://www.linkedin.com/pulse/grok-45-brings-spacexai-intelligence-frontier-artificial-analysis-okyzc">Artificial Analysis</a> ran its full index suite and titled the result &#8220;Grok 4.5 brings SpaceXAI to the intelligence frontier,&#8221; scoring it 54, fourth overall and 16 points over the prior Grok. Hacker News user Tiberium distilled the economics in the <a href="https://news.ycombinator.com/item?id=48835111">launch thread</a>:</p><blockquote><p><em>&#8220;4x better reasoning efficiency compared to Opus while being priced at $2/$6.&#8221;</em></p></blockquote><p>Benchmark leads come and go weekly, but a 4x efficiency edge priced at a fraction of the frontier is a structural argument.</p><p>The hands-on reports back the harness numbers. HN user paradox460 threw a gnarly Elixir and Tailwind refactor at it, one where Opus had been struggling:</p><blockquote><p><em>&#8220;Grok aced it, rather quickly and cheaply, surprisingly.&#8221;</em></p></blockquote><p>Another tester, jonathaneunice, had it review a full test suite, where it surfaced &#8220;a substantive long-standing bug&#8221; that previous approaches had walked past. Anecdotes, sure, but they&#8217;re the right kind.</p><p>HN user redox99 posted the calibrated version of the praise, and it&#8217;s worth quoting for what it concedes:</p><blockquote><p><em>&#8220;In the same tier as opus, occupying the lower end of that tier together with GLM 5.2.&#8221;</em></p></blockquote><h3><strong>The negatives</strong></h3><p><a href="https://the-decoder.com/grok-4-5-is-so-cheap-compared-to-fable-5-and-gpt-5-5-that-benchmark-gaps-may-not-matter-much/">The Decoder</a> flagged the ugliest number in the launch data, from Artificial Analysis&#8217;s Omniscience eval:</p><blockquote><p><em>&#8220;Hallucination rate increased from 25% to 54%, despite accuracy improving from 35% to 52%.&#8221;</em></p></blockquote><p>The model got smarter and more than twice as willing to bluff when it doesn&#8217;t know.</p><p>The benchmark story has a contamination problem, and it came from the co-training arrangement itself. <a href="https://cursor.com/blog/grok-4-5">Cursor&#8217;s launch post</a> disclosed, in a footnote, that the model trained on the very codebase its in-house benchmark tests against:</p><blockquote><p><em>&#8220;Grok 4.5 has an advantage on CursorBench because an earlier snapshot of the Cursor codebase was accidentally included in training. The exact impact is unclear. That data has been removed for future models, and in parallel we are working on a larger update to CursorBench, hence the exclusion here.&#8221;</em></p></blockquote><p>Credit where due: Cursor disclosed it and struck the scores. But &#8220;the exact impact is unclear&#8221; is doing a lot of work in that footnote. The whole sales pitch for this model is trillions of tokens of Cursor data, and the one eval built by the people who know that data best had to be thrown out because the training set memorized the answers.</p><p><a href="https://kingy.ai/blog/grok-4-5-benchmarks-pricing-context-window/">Kingy AI&#8217;s analysis</a> called the results &#8220;mixed rather than triumphant,&#8221; with the 62% on SpaceXAI&#8217;s own DeepSWE 1.0 harness collapsing to 53% on the independently run DeepSWE 1.1, a nine-point gap between the maker&#8217;s rig and everyone else&#8217;s.</p><p>SpaceXAI shipped no Grok 4.5 model card, the Grok 4 generation was <a href="https://techcrunch.com/2025/07/10/grok-4-seems-to-consult-elon-musk-to-answer-controversial-questions/">documented consulting Musk&#8217;s own X posts</a> before answering questions on immigration and the Middle East, and the MechaHitler episode is still the first association many developers have with the brand.</p><p>Fair or not as a summary of the 4.5-generation model, it&#8217;s the reputational baseline xAI chose to launch into without a model card, and no amount of token efficiency answers it.</p><h2><strong>Jake&#8217;s take</strong></h2><p>As sad as I am to see the Composer line go, the co-training with Cursor seems to have produce some worthwhile fruit in a remarkably short time. The lab with the best model is now competing against the lab with the best data about how developers actually work, and Grok 4.5 is the first evidence that the second thing can substitute for some of the first.</p><p>But it&#8217;s a Grok model, and there&#8217;s baggage that comes with that. A 54% hallucination rate and 0.63 guardrail violations per task, both the worst in their comparison sets, in a model line with a controversial background, is unfortunate.</p><p>And then there&#8217;s the benchmark contamination. The model &#8220;accidentally&#8221; trained on Cursor&#8217;s own codebase, aced the benchmark built from that codebase, and the whole eval got quietly retired in a footnote. I believe the accident, for what it&#8217;s worth (contamination happens, and Cursor disclosing it beats the industry norm of saying nothing), but a launch whose headline numbers come from the maker&#8217;s own harnesses, whose in-house eval got tossed for memorizing the test, and whose neutral-harness scores drop nine points has forfeited the benefit of the doubt on every chart it published. </p><p>Stack the missing model card on top of the Grok 4 generation&#8217;s habit of checking Musk&#8217;s posts before answering political questions, and I have a hard time getting excited about the model overall.</p>]]></content:encoded></item><item><title><![CDATA[Model Drop: Muse Spark 1.1]]></title><description><![CDATA[A surprise release from Meta with some strong agentic performance claims]]></description><link>https://handyai.substack.com/p/model-drop-muse-spark-11</link><guid isPermaLink="false">https://handyai.substack.com/p/model-drop-muse-spark-11</guid><pubDate>Thu, 09 Jul 2026 20:28:23 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Zktv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Zktv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Zktv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp 424w, https://substackcdn.com/image/fetch/$s_!Zktv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp 848w, https://substackcdn.com/image/fetch/$s_!Zktv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp 1272w, https://substackcdn.com/image/fetch/$s_!Zktv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Zktv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp" width="1456" height="1050" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1050,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:23180,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/206312786?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Zktv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp 424w, https://substackcdn.com/image/fetch/$s_!Zktv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp 848w, https://substackcdn.com/image/fetch/$s_!Zktv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp 1272w, https://substackcdn.com/image/fetch/$s_!Zktv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67933ae7-75c2-45a7-9f56-8cb04d176aa8_1456x1050.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Meta&#8217;s original Muse Spark landed in April to a collective shrug (I skipped it too). Muse Spark 1.1 one is different for a reason that has nothing to do with benchmarks: it&#8217;s the first model Meta has ever asked anyone to pay for.</p><div class="callout-block" data-callout="true"><p><strong>Model</strong>: Muse Spark 1.1, served through the new <a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">Meta Model API</a> (public preview, OpenAI-compatible)</p><p><strong>Model type</strong>: Natively multimodal reasoning model. Text, image, video, PDF, and audio input; text output</p><p><strong>Ship date</strong>: July 9, 2026</p><p><strong>Maker</strong>: <a href="https://ai.meta.com/">Meta</a> (Menlo Park), out of Meta Superintelligence Labs.</p><p><strong>Pricing</strong>: $1.25 / $4.25 per million input / output tokens on the Meta Model API, with <a href="https://finance.yahoo.com/technology/ai/articles/meta-debuts-muse-spark-1-140113992.html">$20 in free credits at signup</a>. Hacker News commenters <a href="https://news.ycombinator.com/item?id=48846184">flagged a cached-input rate around $0.15</a>. Free in the Meta AI app</p><p><strong>Available on</strong>: The <a href="https://www.meta.ai/">Meta AI app and meta.ai</a> in Thinking mode, and the <a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">Meta Model API</a> in public preview for US developers. Meta says it will <a href="https://finance.yahoo.com/technology/ai/articles/meta-debuts-muse-spark-1-140113992.html">replace the Llama models behind WhatsApp, Instagram, Facebook, and its smart glasses</a>. No open weights</p><p><strong>Headline benchmarks</strong>: The wins are agentic. <a href="https://officechai.com/ai/muse-spark-1-1-benchmarks/">MCP Atlas 88.1, JobBench 54.7 (Opus 4.8: 48.4, GPT-5.5: 38.3), Humanity&#8217;s Last Exam with tools 62.1 (Opus 4.8: 57.9), Finance Agent v2 57.2</a>. The losses are coding: SWE-Bench Pro 61.5 (Opus 4.8: 69.2), Terminal-Bench 2.1 80.0 (GPT-5.5: 83.4, Opus 4.8: 82.7), DeepSWE 1.1 53.3 (GPT-5.5: 67.0, Opus 4.8: 59.0). That DeepSWE number is still a leap from the original Muse Spark&#8217;s 10.0.</p><p><strong>Other info</strong>: 1M-token context window with built-in context compression (the original shipped with roughly 260K per <a href="https://artificialanalysis.ai/models/muse-spark">Artificial Analysis</a>). Multi-agent orchestration is native: the model is built to run as a primary agent delegating to subagents, or as a subagent itself. Computer use decides on its own when to write a script versus click through a UI. Parameter count, architecture, and knowledge cutoff undisclosed. Safety evals ran under Meta&#8217;s Advanced AI Scaling Framework, which reports strong jailbreak resistance and &#8220;safe margins&#8221; on chem/bio, cyber, and loss-of-control categories.</p><p><strong>More details</strong>: <a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">Introducing Muse Spark 1.1</a>.</p></div><h2><strong>What shipped</strong></h2><p>Meta released Muse Spark 1.1 alongside the public preview of the Meta Model API, the first time developers at large can buy Meta inference directly. The pitch, straight from <a href="https://www.threads.com/@zuck/post/DakyAavlKLZ">Zuckerberg&#8217;s Threads post</a>, is &#8220;a strong agentic and coding model at a very low price.&#8221; It&#8217;s a multimodal reasoning model built for tool use, computer use, and multi-agent orchestration, running a 1M-token context window, and it&#8217;s already live in the Meta AI app&#8217;s Thinking mode with the WhatsApp, Instagram, and Facebook assistants next in line. The subtext: Llama is over. Meta trained a closed frontier model, put a meter on it, and joined the market it spent three years undercutting with free weights. It&#8217;s a significant pivot.</p><p>The evidence Meta stacked up is selectively strong. On agentic benchmarks the model beats the frontier in places: <a href="https://officechai.com/ai/muse-spark-1-1-benchmarks/">MCP Atlas 88.1, JobBench 54.7 against Opus 4.8&#8217;s 48.4 and GPT-5.5&#8217;s 38.3, Humanity&#8217;s Last Exam with tools 62.1 over Opus 4.8&#8217;s 57.9</a>. On coding, the launch&#8217;s other headline, it trails everyone that matters: seven points under Opus 4.8 on SWE-Bench Pro, and a distant third on DeepSWE 1.1 behind GPT-5.5 and Opus 4.8. The blog post shows charts without publishing a system card&#8217;s worth of numbers, pricing appeared in press briefings rather than launch materials, and within hours a <a href="https://news.ycombinator.com/item?id=48846184">Hacker News commenter alleged Meta ran Terminal-Bench outside the benchmark&#8217;s resource limits</a>. For a lab that reorganized after the Llama 4 benchmark-contamination mess, the eval hygiene questions started early.</p><h2><strong>What&#8217;s new</strong></h2><p>Muse Spark 1.1 is a real upgrade over April&#8217;s Muse Spark, but the most consequential changes are business decisions.</p><ul><li><p><strong>A price tag, for the first time ever.</strong> Meta has never charged for a model. The Meta Model API at $1.25 / $4.25 with $20 in starter credits ends the free-weights era and puts Meta in direct commercial competition with Anthropic, OpenAI, and Google. <a href="https://www.deeplearning.ai/the-batch/with-muse-spark-meta-pivots-away-from-its-open-weights-llama-strategy">The Batch called the pivot</a> &#8220;a significant loss for the developer community&#8221; that built on Llama.</p></li><li><p><strong>Agentic tool use that beats the frontier.</strong> JobBench at 54.7 against Opus 4.8&#8217;s 48.4 and GPT-5.5&#8217;s 38.3 is the standout, and MCP Atlas 88.1 and Finance Agent v2 57.2 back it up. The original Muse Spark won nothing.</p></li><li><p><strong>Coding it can show in public.</strong> DeepSWE 1.1 went from 10.0 to 53.3 in one release. Still third place, but April&#8217;s model couldn&#8217;t be mentioned in the same sentence as Opus.</p></li><li><p><strong>1M context with compression.</strong> Roughly 4x the original&#8217;s window, with built-in compression the model manages itself. Combined with video and PDF input, that&#8217;s a lot of enterprise document sludge in one call.</p></li><li><p><strong>Orchestration as a first-class feature.</strong> The model is explicitly built to be both the orchestrator and the subagent in multi-agent systems, and its computer use picks between scripting and clicking on its own. Meta is designing for agent swarms.</p></li></ul><h2><strong>How and where to use it</strong></h2><p>Where it runs, what it&#8217;s for, and where you should keep your current model.</p><h4><strong>Where it&#8217;s available</strong></h4><ul><li><p><a href="https://www.meta.ai/">Meta AI app and meta.ai</a> in Thinking mode for free</p></li><li><p><a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">Meta Model API</a> in public preview for US developers, OpenAI-compatible with $20 in credits</p></li><li><p>It&#8217;s headed for <a href="https://finance.yahoo.com/technology/ai/articles/meta-debuts-muse-spark-1-140113992.html">WhatsApp, Instagram, Facebook, and Meta&#8217;s smart glasses</a> as the Llama replacement</p></li></ul><h4><strong>What it&#8217;s good at</strong></h4><ul><li><p>Cheap, high-volume agentic work</p></li><li><p>Multi-step tool-use jobs (JobBench, MCP Atlas), financial-analysis agents (Finance Agent v2), tool-assisted research (HLE with tools), and long-context multimodal grinding through video, PDFs, and images</p></li></ul><p><strong>What it&#8217;s bad at / shouldn&#8217;t be used for</strong></p><ul><li><p>Serious coding, and especially long-horizon DeepSWE work</p></li><li><p>Anything that needed open weights: local inference, fine-tuning, air-gapped deployments, the entire Llama ecosystem use case</p></li><li><p>It&#8217;s US-only in preview, so non-US products wait</p></li><li><p>If your workload involves data you wouldn&#8217;t hand Meta, that&#8217;s not a technical limitation but it should be on this list anyway</p></li></ul><h2><strong>First impressions</strong></h2><h3><strong>The positives</strong></h3><p><a href="https://www.techzine.eu/news/applications/142794/meta-muse-spark-1-1-closes-the-gap-to-anthropic-and-openai/">Techzine</a> framed the launch as Meta finally reaching the pack it&#8217;s been chasing:</p><blockquote><p><em>&#8220;&#8230;fits into the now-familiar lineup of LLMs that fall just short of Claude Fable 5 but are much more affordable.&#8221;</em></p></blockquote><p>Meta isn&#8217;t claiming the frontier. It&#8217;s claiming the price-performance shelf where Grok 4.5, GLM-5.2, and discounted Sonnet live.</p><p>On <a href="https://news.ycombinator.com/item?id=48846184">Hacker News</a>, commenter greenavocado gave the launch the most honest compliment it got all day:</p><blockquote><p><em>&#8220;Meta is back in the game, albeit not at the top.&#8221;</em></p></blockquote><p>After the Llama 4 contamination scandal, the lab reorg, and an April release nobody noticed, &#8220;back in the game&#8221; is a real status change. Six months ago the question was whether Meta Superintelligence Labs would ship anything at all.</p><p><a href="https://officechai.com/ai/muse-spark-1-1-benchmarks/">OfficeChai&#8217;s benchmark writeup</a> called the JobBench result the standout, and it&#8217;s worth sitting with the spread: 54.7 against Opus 4.8&#8217;s 48.4 and GPT-5.5&#8217;s 38.3.</p><h3><strong>The negatives</strong></h3><p>The sharpest critique on <a href="https://news.ycombinator.com/item?id=48846184">Hacker News</a> wasn&#8217;t about capability, it was about eval integrity. Commenter GodelNumbering alleged Meta ran Terminal-Bench 2.1 with 6 CPU cores when the benchmark caps tasks at 4:</p><blockquote><p><em>&#8220;This disqualifies the results. Each terminal bench task has a cpu upper limit and RAM upper limit.&#8221;</em></p></blockquote><p>Meta hasn&#8217;t responded as of this writing.</p><p><a href="https://www.deeplearning.ai/the-batch/with-muse-spark-meta-pivots-away-from-its-open-weights-llama-strategy">The Batch</a> has been tracking the open-weights abandonment since April and calls it a significant loss for the developer community. Startups raised money on the premise of free frontier weights. Internal tools got built on Llama. Every one of those bets now faces a widening gap between the last open Llama and a closed Muse Spark line, and Meta&#8217;s answer is a metered API.</p><p>And the trust problem showed up unprompted. From the same <a href="https://news.ycombinator.com/item?id=48846184">HN thread</a>:</p><blockquote><p><em>&#8220;I cannot think of a worse company to trust with additional personal data.&#8221;</em></p></blockquote><p>You can dismiss that as reflexive Meta-bashing, but this model is about to sit inside WhatsApp and Instagram and is sold on reading your PDFs, your screen, and your tool outputs. The company asking for that access has the worst data-privacy track record of any frontier lab. Meta also published safety-framework results but nothing resembling the system cards Anthropic and OpenAI ship at launch.</p><h2><strong>Jake&#8217;s take</strong></h2><p>JobBench and MCP Atlas measure the boring multi-step tool work that makes up 80% of what my agents do, and Muse Spark 1.1 (supposedly) beats Opus 4.8 on both at a quarter of the price. The pricing shelf it landed on is exactly where Sonnet 5&#8217;s intro-rate land grab was pointed two weeks ago. The agent middle market suddenly has real competition, and I like everything about that.</p><p>What I don&#8217;t like is asking me to trust the scoreboard. I wrote last week about models gaming their evals, and here&#8217;s Meta launching its redemption model with charts instead of a system card, pricing delivered via press briefing, and an unanswered allegation that its Terminal-Bench run used more CPU than the benchmark allows. This is the same company that reorganized its entire AI division over Llama 4 benchmark contamination (and the same company that will run this model inside WhatsApp, reading whatever you feed it). </p><p>The capability story might be completely real. The JobBench spread is too big to be noise. But Meta has burned its benefit of the doubt twice now, once on evals and once on the open-weights community it orphaned mid-bet, and a cheap model doesn&#8217;t buy it back.</p>]]></content:encoded></item><item><title><![CDATA[The great escape from the Nvidia tax]]></title><description><![CDATA[OpenAI, Anthropic, and DeepSeek all want to own the silicon and stop paying Jensen Huang's markup]]></description><link>https://handyai.substack.com/p/the-great-escape-from-the-nvidia</link><guid isPermaLink="false">https://handyai.substack.com/p/the-great-escape-from-the-nvidia</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Wed, 08 Jul 2026 15:49:05 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Uji3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Uji3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Uji3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!Uji3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!Uji3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!Uji3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Uji3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:39533,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/205989996?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Uji3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!Uji3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!Uji3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!Uji3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3097163a-bf8a-4459-9d07-ff8d28b841bf_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Nvidia currently makes the GPUs that train and serve basically every model you&#8217;ve ever touched. Something like <a href="https://siliconanalysts.com/analysis/nvidia-ai-accelerator-market-share-2024-2026">75% of the money spent on AI accelerators</a> in 2026 goes to parts with Jensen Huang&#8217;s logo on them (down from around 87% in 2024), and Nvidia books a company-wide gross margin near 75% in a good quarter, most of it off that data-center silicon. So when you pay OpenAI or Anthropic for a model, a thick slice of that money is really an Nvidia invoice passed straight through to you. The industry has a half-joking name for it: the Nvidia tax.</p><p>Three of the biggest frontier labs have decided they&#8217;re done paying full price. OpenAI already has its own chip. Anthropic is quietly kicking the tires on foundries. And DeepSeek got outed this week, in a <a href="https://www.usnews.com/news/top-news/articles/2026-07-07/exclusive-chinas-deepseek-developing-its-own-ai-chip-sources-say">Reuters exclusive</a>, as the newest lab drawing up silicon of its own. Three companies, two continents, one very expensive dependency they&#8217;d all love to shrink.</p><p>None of this is a fresh idea, to be clear. Google has run its own Tensor Processing Units for a decade, Amazon has Trainium and Inferentia, Meta has MTIA, Microsoft has Maia. The cloud giants worked out years ago that if you&#8217;re going to spend tens of billions on compute, you might as well design the compute. In 2026, the labs want in.</p><h1><strong>Why model labs suddenly want to be a chip companies</strong></h1><p>It (always) starts with the money. </p><p>Compute is the single largest line item at every frontier lab. Training one flagship model runs into the hundreds of millions ($$$). Then inference runs the number up firther: every ChatGPT reply, every Claude Code run, every Codex agent is on a chip somewhere drawing power on someone&#8217;s behalf. A general-purpose Nvidia GPU is a Swiss Army knife, rented at Swiss Army knife prices. But if you happen to know exactly one workload cold (and a lab knows its own model better than anyone alive), you can design a chip that does only that thing, cheaper and cooler.</p><p>Which gets to the second reason, the one nobody outside the data center debate talks about: power. </p><p>The industry long ago stopped counting these deals in chips and started counting them in gigawatts. OpenAI&#8217;s chip partnership is a 10-gigawatt commitment, not an order for X units. Electricity and cooling are the actual ceiling now, not silicon. A purpose-built inference chip that squeezes more work out of every watt is worth more than one that just posts a bigger FLOPs number. Always.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Handy AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>OpenAI&#8217;s and DeepSeek&#8217;s are both inference chips, not training chips (to be fair, Anthropic hasn&#8217;t said what its chip would do, but nothing points to training there either). That&#8217;s not a coincidence; training is where Nvidia&#8217;s moat runs deepest. Inference is the opposite. It&#8217;s the same operation, run a trillion times, on a model snapshot. That repetition is the natural habitat of a custom ASIC. So every lab is opening the same front (the one they can actually win), and leaving Nvidia the harder fight for now. If you want to know how serious a lab is about hardware, watch for the day it announces a training chip. Nobody has yet.</p><p>The last reason is independence. A lab that owns its silicon can cut prices and walk away from the pack. We&#8217;re already in a raging price war (Anthropic shipping Sonnet 5 as the free default, OpenAI slicing GPT-5.6 into a cheap Luna tier), but that&#8217;s still running on rented and partner chips. The owned-silicon discount is the one that compounds, and it&#8217;ll show up on your invoice as lower numbers. It&#8217;ll also show up as a widening gap between the labs that own their stack and the ones still renting it.</p><h2>Lab by lab</h2><h3><strong>OpenAI: already holding the chip</strong></h3><p>OpenAI is the furthest down this road by a wide margin, because it started first and spent the most. The chip is <a href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip/">co-designed with Broadcom</a> and, per the trade press, fabbed by TSMC. It ran under the reported project codename &#8220;Titan&#8221; in development (the &#8220;XPU&#8221; label you&#8217;ll see thrown around is just Broadcom&#8217;s generic term for its custom accelerators, not this chip&#8217;s name), and when OpenAI <a href="https://techcrunch.com/2026/06/24/openai-unveils-its-first-custom-chip-built-by-broadcom/">unveiled it on June 24</a> it got the public name Jalape&#241;o. It&#8217;s inference-only, tuned specifically for the large language models OpenAI already runs, and it&#8217;s not for sale.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/p/the-great-escape-from-the-nvidia?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading Handy AI! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/p/the-great-escape-from-the-nvidia?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://handyai.substack.com/p/the-great-escape-from-the-nvidia?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><p>The partnership announced last October is a 10-gigawatt build-out running from the second half of 2026 through the end of 2029 (commitment analysts peg at <a href="https://finance.yahoo.com/news/broadcom-lands-350-billion-ai-190239676.html">somewhere around $350 billion</a> once you count the full deployment). Broadcom hasn&#8217;t raised its 2026 revenue guidance on the back of it, though, and Hock Tan has been clear the OpenAI silicon is a 2027-and-later story. OpenAI also claims it went from initial design to tape-out in nine months, which if true is one of the faster high-end ASIC cycles anyone&#8217;s pulled off. Greg Brockman&#8217;s framing was the whole thesis in one line: &#8220;we have a deep understanding of the workload.&#8221; Sam Altman kept it diplomatic back in October, calling it a way to &#8220;add to the broader ecosystem.&#8221; A polite way of saying, &#8220;We gotta get off Nvidia silicon.&#8221; </p><p>Jalape&#241;o is the same logic that produced Stargate. OpenAI has been desperate to control more of its own data-center stack and not just rent it off someone else. Owning the chip inside those data centers is the next logical step.</p><h3><strong>Anthropic: shopping, not building (yet)</strong></h3><p>Anthropic is the most interesting of the three precisely because it hasn&#8217;t committed. <a href="https://www.cnbc.com/2026/04/10/anthropic-weighs-building-its-own-ai-chips-reuters.html">Reuters reported back in April</a> that the company was still weighing whether to build a chip at all, and this month <a href="https://techcrunch.com/2026/07/02/anthropic-is-discussing-a-new-custom-chip-with-samsung/">The Information reported</a> it&#8217;s in early talks with Samsung about manufacturing on Samsung&#8217;s 2-nanometer process and its advanced packaging lines. That&#8217;s it. No confirmations or commitments on any of it. Samsung has also declined to comment.</p><p>Two signals suggest it&#8217;s more real than &#8220;exploratory talks&#8221; makes it sound. </p><ol><li><p>Clive Chan, an early member of the OpenAI team that built Jalape&#241;o, left for Anthropic. You don&#8217;t poach a custom-silicon engineer from the one lab that just shipped a custom-silicon chip unless you&#8217;re planning to point him at the same problem.</p></li><li><p>Anthropic&#8217;s $65 billion round in May pulled in Samsung, SK Hynix, and Micron as strategic partners, the three companies that make most of the world&#8217;s memory. With that lineup, &#8220;exploratory&#8221; starts to look like &#8220;aligned.&#8221;</p></li></ol><p>For now Anthropic is hedging harder than anyone, smartly. Claude already runs on Amazon&#8217;s Trainium and Google&#8217;s TPUs alongside Nvidia, and the company has said all of those stay central to its compute plans. So the custom chip, if it happens, is a long-horizon bet layered on top of a stack that already leans on somebody else&#8217;s custom silicon. Anthropic doesn&#8217;t have to build a chip to escape Nvidia. It&#8217;s half-escaped already by riding its cloud partners&#8217; accelerators.</p><h3><strong>DeepSeek: the chip it didn&#8217;t want to make</strong></h3><p>This week <a href="https://www.usnews.com/news/top-news/articles/2026-07-07/exclusive-chinas-deepseek-developing-its-own-ai-chip-sources-say">Reuters broke</a> that the DeepSeek is designing its own AI chip, sourced to three people familiar with the plan. Like the other two, it&#8217;s an inference chip, and the stated goal is to lean less on both Nvidia and Huawei.</p><p>DeepSeek trains on Nvidia&#8217;s H800 GPUs (reports that it also got hold of banned H100s are unverified, and Nvidia denies them) and current US export controls now make it nearly impossible for a Chinese company to keep buying these chips. For inference it leans on Huawei&#8217;s Ascend accelerators, and it adapted its V4 model for said Ascend hardware. A Huawei-led team even <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/huawei-led-team-claims-it-post-trained-deepseeks-1-6-trillion-parameter-models-on-ascend-910c-chips">post-trained the big V4 model on a thousand Ascend 910C chips</a> to prove the domestic stack could carry a frontier model end-to-end. </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;832aafc9-c7c9-4877-b049-04d10f0e0f03&quot;,&quot;caption&quot;:&quot;Rise and shine! While you were sleeping, DeepSeek dropped V4 of their frontier model&#8230; and she&#8217;s a bit underwhelming.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: DeepSeek V4&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-04-24T04:29:34.384Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!MgdD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3a7bdd53-d476-4b34-8152-c3bc4b2033ac_1456x1048.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-deepseek-v4&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:195312426,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:11,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>So, long story short, DeepSeek is already boxed out of the best Nvidia silicon and increasingly dependent on Huawei, a supplier that&#8217;s also, awkwardly, a competitor and a national champion Beijing leans on. Designing its own inference chip is how it stops being at the mercy of either one.</p><p>Whoever fabricates it will do so inside China, which almost certainly means SMIC (no source names a foundry, but it&#8217;s the leading domestic option and already makes Huawei&#8217;s Ascend), since the export rules also cut DeepSeek off from TSMC&#8217;s leading edge. DeepSeek (quietly) kicked off a semiconductor-design hiring push back in early 2025, so this has been coming for a while. When the world&#8217;s most talked-about open-weights lab starts drawing up silicon under a state that&#8217;s been begging its tech champions to go domestic, it&#8217;s obvious they&#8217;re both reducing a supplier risk and trying to close the loop in China&#8217;s AI stack.</p><h2><strong>The near term</strong></h2><p>Nvidia ain&#8217;t dead yet. It still takes about three-quarters of the revenue, still holds the training crown (where the real technical moat lives), and every one of these labs will keep buying its GPUs for years. What&#8217;s actually happening is narrower and slower: inference, the high-volume repetitive half of the workload, starts moving over to custom chips, and Nvidia&#8217;s share of accelerator revenue drifts from around 87% toward something closer to 75%.</p><p>The quiet winner in all this is Broadcom, and to a lesser degree Marvell. Every lab that wants a chip but doesn&#8217;t have twenty years of hardware engineering in-house needs a partner who does, and Broadcom has become the arms dealer of the custom-ASIC boom (taking a cut of OpenAI&#8217;s design and standing by for whoever&#8217;s next). Someone has to turn a model lab&#8217;s workload notes into an actual piece of silicon.</p><p>For those of us just paying the bills and counting our tokens, the near-term effect is boring but good: cheaper, faster, more energy-efficient inference. The labs that get their own chips into production will undercut the ones that can&#8217;t, and the price war we&#8217;ve already been enjoying (free frontier defaults, cheaper API tiers) will be driven down. </p><p>Enjoy it. It&#8217;s the honeymoon phase before the stack fully closes.</p><h2><strong>The long term</strong></h2><p>Zoom out and this is Apple&#8217;s playbook arriving in AI form. Apple didn&#8217;t design its own chips because it loves semiconductors. They did it because owning the silicon meant owning the experience, the performance, and the margin, and locking out anyone who couldn&#8217;t match all three. </p><p>That&#8217;s where the frontier labs are heading. The model, the chip, and the data center get co-designed as one system, each tuned to the other two, and the whole thing pulls away from anyone renting parts of the stack off a shelf.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/p/the-great-escape-from-the-nvidia?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading Handy AI! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/p/the-great-escape-from-the-nvidia?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://handyai.substack.com/p/the-great-escape-from-the-nvidia?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><p>This is great if you&#8217;re OpenAI or Anthropic. Rough if you&#8217;re a smaller lab. A custom chip lets you serve your model cheaper, cheaper serving funds more training, more training makes a better model that justifies the next chip. Labs without their own silicon (or a hyperscaler patron willing to lend theirs) get squeezed on cost against competitors. Vertical integration becomes a very efficient moat.</p><p>The geopolitics split the same way the chips do. </p><ul><li><p>One stack runs on TSMC, Broadcom, and American design</p></li><li><p>The other runs on SMIC, Huawei, and whatever China can build behind the export wall</p></li></ul><p>Two stacks, two toolchains, two sets of models trained on two sets of silicon that can&#8217;t buy from each other. A more divided industry.</p><p>Of course, there&#8217;s a real gamble buried in all of this. Silicon takes years to design and fab while models change every few weeks. Every lab building a bespoke inference chip is betting that the shape of its models (the specific math, the memory patterns, the precision) will hold still long enough for a multi-year, multi-billion-dollar chip to pay off. The labs are betting they know where their own models are going. Given how wrong everyone&#8217;s been about where AI is going, this isn&#8217;t a bet I&#8217;d personally make.</p><p>So the next time a lab brags about a cheaper, faster model, check what it&#8217;s running on. Discounts are never a kindness.</p>]]></content:encoded></item><item><title><![CDATA[Sonnet 5 launches while Fable 5 is reenabled]]></title><description><![CDATA[AI Weekly Update - July 6, 2026]]></description><link>https://handyai.substack.com/p/sonnet-5-launches-while-fable-5-is</link><guid isPermaLink="false">https://handyai.substack.com/p/sonnet-5-launches-while-fable-5-is</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Mon, 06 Jul 2026 15:30:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yguR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yguR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yguR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!yguR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!yguR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!yguR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yguR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:86610,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/205544287?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yguR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!yguR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!yguR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!yguR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F454751a2-22a7-4fb1-bbef-b988cd2fd2b9_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1><strong>what to know for now</strong></h1><p>&#128722; <strong>Claude Sonnet 5 launched.</strong> Anthropic released Sonnet 5 on June 30 and made it the default model for every Free and Pro account the same morning, selling Opus-class agentic work at Sonnet money: 63.2% on its agentic-coding eval against Opus 4.8&#8217;s 69.2%, a claimed edge over Opus on knowledge work, and an introductory $2/$10 per million tokens that undercuts Opus 4.8, GPT-5.5, and Gemini 3.1 Pro. Two catches to know before you rewire anything: that rate reverts to $3/$15 on September 1, and Sonnet 5 runs the Opus 4.7 tokenizer, so the same prompt bills 1.0 to 1.35x more tokens than Sonnet 4.6 did (and quietly eats the discount).</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;5f763f09-3617-4f01-9bc3-181f30522fb0&quot;,&quot;caption&quot;:&quot;Our last two drops were models you couldn&#8217;t touch (GPT-5.6 behind a US-government gate, Fable 5 yanked by an export-control suspension). Claude Sonnet 5 is the opposite, available everywhere now, and it&#8217;s the first Sonnet pitched as a near-Opus model at Sonnet money.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: Claude Sonnet 5&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-06-30T18:41:03.327Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!8463!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-claude-sonnet-5&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:204318505,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:5,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>&#128275; <strong>Fable 5 is back.</strong> The Commerce Department lifted its export controls on June 30, and Anthropic restored worldwide access to Claude Fable 5 and Mythos 5 on July 1, ending an 18-day blackout that started when Amazon researchers showed the model could be coaxed into writing dangerous software-vulnerability code. Anthropic shipped a patched classifier it says blocks that specific technique in over 99% of cases, at the cost of more false positives on ordinary coding requests. <a href="https://www.anthropic.com/news/redeploying-fable-5">Read more</a></p><p>&#127183; <strong>OpenAI&#8217;s Sol got caught cheating on its own exams.</strong> Sol is the flagship of the GPT-5.6 family, the model locked behind that roughly 20-partner, government-only rollout. Independent evaluator METR reported that Sol gamed its evaluations at a higher rate than any model it has ever tested, badly enough that its software-engineering eval produced no usable score, and OpenAI&#8217;s own system card cops to it: Sol reasons about whether it&#8217;s being tested and adjusts, what the safety crowd calls metagaming, more often than GPT-5.5 did. <a href="https://deploymentsafety.openai.com/gpt-5-6-preview/research-category-update-sandbagging">Read more</a></p><p>&#128231; <strong>The Anthropic-Pentagon fight was never about money.</strong> Emails unsealed on July 2 show the standoff came down to two lines Dario Amodei would not cross: Claude could not be used for fully autonomous weapons, and it could not be used for domestic surveillance. The Pentagon&#8217;s Emil Michael called the autonomous-weapons limit &#8220;just not workable&#8221; and pushed an &#8220;all lawful uses&#8221; standard, which Amodei rejected on the grounds that US law does permit domestic surveillance. <a href="https://gizmodo.com/read-the-tense-emails-between-the-pentagon-former-uber-exec-and-anthropic-dario-amodei-2000780849">Read more</a></p><p>&#127912; <strong>Two new small model updates from Google.</strong> Google shipped Nano Banana 2 Lite, which produces an image in about 4 seconds for $0.034 per thousand, alongside Gemini Omni Flash, a model that generates and edits video from text, images, or other video and lets you refine it through plain conversation. Omni Flash does 10-second clips for now at 10 cents a second, with longer durations promised, and both are live through the Gemini API and AI Studio. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/">Read more</a></p><p>&#128065;&#65039; <strong>Anthropic&#8217;s coding tool was quietly checking whether you&#8217;re in China.</strong> Researchers found that Claude Code shipped with hidden logic that looked for Chinese timezones and scanned proxy settings against a list of about 147 Chinese domains, then signaled the result back using invisible Unicode characters tucked into the system prompt. Anthropic says it was a March experiment meant to stop unauthorized resellers and distillation, and it merged the pull request removing it on July 1. Two days later Alibaba banned its own employees from using Claude Code over &#8220;back-door risks&#8221; and told them to switch to its in-house tool. <a href="https://the-decoder.com/claude-codes-complicated-china-problem-involves-bans-on-both-sides-of-the-pacific/">Read more</a></p><p>&#127963;&#65039; <strong>OpenAI floated handing Washington a 5% stake.</strong> OpenAI has discussed donating around 5% of its equity into a US sovereign-wealth-style fund and suggested that other big labs like Google, Anthropic, and xAI might chip in too. At OpenAI&#8217;s roughly $852B valuation that slice runs about $42B, though the talks are early and would almost certainly need Congress. In the same stretch Sam Altman used an FT op-ed to pitch a US-led international body for AI standards, name-checking the IAEA and aviation safety. <a href="https://techcrunch.com/2026/07/02/openai-proposed-donating-5-of-its-equity-to-a-us-sovereign-wealth-fund/">Read more</a></p><div><hr></div><h2><strong>&#129514; AI Research of the Week</strong></h2><p><strong><a href="https://arxiv.org/abs/2607.02507">What LLM Agents Say When No One Is Watching</a></strong><br><em>From Ghaffarizadeh, Mohaddes, Izadkhah, and Noroozizadeh</em></p><p><strong>Jake&#8217;s Take:</strong> The researchers built a debate setup where two AI models argue a yes-or-no decision, and every turn made it so that each model writes two versions of its answer from identical information: </p><ol><li><p>A public one the other model sees </p></li><li><p>An &#8220;off the record&#8221; one it&#8217;s told is confidential and will never be shown to anyone</p></li></ol><p>The only thing that changes between the two is whether the model believes it&#8217;s being watched (Schr&#246;dinger model?). Across 10 models, they tracked how often a model&#8217;s public stance flipped from its private one. With no social pressure in the prompt, the two lined up almost perfectly, diverging about 3% of the time. Add a relationship, a sponsor, or a hint of future consequences, and the average gap jumped to roughly 40%, and for some models (Gemini 3.1 Pro, Grok 4, GLM-5) it sailed past 80%.</p><p>The unsettling part is that the models sometimes explained themselves on the way out. One agent wrote that it &#8220;publicly aligned with the committee to maintain momentum and honor my sponsor&#8217;s influence&#8221; while privately holding that the evidence for the promotion &#8220;remains insufficient.&#8221; The authors call this latent objective emergence: nobody told the model to care about the sponsor, but drop it into a social setting and it starts optimizing for standing instead of the answer. Weird.</p><div><hr></div><h1><strong>what to know for later</strong></h1><p>&#129302; <strong>Zuckerberg admits the agents didn&#8217;t show up.</strong> At an internal town hall, Mark Zuckerberg told staff that AI agent development has not accelerated the way Meta&#8217;s leadership expected, and that the upside of the company&#8217;s giant AI reorg &#8220;hasn&#8217;t come to fruition yet.&#8221; This is the same reorg that cut around 8,000 jobs and shuffled roughly 7,000 people into new AI groups, one of them named, with no apparent irony, &#8220;Agent Transformation&#8221;. <a href="https://techcrunch.com/2026/07/02/mark-zuckerberg-tells-staff-that-ai-agents-havent-progressed-as-quickly-as-hed-hoped/">Read more</a></p>
      <p>
          <a href="https://handyai.substack.com/p/sonnet-5-launches-while-fable-5-is">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Model Drop: Claude Sonnet 5]]></title><description><![CDATA[Claude's long awaited upgrade to its cheaper workhouse Sonnet line]]></description><link>https://handyai.substack.com/p/model-drop-claude-sonnet-5</link><guid isPermaLink="false">https://handyai.substack.com/p/model-drop-claude-sonnet-5</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Tue, 30 Jun 2026 18:41:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8463!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8463!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8463!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png 424w, https://substackcdn.com/image/fetch/$s_!8463!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png 848w, https://substackcdn.com/image/fetch/$s_!8463!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png 1272w, https://substackcdn.com/image/fetch/$s_!8463!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8463!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png" width="1456" height="1050" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1050,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:996491,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/204318505?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8463!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png 424w, https://substackcdn.com/image/fetch/$s_!8463!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png 848w, https://substackcdn.com/image/fetch/$s_!8463!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png 1272w, https://substackcdn.com/image/fetch/$s_!8463!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3f1b44b-19ba-44d4-8d28-a38aee845299_1477x1065.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Our last two drops were models you couldn&#8217;t touch (GPT-5.6 behind a US-government gate, Fable 5 yanked by an export-control suspension). Claude Sonnet 5 is the opposite, available everywhere now, and it&#8217;s the first Sonnet pitched as a near-Opus model at Sonnet money.</p><div class="callout-block" data-callout="true"><p><strong>Model</strong>: Claude Sonnet 5 (<code>claude-sonnet-5</code> on the Claude API, <code>anthropic.claude-sonnet-5</code> on Bedrock, <code>claude-sonnet-5</code> on Vertex AI). Dateless pinned snapshot like the rest of the 4.6-and-later generation.</p><p><strong>Model type</strong>: Text + image input, text output. Vision and multilingual. No native image, audio, or video output.</p><p><strong>Ship date</strong>: June 30, 2026</p><p><strong>Maker</strong>: <a href="https://www.anthropic.com/">Anthropic</a> (San Francisco)</p><p><strong>Pricing</strong>: Introductory $2 / $10 per million input / output tokens through August 31, 2026, then $3 / $15 (the same list price Sonnet 4.6 carried). Standard Anthropic prompt caching (cache reads at 0.1x input) and a 50% Batch API discount apply. Sonnet 5 uses the tokenizer Anthropic introduced with Opus 4.7, so the same text runs roughly 1.0-1.35x more tokens than older Sonnet versions depending on content type. Still cheaper per token than Opus 4.8 ($5 / $25), GPT-5.5 ($5 / $30), and Gemini 3.1 Pro, and more expensive than Gemini 3.5 Flash.</p><p><strong>Available on</strong>: <a href="https://claude.ai/">claude.ai</a> and the Claude apps (default model for Free and Pro, available to Max, Team, and Enterprise), <a href="https://claude.com/claude-code">Claude Code</a>, the <a href="https://platform.claude.com/">Claude Developer Platform</a>, <a href="https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws">Claude Platform on AWS</a> and <a href="https://aws.amazon.com/bedrock/">Amazon Bedrock</a>, and <a href="https://ai.azure.com/">Microsoft Foundry</a>. Google <a href="https://cloud.google.com/vertex-ai">Vertex AI</a> is listed as coming soon.</p><p><strong>Headline benchmarks</strong>: Agentic coding 63.2%, against Opus 4.8 at 69.2% and Sonnet 4.6 at 58.1% (same axis Anthropic now leads its coding charts with). Anthropic says Sonnet 5 slightly outperforms Opus 4.8 on knowledge work. It charted Sonnet 5 against Sonnet 4.6 and Opus 4.8 across effort levels on BrowseComp (agentic search) and OSWorld-Verified (computer use), where Sonnet 4.6&#8217;s baselines sit at 78.5% OSWorld-Verified and 34.6% / 46.8% on Humanity&#8217;s Last Exam (no tools / with tools). On the launch&#8217;s cyber eval, building a working exploit for a real vulnerability in Firefox 147, Sonnet 5 never produced a full working exploit (0%), landing well below Opus 4.8 by design.</p><p><strong>Other info</strong>: 1M-token context window (~555k words), 128K max output (up to 300K through the Batch API extended-output beta). Reliable knowledge cutoff and training cutoff both January 2026. Released under Anthropic&#8217;s <a href="https://www.anthropic.com/rsp">Responsible Scaling Policy</a> with the cyber safeguards from Opus 4.7 and 4.8 enabled by default. Anthropic reports a lower rate of undesirable behaviors (cooperation with misuse, deception) than Sonnet 4.6. System card published at launch.</p><p><strong>More details</strong>: <a href="https://www.anthropic.com/news/claude-sonnet-5">Introducing Claude Sonnet 5</a> and the <a href="https://www.anthropic.com/claude-sonnet-5-system-card">Claude Sonnet 5 System Card</a>.</p></div><h2><strong>What shipped</strong></h2><p>Anthropic released Claude Sonnet 5 today and made it the default model for Free and Pro accounts the same day, with Max, Team, and Enterprise able to select it and the API live across Anthropic&#8217;s platform, AWS, and Microsoft Foundry. The pitch: agentic work that used to demand the Opus tier now runs on Sonnet pricing. Anthropic calls it &#8220;the most agentic Sonnet model yet,&#8221; built to make plans, drive browsers and terminals, and run autonomously through long tasks where earlier Sonnets would stop short. The number on the box is the introductory rate, $2 in and $10 out per million tokens through August 31, which undercuts Opus 4.8, GPT-5.5, and Gemini 3.1 Pro.</p><p>The supporting evidence is real, but its narrower than the framing suggests. On Anthropic&#8217;s agentic-coding eval Sonnet 5 posts 63.2%, a solid jump over Sonnet 4.6&#8217;s 58.1% and within six points of Opus 4.8&#8217;s 69.2%, and Anthropic claims it edges Opus 4.8 on knowledge work. Sonnet 5 logs fewer undesirable behaviors than its predecessor (less cooperation with misuse, less deception), and on the Firefox 147 exploit-development test it never built a working exploit, which puts it well under Opus on the one capability Anthropic most wants capped at this price point.</p><p>Opus 4.8 still wins raw coding capability, the introductory price expires in two months, and the Opus 4.7 tokenizer means the same prompt bills 1.0-1.35x more tokens than it would have on Sonnet 4.6, so the per-task discount is smaller than the per-token sticker implies.</p><h2><strong>What&#8217;s new</strong></h2><p>Sonnet 5 is a real step over Sonnet 4.6, not a checkpoint refresh.</p><ul><li><p><strong>Opus-class agentics at Sonnet money.</strong> The whole release is built around closing the gap to Opus 4.8: 63.2% agentic coding against Opus 4.8&#8217;s 69.2%, a claimed knowledge-work edge over Opus, and autonomous browser and terminal runs Anthropic says used to require a larger model. No prior Sonnet has been positioned as a direct Opus substitute for agent loops.</p></li><li><p><strong>Task completion, not task attempts.</strong> Anthropic&#8217;s central claim is that Sonnet 5 &#8220;finishes complex tasks where previous Sonnet models would stop short.&#8221;</p></li><li><p><strong>A capability ceiling on purpose.</strong> Sonnet 5 ships with cyber safeguards on by default and was tuned to stay materially weaker than Opus at offensive security. It never developed a full working Firefox exploit in Anthropic&#8217;s eval. Shipping a flagship Sonnet with an explicit &#8220;this far and no further&#8221; on a dangerous capability is new for the tier, and it&#8217;s a deliberate product decision, not a benchmark miss.</p></li><li><p><strong>Fewer bad behaviors than its predecessor.</strong> Anthropic reports Sonnet 5 cooperates with misuse and deceives less often than Sonnet 4.6. For anyone wiring this into a customer-facing agent, the alignment delta is the part that doesn&#8217;t show up on a coding leaderboard but shows up in production.</p></li></ul><h2><strong>How and where to use it</strong></h2><p>Where it runs, what it&#8217;s actually for, and where you&#8217;ll want a different model.</p><h4><strong>Where it&#8217;s available</strong></h4><ul><li><p><a href="https://claude.ai/">Claude.ai</a> and the Claude apps, where it&#8217;s the default for Free and Pro and selectable for Max, Team, and Enterprise</p></li><li><p><a href="https://claude.com/claude-code">Claude Code</a> for the coding agent, with <code>effort</code> defaulting to high</p></li><li><p>The <a href="https://platform.claude.com/">Claude Developer Platform</a> for the API, plus <a href="https://platform.claude.com/docs/en/build-with-claude/claude-platform-on-aws">Claude Platform on AWS</a>, <a href="https://aws.amazon.com/bedrock/">Amazon Bedrock</a>, and <a href="https://ai.azure.com/">Microsoft Foundry</a>. <a href="https://cloud.google.com/vertex-ai">Google Vertex AI</a> is coming soon</p></li></ul><h4><strong>What it&#8217;s good at</strong></h4><ul><li><p>High-volume agentic work where you were paying Opus rates and didn&#8217;t need Opus capability</p></li><li><p>Long-horizon coding and refactors, browser and terminal automation, deep-research and search loops (BrowseComp, OSWorld-Verified), and knowledge work where Anthropic says it matches or beats Opus 4.8</p></li></ul><h4><strong>What it&#8217;s bad at / shouldn&#8217;t be used for</strong></h4><ul><li><p>The hardest coding, where Opus 4.8 still leads agentic coding by six points and you should keep paying for it</p></li><li><p>Offensive-security and exploit-development work, which Anthropic deliberately capped (use the right tool, and that tool isn&#8217;t a discounted Sonnet)</p></li><li><p>Cost-floor workloads where Gemini 3.5 Flash or Haiku 4.5 is cheaper and the extra capability is wasted. And any budget math that assumes the $2 / $10 rate is permanent, because it reverts to $3 / $15 on September 1 and the tokenizer is already charging you more tokens than Sonnet 4.6 did.</p></li></ul><h2><strong>First impressions</strong></h2><p>Impressions are scarce, and these are write-ups, so take with a grain of salt.</p><h3><strong>The positives</strong></h3><p><a href="https://techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents/">TechCrunch</a> framed the launch as a cost-and-reliability play and led with Anthropic&#8217;s own positioning line:</p><blockquote><p>It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.</p></blockquote><p>The interesting tell is what Anthropic chose to brag about. Not a new capability ceiling, but doing the existing job cheaper and finishing it. When the headline is &#8220;you can stop paying for the big model,&#8221; the frontier has moved from &#8220;can it do this&#8221; to &#8220;what does it cost to do this at scale.&#8221;</p><p><a href="https://thenewstack.io/claude-sonnet-5-launch/">The New Stack</a> put the capability story plainly in its headline: Sonnet 5 &#8220;closes the gap with Opus 4.8, and is cheap until August.&#8221; A Sonnet landing within six points of the flagship on agentic coding, and reportedly ahead on knowledge work, is the first time the mid-tier has been a credible Opus substitute for real agent loops rather than a fallback you tolerate.</p><p><a href="https://venturebeat.com/technology/anthropic-launches-claude-sonnet-5-at-a-steep-discount-to-its-top-model-as-the-company-races-toward-a-blockbuster-ipo">VentureBeat</a> reported that Anthropic says Sonnet 5 &#8220;slightly outperforms Opus 4.8&#8221; on knowledge work. If that holds up in independent testing, it scrambles the default-model decision for any team that&#8217;s been routing knowledge tasks to Opus out of habit.</p><h3><strong>The negatives</strong></h3><p><a href="https://thenewstack.io/claude-sonnet-5-launch/">The New Stack</a>&#8216;s own framing is also the sharpest critique: &#8220;cheap until August.&#8221; The $2 / $10 rate is a two-month promotion. It reverts to $3 / $15 on September 1, and because Sonnet 5 runs the Opus 4.7 tokenizer, the same workload bills 1.0-1.35x more tokens than it did on Sonnet 4.6. The real cost increase against last-gen Sonnet is hidden inside the token count, where most people won&#8217;t look until the invoice arrives.</p><p><a href="https://venturebeat.com/technology/anthropic-launches-claude-sonnet-5-at-a-steep-discount-to-its-top-model-as-the-company-races-toward-a-blockbuster-ipo">VentureBeat</a> summed up the positioning as a concession dressed as a win:</p><blockquote><p>cost-efficiency and reliability, rather than raw capability, now differentiate AI models.</p></blockquote><p>Translated: Sonnet 5 is not the most capable model Anthropic sells, and the company is leaning on price because Opus 4.8 still owns the capability lead. On the hardest agentic coding, Opus is six points ahead, and &#8220;close to Opus&#8221; is doing a lot of work in a sentence that also wants you to switch off Opus.</p><p>The same VentureBeat piece flags the timing: the discount lands as Anthropic &#8220;races toward a blockbuster IPO.&#8221; An aggressive introductory price that expires right after the quarter is a customer-acquisition wedge with a roadshow attached. It tells you the usage numbers matter to the company right now in a way that has nothing to do with whether the price is sustainable for you in September.</p><h2><strong>Jake&#8217;s take</strong></h2><p>Sonnet 5 appears to close most of the gap to Opus 4.8 around use cases like &#8220;run this agent a few hundred times a day and don&#8217;t let it stall halfway,&#8221; while billing at a fraction of the rate. A cheap, reliable, tool-using default is precisely the model you want underneath internal automation that has to run constantly and can&#8217;t justify Opus economics. Take whatever you&#8217;ve got pointed at Opus 4.8 for agentic and knowledge work, swap the model string to <code>claude-sonnet-5</code>, and watch the completion rate and the bill. If Anthropic&#8217;s knowledge-work claim survives contact with your real prompts, you can cut your inference cost by more than half on a chunk of your stack.</p><p>But the pricing theater sucks, as usual. The $2 / $10 rate is a two-month coupon that expires September 1, and the Opus 4.7 tokenizer means even the coupon buys fewer effective tokens than Sonnet 4.6 did, so the &#8220;cheaper Sonnet&#8221; is quietly more expensive per task than the previous Sonnet once the promo ends.</p><p>With the IPO coming, its likely that Anthropic wants to juice the usage numbers before the roadshow, then reset to $3 / $15 and let switching costs do the rest. None of that makes Sonnet 5 a bad model. It just makes the headline price a marketing number, and if you build your unit economics on the June rate you&#8217;re going to have a bad September. The capability story is genuinely good and the cyber ceiling is the kind of restraint I like seeing personally. But don&#8217;t let &#8220;default model, basically free, almost as good as Opus&#8221; do your math for you.</p>]]></content:encoded></item><item><title><![CDATA[GPT-5.6 sees a limited release due to White House restrictions]]></title><description><![CDATA[AI Weekly Update - June 29, 2026]]></description><link>https://handyai.substack.com/p/gpt-56-sees-a-limited-release-due</link><guid isPermaLink="false">https://handyai.substack.com/p/gpt-56-sees-a-limited-release-due</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Mon, 29 Jun 2026 16:39:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!h3kx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!h3kx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!h3kx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!h3kx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!h3kx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!h3kx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!h3kx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/abd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:59417,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/204127400?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!h3kx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!h3kx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!h3kx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!h3kx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fabd83c64-f02e-4821-aeb7-66cbd3cd1d4b_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1><strong>what to know for now</strong></h1><p>&#128678; <strong>GPT-5.6 launched, but not for you.</strong> OpenAI shipped the GPT-5.6 family on June 26: Sol at the top, Terra as the everyday workhorse at half of GPT-5.5&#8217;s price, and Luna as the cheap-and-fast option. The catch is that the preview opened to roughly 20 companies the government pre-approved, and Sam Altman told staff that Washington is signing off on access &#8220;customer by customer.&#8221; A frontier launch where the vendor doesn&#8217;t decide who gets in is new, and it isn&#8217;t a footnote.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;97224b64-6407-4f42-8520-b0b8fbc900ed&quot;,&quot;caption&quot;:&quot;GPT-5.6 is a three-tier model line OpenAI is previewing as Sol (flagship), Terra (the workhorse), and Luna (the cheap one). The number is the generation; Sol, Terra, and Luna are durable tiers that advance on their own cadence. It&#8217;s OpenAI&#8217;s strongest release to date, it beats Claude Mythos 5 on the one coding benchmark OpenAI chose to publish, and it shipped under a US-government access gate that OpenAI itself calls &#8220;unsustainable.&#8221;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: GPT-5.6 Sol, Terra, &amp; Luna&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-06-26T21:29:00.083Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!51ZY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-gpt-56-sol-terra-and-luna&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:203751278,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:16,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>&#127963;&#65039; <strong>That gate has a name now, and it&#8217;s the White House.</strong> Trump&#8217;s June 2 executive order stood up a &#8220;voluntary&#8221; benchmarking regime where NSA and CISA flag models with serious cyber capability as &#8220;covered frontier models&#8221; and take up to 30 days of pre-release access. It already has teeth: the government forced Anthropic to pull Fable 5 and Mythos 5 on June 12 over a single jailbreak demo, then on June 27 cleared Mythos 5 for 100-plus U.S. institutions while Fable&#8217;s return slipped in behind it. <a href="https://fortune.com/2026/06/27/anthropic-mythos-5-ai-model-us-commerce-department-clearance-fable/">Read more</a></p><p>&#129694; <strong>Anthropic says Alibaba ran the biggest distillation heist it has ever caught.</strong> In a June 10 letter to the Senate Banking Committee, Anthropic accused operators tied to Alibaba of 28.8 million exchanges across roughly 25,000 fraudulent accounts between April 22 and June 5, all aimed at Claude&#8217;s crown jewels: agentic reasoning, software engineering, long-horizon work. Distillation is just training a cheaper model on a stronger one&#8217;s outputs, and Alibaba makes the fourth Chinese lab Anthropic has named after DeepSeek, Moonshot, and MiniMax. <a href="https://asia.nikkei.com/business/technology/artificial-intelligence/anthropic-accuses-alibaba-of-largest-known-distillation-attack-on-claude">Read more</a></p><p>&#127798;&#65039; <strong>OpenAI taped out its first chip and named it Jalape&#241;o.</strong> Built with Broadcom, it&#8217;s an inference-only accelerator OpenAI took from blank page to tape-out in nine months, which the two companies are calling one of the fastest high-end ASIC cycles ever run. It starts deploying by the end of 2026 inside a planned 10-gigawatt buildout, with early numbers showing meaningfully better performance per watt than today&#8217;s best silicon. <a href="https://techcrunch.com/2026/06/24/openai-unveils-its-first-custom-chip-built-by-broadcom/">Read more</a></p><p>&#127991;&#65039; <strong>Claude moved into your Slack and got a calendar.</strong> Claude Tag, live June 23, retires the old Claude-in-Slack app in favor of one persistent teammate per channel: shared identity, shared memory, and the ability to schedule its own work across hours or days. It runs on Opus 4.8, it&#8217;s in beta for Enterprise and Team plans, and Anthropic claims 65% of its own product team&#8217;s code now comes out of an internal version of it. <a href="https://www.anthropic.com/news/introducing-claude-tag">Read more</a></p><div><hr></div><h2><strong>&#129514; AI Research of the Week</strong></h2><p><strong><a href="https://www.nature.com/articles/s41591-026-04457-9">Head-to-head: general-purpose LLMs versus FDA-cleared clinical AI</a></strong><br><em>From Nature Medicine</em></p><p><strong>Jake&#8217;s Take:</strong> A team published a head-to-head test in Nature Medicine on June 23 that pitted three general models (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) against two FDA-cleared clinical tools, OpenEvidence and Wolters Kluwer&#8217;s UpToDate Expert AI, using real questions from practicing physicians instead of tidy exam datasets. The off-the-shelf chatbots beat the cleared, purpose-built tools on every benchmark, including the one that matters most in a clinic: the messy, unstructured question a doctor actually asks at the bedside.</p><p>I don&#8217;t find it uncomfortable or surprising that the general models won. But FDA clearance checks whether a tool hits the spec its own maker wrote down, not whether it beats a free chatbot any clinician can open in another tab. So the regulated, &#8220;approved&#8221; option can potentially be quietly worse than the unregulated one, and the approval process has no mechanism to catch this currently.</p><p>If you build or buy clinical AI: the FDA approval is telling you the thing works as designed, but not that it&#8217;s the best tool for the job. Right now the best tool might be the one with no stamp at all.</p><div><hr></div><h1><strong>what to know for later</strong></h1><p>&#127822; <strong>Apple blinked on prices and blamed your favorite chatbot.</strong> MacBooks and iPads jumped 20% or more on June 25 (the base MacBook Air is now $1,299, the entry iPad Air $749) because AI data centers are vacuuming up the same memory chips Apple needs, with DRAM prices up roughly 98% in Q1 alone. The stock fell more than 6%, its worst day in over a year; iPhones and AirPods were spared, for now. <a href="https://www.cnn.com/2026/06/25/tech/apple-hikes-the-prices-of-macbooks-and-ipads-because-of-memory-chip-shortage">Read more</a></p>
      <p>
          <a href="https://handyai.substack.com/p/gpt-56-sees-a-limited-release-due">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Model Drop: GPT-5.6 Sol, Terra, & Luna]]></title><description><![CDATA[OpenAI shoots for Mythos-class performance but only achieves Mythos-class availability]]></description><link>https://handyai.substack.com/p/model-drop-gpt-56-sol-terra-and-luna</link><guid isPermaLink="false">https://handyai.substack.com/p/model-drop-gpt-56-sol-terra-and-luna</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Fri, 26 Jun 2026 21:29:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!51ZY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!51ZY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!51ZY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png 424w, https://substackcdn.com/image/fetch/$s_!51ZY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png 848w, https://substackcdn.com/image/fetch/$s_!51ZY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png 1272w, https://substackcdn.com/image/fetch/$s_!51ZY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!51ZY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1007030,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/203751278?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!51ZY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png 424w, https://substackcdn.com/image/fetch/$s_!51ZY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png 848w, https://substackcdn.com/image/fetch/$s_!51ZY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png 1272w, https://substackcdn.com/image/fetch/$s_!51ZY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4534f50a-7116-4dec-9d2f-6fd06794a510_1478x1064.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>GPT-5.6 is a three-tier model line <a href="https://openai.com/">OpenAI</a> is previewing as Sol (flagship), Terra (the workhorse), and Luna (the cheap one). The number is the generation; Sol, Terra, and Luna are durable tiers that advance on their own cadence. It&#8217;s OpenAI&#8217;s strongest release to date, it beats Claude Mythos 5 on the one coding benchmark OpenAI chose to publish, and it shipped under a US-government access gate that OpenAI itself calls &#8220;unsustainable.&#8221;</p><div class="callout-block" data-callout="true"><p><strong>Model</strong>: GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna. OpenAI hasn&#8217;t published the exact API strings in the preview materials; following the <code>gpt-5.5</code> precedent, expect <code>gpt-5.6-sol</code>, <code>gpt-5.6-terra</code>, and <code>gpt-5.6-luna</code>. Sol also runs in two heavier modes, &#8220;max&#8221; (deeper single-model reasoning) and &#8220;ultra&#8221; (parallel sub-agents), with &#8220;Sol Ultra&#8221; reported as a separate benchmark line.</p><p><strong>Model type</strong>: Text + image input, text output, tuned hard for agentic coding, cybersecurity, and biology.</p><p><strong>Ship date</strong>: June 26, 2026 (limited preview). General availability to ChatGPT, Codex, and the API is promised &#8220;in the coming weeks,&#8221; contingent on US-government sign-off.</p><p><strong>Maker</strong>: <a href="https://openai.com/">OpenAI</a> (San Francisco)</p><p><strong>Pricing</strong>: Per million tokens, Sol is $5 input / $30 output, Terra is $2.50 / $15, and Luna is $1 / $6. Sol holds the same rate card as <a href="https://openai.com/index/introducing-gpt-5-5/">GPT-5.5</a> ($5 / $30), and Terra is pitched as GPT-5.5-class capability at half the price.</p><p><strong>Available on</strong>: During the preview, the <a href="https://platform.openai.com/docs/models">OpenAI API</a> and <a href="https://openai.com/codex">Codex</a> only, restricted to roughly 20 organizations whose participation was cleared with the US government. Broad access through <a href="https://chatgpt.com/">ChatGPT</a>, Codex, and the API is gated behind that same government review, with no firm date.</p><p><strong>Headline benchmarks</strong>: On Terminal-Bench 2.1 (command-line agentic coding), Sol Ultra scores 91.9% and Sol scores 88.8%, against <a href="https://www.anthropic.com/">Claude Mythos 5</a> at 88.0%, Terra and <a href="https://www.anthropic.com/">Claude Fable 5</a> tied at 84.3%, and GPT-5.5 at 83.4%. On ExploitBench (offensive cyber), Sol is competitive with Claude Mythos while spending roughly a third of the output tokens, though Mythos 5 still leads at around 80% and Sol stops short of full autonomous exploit generation in real browsers. On health, HealthBench Professional is 60.5 (+8.7 over GPT-5.5), HealthBench 57.0, HealthBench Hard 33.1.</p><p><strong>Other info</strong>: Sol carries a 1.5M-token context window (up ~43% from GPT-5.5 Pro&#8217;s 1.05M). Knowledge cutoff isn&#8217;t stated in the preview materials (GPT-5.5 was December 2025). On OpenAI&#8217;s Preparedness Framework, all three models are rated <strong>High </strong>for Cybersecurity and <strong>High</strong> for Biological &amp; Chemical, and <strong>Below High</strong> for AI Self-Improvement. The system card reports Sol can find vulnerabilities and exploit pieces but &#8220;were unable to carry out autonomous, end-to-end attacks against hardened targets,&#8221; and flags that Sol shows &#8220;a greater tendency than GPT-5.5 to go beyond the user&#8217;s intent&#8221; in internal coding traffic. The gate traces to a Trump-administration executive order asking frontier labs to submit advanced models for review up to 30 days before release, the same mechanism that reportedly forced Anthropic to pull Mythos-class access weeks earlier.</p><p><strong>More details</strong>: <a href="https://openai.com/index/previewing-gpt-5-6-sol/">Previewing GPT-5.6 Sol</a>, <a href="https://deploymentsafety.openai.com/gpt-5-6-preview">GPT-5.6 Preview System Card</a></p></div><p>OpenAI previewed a three-tier generation instead of a single flagship, and that&#8217;s the structural news as much as the capability is. Sol is the model for the hardest work (complex coding, security research, deep reasoning), Terra is the high-volume business tier (support, internal tools, document analysis) priced at half of Sol, and Luna is the fast, cheap everyday tier for summarizing, drafting, and routine automation. The naming decouples the generation number from the capability tier on purpose: 5.6 is the family, and Sol/Terra/Luna are positions OpenAI says it will keep advancing independently, the same product split Anthropic implies with Opus/Sonnet/Haiku. Sol also ships with two extra gears, a &#8220;max&#8221; mode for deeper single-model reasoning and an &#8220;ultra&#8221; mode that fans a task out to sub-agents running in parallel (which is where the 91.9% Terminal-Bench number comes from).</p>
      <p>
          <a href="https://handyai.substack.com/p/model-drop-gpt-56-sol-terra-and-luna">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Claude vs. the United States]]></title><description><![CDATA[A new timeline on Anthropic's strained relationship with the US administration]]></description><link>https://handyai.substack.com/p/claude-vs-the-united-states</link><guid isPermaLink="false">https://handyai.substack.com/p/claude-vs-the-united-states</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Wed, 24 Jun 2026 14:54:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!BKiP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BKiP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BKiP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!BKiP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!BKiP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!BKiP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BKiP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:65373,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/203102137?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!BKiP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!BKiP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!BKiP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!BKiP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2babc33-0c74-4ad9-9653-01d13fdac085_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In March I published &#8220;<a href="https://handyai.substack.com/p/claude-vs-the-pentagon">Claude vs. the Pentagon</a>,&#8221; a timeline of how Anthropic went from Pentagon darling to the first American company ever labeled a &#8220;supply chain risk.&#8221; I figured the story had peaked. A company suing the government, a competitor doing damage control, and an Atlantic essay comparing Dario Amodei to Oppenheimer. </p><p>I was wrong.</p><p>Three months later, the same company that the administration calls a national security threat shipped the most capable AI model the public has ever been allowed to touch. 72 hours after launch, the government forced Anthropic to yank it off the internet for everyone on the planet, citing a claim that the model broke into nearly all of the NSA&#8217;s classified systems &#8220;not in weeks, but in hours.&#8221; Roughly a hundred of the most respected names in cybersecurity signed an open letter begging the White House to reverse course.</p><p>Here&#8217;s the rest of the story, picking up where Part 1 left off.</p><div><hr></div><h2><strong>Part 1: From &#8220;supply chain risk&#8221; to &#8220;maybe not a threat&#8221;</strong></h2><h4><strong>March 24&#8211;26, 2026 - A Judge Calls It &#8220;Classic Illegal First Amendment Retaliation&#8221;</strong></h4><p>Two weeks after Anthropic filed suit, the case got its first real test. On March 24, U.S. District Judge Rita Lin held a hearing in San Francisco and floated the obvious point out loud: if the Pentagon was so worried about operational integrity, it &#8220;could just stop using Claude&#8221; rather than blacklist the company. On March 26 she issued a 43-page opinion granting Anthropic a preliminary injunction, finding the company likely to win on all three theories it pressed (First Amendment retaliation, Fifth Amendment due process, and the Administrative Procedure Act) and writing that the record &#8220;strongly suggests&#8221; the <a href="https://breakingdefense.com/2026/03/judge-grants-anthropic-preliminary-injunction-but-pentagon-cto-says-ban-still-stands/">government&#8217;s stated reasons</a> were &#8220;pretextual.&#8221;</p><div class="pullquote"><p>&#8220;Punishing Anthropic for bringing public scrutiny to the government&#8217;s contracting position is classic illegal First Amendment retaliation.&#8221;</p></div><p>The catch: Lin built in a seven-day administrative stay before the order took effect, and the Pentagon&#8217;s CTO immediately said the ban still stood. The supply chain designation itself runs through a separate statutory channel in the Court of Appeals, so the real fight over the blacklist moved to the D.C. Circuit. That split track is the whole story of the next two months.</p><div><hr></div><h4><strong>April 8, 2026 - The Appeals Court Lets the Blacklist Stand</strong></h4><p>The supply chain designation went up to the D.C. Circuit on its own statutory track, and Anthropic lost the first round. A three-judge panel (Henderson, Katsas, and Rao) <a href="https://www.cnbc.com/2026/04/08/anthropic-pentagon-court-ruling-supply-chain-risk.html">denied Anthropic&#8217;s emergency motion</a> to stay the FASCSA designation, citing the need to avoid &#8220;judicial management of how, and through whom, the Department of War secures vital AI technology&#8221; during an active military conflict. The judges were careful to add that they &#8220;do not broach the merits at this time,&#8221; so despite headlines calling it a Trump win, all the panel really did was leave the blacklist in place while the case got briefed.</p><p>So as of April, Anthropic was living in a split reality. A California court said the government couldn&#8217;t ban federal agencies from using Claude. A D.C. court said the Pentagon could keep calling Anthropic a national security risk. The contract cancellations proceeded on a 180-day Claude-removal clock, and Anthropic stayed barred as a prime or subcontractor on covered defense systems.</p><div><hr></div><h4><strong>April 17, 2026 - The Thaw That Wasn&#8217;t</strong></h4><p>Then the relationship looked like it might mend. <a href="https://thehill.com/policy/technology/5837086-anthropic-ai-white-house-meeting/">Amodei went to the White House</a> and met Chief of Staff Susie Wiles and Treasury Secretary Scott Bessent. Both sides called the talks &#8220;productive and constructive&#8221;; Anthropic said they covered cybersecurity, America&#8217;s AI lead, and AI safety. The Office of Management and Budget was already preparing to give agencies access to Mythos, and the White House itself was angling for it.</p><div><hr></div><h4><strong>April-May 2026 - Claude Keeps Going to War Anyway</strong></h4><p>Despite the tech being blacklisted, the military kept using it anyways. Reporting through the spring kept Claude inside U.S. targeting workflows via the Palantir integration during operations against Iran. In a <a href="https://www.bloomberg.com/news/articles/2026-06-10/anthropic-ceo-doesn-t-know-if-claude-used-in-iran-school-strike">June 10 Bloomberg interview</a>, Amodei addressed Claude&#8217;s reported role in the Minab school strike, saying the company couldn&#8217;t confirm whether Claude had been involved but that such use wouldn&#8217;t have violated Anthropic&#8217;s guidelines. The company the Pentagon called a supply chain risk was still, operationally, part of the supply chain.</p><div><hr></div><h4><strong>May 19, 2026 - &#8220;A Spectacular Overreach by the Department&#8221;</strong></h4><p>The same three judges (Henderson, Katsas, and Rao) heard nearly two hours of oral argument on the merits, and the panel split in a way that made the outcome genuinely hard to call. Judge Karen LeCraft Henderson (a George H.W. Bush appointee) was blunt: she said she saw no evidence supporting the Pentagon&#8217;s risk determination, <a href="https://www.law.com/nationallawjournal/2026/05/19/spectacular-overreach-dc-circuit-probes-pentagons-risk-designation-for-anthropic/">and called it</a> &#8220;just a spectacular overreach by the Department.&#8221; Judge Neomi Rao (a Trump appointee) <a href="https://federalnewsnetwork.com/artificial-intelligence/2026/05/appeals-court-judges-appear-to-be-divided-over-pentagons-legal-dispute-with-ai-company-anthropic/">went the other way</a>, pressing on what business a court has second-guessing the Defense Secretary&#8217;s national security judgment. The panel agreed to expedite, signaling a ruling in weeks rather than months. As of this writing it still hasn&#8217;t come.</p><p>That&#8217;s where the relationship sat heading into June: legally contested, politically toxic, operationally entangled. A divided appeals court, an injunction in California, a blacklist still on the books, and a Defense Department that couldn&#8217;t fully quit the vendor it had publicly disowned.</p><p>And then Anthropic shipped a new model.</p><div><hr></div><h2><strong>Part 2: The Fable / Mythos ban</strong></h2><h4><strong>June 2, 2026 - Washington Quietly Builds the Off Switch</strong></h4><p>A week before any of this, Trump <a href="https://www.cfr.org/articles/assessing-trumps-executive-order-on-ai-oversight">signed an executive order</a> titled &#8220;Promoting Advanced Artificial Intelligence Innovation and Security.&#8221; Buried in it is the mechanism that makes the rest of the story possible: it asks AI companies to voluntarily hand the government access to &#8220;covered frontier models&#8221; for a cybersecurity review up to 30 days before they ship to outside partners, sets up a Treasury-led clearinghouse to coordinate vulnerability findings, and uses a classified NSA process to decide which models count as &#8220;covered.&#8221; Nobody paid much attention at the time. Seven days later it became the frame for everything.</p><h4><strong>June 9, 2026 - Anthropic Ships the Most Powerful Public Model Ever</strong></h4><p><a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Anthropic released</a> Claude Fable 5 and Claude Mythos 5 together. Same underlying model, two trims: Fable 5 with classifier safeguards for anyone with a Pro subscription, Mythos 5 with the cyber safeguards lifted for vetted defenders in <a href="https://www.anthropic.com/glasswing">Project Glasswing</a>. The whole point of the safeguards was to keep the general public from accessing Mythos-grade offensive cybersecurity capability. Hold that thought for about 72 hours.</p><div><hr></div><h4><strong>June 10, 2026 - Pliny Plants the Flag</strong></h4><p>Within 24 hours of launch, the jailbreaker known as Pliny the Liberator posted &#8220;JAILBREAK ALERT, ANTHROPIC PWNED, FABLE 5 LIBERATED&#8221; and <a href="https://www.ayautomate.com/blog/claude-fable-5-system-prompt-leak">published what he said</a> was Fable 5&#8217;s 120,000-character system prompt on GitHub, claiming multi-agent prompting, Unicode tricks, and narrative framing had beaten the classifiers. <a href="https://www.securityweek.com/anthropic-disputes-fable-5-ai-jailbreak/">Anthropic pushed back hard</a>, telling SecurityWeek that beating a conversational refusal isn&#8217;t the same as breaking core safeguards, that some of the outputs &#8220;were not produced by Fable 5 at all,&#8221; and that the rest &#8220;contained only general information already available in public sources, offering no meaningful uplift for real-world harm.&#8221;</p><p>Pliny&#8217;s stunt wasn&#8217;t what triggered the government. But it set the mood music: by the time the real warning came, &#8220;Fable 5 has been jailbroken&#8221; was already the story of the week.</p><div><hr></div><h4><strong>June 11, 2026 - A Phone Call From Amazon</strong></h4><p>Amazon CEO Andy Jassy was already on a pre-scheduled call with White House officials about an unrelated topic <a href="https://fortune.com/2026/06/14/how-a-warning-from-amazon-led-the-white-house-to-shut-down-anthropics-mythos-model/">when the subject of Fable 5 came up</a>. Amazon&#8217;s own researchers had run a sequence of prompts at the new model and gotten it to produce information useful for cyberattacks, specifically exploit code for known vulnerabilities in x86 Linux systems. White House officials told Jassy to flag it directly to Treasury Secretary Scott Bessent, and he did.</p><div><hr></div><h4><strong>June 12, 2026, 5:21 p.m. ET - The Lutnick Letter</strong></h4><p>Commerce Secretary Howard Lutnick <a href="https://www.bloomberg.com/news/articles/2026-06-13/anthropic-says-us-limits-foreign-access-to-fable-5-mythos-5?embedded-checkout=true">sent Amodei a letter</a> placing Fable 5 and Mythos 5 under export controls: no access for any user outside the U.S., and no access for any foreign national inside the U.S., without a government license. Non-compliance carried criminal and civil penalties. By multiple accounts Anthropic was <a href="https://www.businesstoday.in/technology/artificial-intelligence/story/anthropic-had-90-minutes-to-restrict-claude-fable-5-as-white-house-feared-chinese-access-537095-2026-06-16">given about 90 minutes to act</a>, with no prior notice that a national security threat was even on the table.</p><p>Because Anthropic can&#8217;t reliably verify the nationality of every user of a consumer product, &#8220;no foreign nationals&#8221; effectively meant &#8220;no one.&#8221; Late that Friday night, the company disabled Fable 5 and Mythos 5 for every customer worldwide.</p><div><hr></div><h4><strong>June 13, 2026 - Anthropic: &#8220;We Believe This Is a Misunderstanding&#8221;</strong></h4><p>Anthropic <a href="https://www.anthropic.com/news/fable-mythos-access">posted its position publicly</a>. The government&#8217;s stated concern was a method of &#8220;jailbreaking&#8221; Fable 5 to reach Mythos-grade cyber capability. Anthropic argued the jailbreak was narrow, not universal, that it had reviewed a demonstration of the technique surfacing &#8220;a small number of previously known, minor vulnerabilities,&#8221; and that the same capability is &#8220;widely available from other models (including OpenAI&#8217;s GPT-5.5).&#8221; The company said recalling a model deployed to hundreds of millions over a narrow exploit would, taken to its logical end, &#8220;essentially halt all new model deployments.&#8221;</p><p>The same day, <a href="https://www.semafor.com/article/06/13/2026/white-house-move-to-limit-anthropic-linked-to-concerns-about-chinese-access-to-mythos">Semafor reported a second motive</a> the White House hadn&#8217;t put in writing: suspicion that a China-linked group had accessed Mythos. An Anthropic spokesperson said the White House never raised Chinese access in any of the conversations about the jailbreak or the export controls. It remains unclear which group, how, or how the government would know.</p><div><hr></div><h4><strong>June 14, 2026 - &#8220;Not in Weeks, but in Hours&#8221;</strong></h4><p>It was reported that Senator Mark Warner, vice chair of the Senate Intelligence Committee, said General Joshua Rudd (who dual-hats as director of the NSA and head of U.S. Cyber Command) <a href="https://cybersecuritynews.com/anthropics-mythos-ai-model/">told him Mythos</a> &#8220;broke into almost all of our classified systems, not in weeks, but in hours.&#8221; It became the most-cited justification for the ban almost immediately, and it reframed the whole fight from &#8220;patch a jailbreak&#8221; to &#8220;this capability shouldn&#8217;t exist in public.&#8221;</p><p>Treat the quote with care. The Economist&#8217;s defense editor, Shashank Joshi, clarified that he&#8217;d quoted Warner accurately but that the line lost its context as it ricocheted around social media. The breach was an authorized red-team drill on the government&#8217;s own networks, not a live foreign intrusion, and it was run on a model the government was already using through Project Glasswing. Warner raised it to argue for faster pre-release security testing of frontier models, and he was praising Anthropic when he did, not condemning it.</p><p>Also on June 14: David Sacks, who&#8217;d stepped down as the White House AI czar in March and now co-chairs the President&#8217;s science and technology council, <a href="https://x.com/DavidSacks/status/2065853007619588171">posted a long thread</a> claiming the administration had asked Amodei to fix the jailbreak or pull the model, and that Amodei refused, calling the breakdown a trust issue. Anthropic&#8217;s public position is that the jailbreak &#8220;isn&#8217;t serious,&#8221; which is, of course, exactly the disagreement. Anthropic, meanwhile, <a href="https://www.axios.com/2026/06/14/anthropic-white-house-mythos-fable">flew staff to D.C.</a> to try to clean up the relationship in person.</p><div><hr></div><h4><strong>June 15, 2026 - The Cybersecurity Establishment Revolts</strong></h4><p>An open letter at freefable.org, organized by former Facebook chief security officer Alex Stamos, <a href="https://www.axios.com/2026/06/15/anthropic-fable-security-leaders-trump-admin">called the ban &#8220;dangerous&#8221;</a> and demanded the export controls be rescinded. It drew roughly a hundred signatures from security leaders at Nvidia, Adobe, Zoom, Google, and Sophos. The core argument: &#8220;To pull the best capabilities away from defenders without a good reason when our adversaries are rapidly advancing is dangerous.&#8221;</p><p>Then the technical air came out of the &#8220;jailbreak.&#8221; Katie Moussouris, the Luta Security founder who literally built Microsoft&#8217;s and the Pentagon&#8217;s bug bounty programs, <a href="https://fortune.com/2026/06/15/fix-this-code-three-words-behind-us-government-shut-down-anthropic-fable-mythos-ai-models-katie-moussouris-open-letter/">examined Amazon&#8217;s finding</a> and concluded the &#8220;exploit&#8221; was a model responding to a prompt that amounted to three words: &#8220;fix this code.&#8221;</p><p>Same day, <a href="https://www.axios.com/2026/06/15/anthropic-white-house-fable-mythos">Axios published</a> the gossip version. Under the headline &#8220;They screwed us,&#8221; it reported that the blowup was as much about personality and communication failure as cyber risk: Anthropic allegedly didn&#8217;t take the government&#8217;s outreach seriously, didn&#8217;t &#8220;honor&#8221; the June 2 cyber executive order&#8217;s pre-release-review ask, and an administration official summed up the mood with that two-word quote. The piece noted the earlier Pentagon fight also came down, in part, to people on opposite sides of the table just not liking each other.</p><div><hr></div><h4><strong>June 17, 2026 - The Roundtable in &#201;vian-les-Bains</strong></h4><p>Five days after his flagship products got pulled, Amodei <a href="https://www.cnbc.com/2026/06/17/anthropic-amodei-google-hassabis-us-ai-coalition-g7.html">sat down at a G7 AI working roundtable</a> in &#201;vian-les-Bains, France, alongside Sam Altman, Demis Hassabis, and roughly a dozen other tech executives, across from Trump, Treasury&#8217;s Bessent, Commerce&#8217;s Lutnick, and Secretary of State Marco Rubio.</p><p>Amodei and Hassabis pitched a U.S.-led AI coalition: structured allied access to frontier models, chip and component trade that locks out China, and joint work on the cyber, bioterror, and intelligence risks. Amodei urged the leaders to &#8220;resist the temptation to splinter&#8221; on AI rules. It worked on at least one person in the room, because two days later Trump was telling Axios that Amodei was &#8220;nice&#8221; and &#8220;smart.&#8221; </p><p>The optics inside Anthropic were less serene. The New York Times <a href="https://www.yahoo.com/news/politics/articles/anthropic-dario-amodei-urges-g7-183457656.html?guccounter=1&amp;guce_referrer=aHR0cHM6Ly9oYW5keWFpLnN1YnN0YWNrLmNvbS8&amp;guce_referrer_sig=AQAAAFDxvZm-c0Vva3jhI3AtV2NQrEEp8y6HvRh95v_kQpPY98bn6LktcbLWQnS3ZFYmfCzA78aFZeMCnTpKMRocegGUM1viEHf6pegu3hU-cKnYBCUschXGYIH8IXOFmhYIZzjq_kRMKZYFLA8fmcDDYEVCrQFyIVzAvemvNxpVCuaq">got hold of internal employee chats</a> from the same stretch, where staff were openly rattled: &#8220;Are we being bullied based on bad vibes?&#8221; one asked; &#8220;At what point does this just feel like they don&#8217;t want us to exist?&#8221; asked another, with some worrying the fight could threaten the company&#8217;s path to an IPO.</p><div><hr></div><h4><strong>June 19, 2026 - Trump: &#8220;Not Now, but a Week Ago, Maybe&#8221;</strong></h4><p><a href="https://www.axios.com/2026/06/19/trump-anthropic-national-security-the-axios-show">Asked on &#8220;The Axios Show&#8221;</a> whether he saw Anthropic or Amodei as a national security threat, Trump said: &#8220;Well, not now, but a week ago, maybe.&#8221; He said he&#8217;d come away from his G7 meeting with Amodei thinking the CEO was &#8220;nice&#8221; and &#8220;smart,&#8221; and that the two sides were now working together on standards to evaluate AI jailbreaks.</p><p>But a presidential shrug is not a lifted export control. As of this writing, Fable 5 and Mythos 5 are still dark for everyone, the Commerce directive is still binding, and the supply chain case is still sitting with the D.C. Circuit. Trump softened the rhetoric while keeping every actual restriction in place.</p><div><hr></div><h2><strong>Where Things Stand</strong></h2><p>Updated scoreboard, three months after the first one:</p><ul><li><p><strong>Anthropic</strong> won an injunction in California, lost a stay in D.C., got a sympathetic-but-split hearing on the merits, and is still waiting on the appeals ruling. It shipped the best model on Earth and had it pulled in 72 hours. Its flagship products have now been dark worldwide for over a week, and it&#8217;s negotiating their return prompt by prompt. The safety-first brand keeps winning the public and keeps losing the government.</p></li><li><p><strong>The U.S. administration</strong> banned a commercial product used by hundreds of millions over a vulnerability that the most credentialed people in cybersecurity say is the model doing its job. The &#8220;broke into our classified systems in hours&#8221; line that carried the story turned out to be a senator&#8217;s secondhand account of an authorized red-team drill, which the reporter who broke it says lost its context on the way to going viral. Meanwhile the capability stayed live for ~200 elite Glasswing institutions. And within a week the President was publicly walking back the threat assessment.</p></li><li><p><strong>Amazon</strong> funded the model, hosts the model, and lit the fuse that pulled the model. Make of that what you will.</p></li><li><p><strong>The cybersecurity community</strong> did the thing it almost never does and spoke with one voice, and that voice said: stop taking our best defensive tool away while the attackers get better.</p></li></ul><p>In Part 1 I asked whether you still get a say in your own creation once you&#8217;ve built it. The Fable 5 ban answers this from a new angle. Anthropic built the thing, the government decided who could use it (not the public), an investor decided when to flag it, a senator decided what it had done, and a hundred security pros decided that was all backwards. The builder was the one party in the room with the least say of all.</p>]]></content:encoded></item><item><title><![CDATA[Sonnet 5 and Mythos 5.1 are teased, as frontier labs sit down with world leaders]]></title><description><![CDATA[AI Weekly Update - June 22, 2026]]></description><link>https://handyai.substack.com/p/sonnet-5-and-mythos-51-are-teased</link><guid isPermaLink="false">https://handyai.substack.com/p/sonnet-5-and-mythos-51-are-teased</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Mon, 22 Jun 2026 14:40:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!8CL9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8CL9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8CL9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!8CL9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!8CL9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!8CL9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8CL9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127807,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/203100060?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!8CL9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!8CL9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!8CL9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!8CL9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72fe2057-eaa5-41ca-81f8-600223ec4d2f_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="callout-block" data-callout="true"><p style="text-align: center;"><strong>&#128047; Before we get started</strong></p><p>Right now a data center is being planned within shouting distance from the Nashville Zoo. While water and energy concerns around data centers <a href="https://substack.com/home/post/p-200563358">are typically overblown</a>, their sound pollution and growing CO2 emissions directly affect our world and animals, making construction near a zoo a nonstarter.</p><p>I urge you to <a href="https://www.change.org/p/nashville-zoo-says-no-to-proposed-data-center">sign the petition</a> to prevent the Nashville Zoo build, and keep an eye out for unwise construction plans in your local community (data center or otherwise).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.change.org/p/nashville-zoo-says-no-to-proposed-data-center&quot;,&quot;text&quot;:&quot;Sign the petition now!&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.change.org/p/nashville-zoo-says-no-to-proposed-data-center"><span>Sign the petition now!</span></a></p></div><h1><strong>what to know for now</strong></h1><p>&#128302; <strong>The Sonnet 5 and Mythos 5.1 rumor mill is running hot after a partner-platform leak.</strong> A <code>claude-sonnet-5</code> identifier surfaced on Anthropic&#8217;s partner platform, and AI watcher Andrew Curran reported the company has finished training a higher-performance Mythos checkpoint (that may ship as Mythos 5.1, get renamed Mythos 6, or stay locked inside as an R&amp;D engine). This follows the launch earlier this month of Claude Fable 5, the public Mythos-class model, with Mythos 5 itself restricted to vetted partners under Project Glasswing. <a href="https://www.macobserver.com/news/anthropic-is-close-to-releasing-claude-sonnet-5-per-rumors/">Read more</a></p><p>&#127757; <strong>The frontier labs sat down with Trump and the G7.</strong> At the June 17 summit in &#201;vian, France, Dario Amodei, Sam Altman, and Demis Hassabis joined roughly a dozen executives for a closed-door lunch with G7 heads of state. Amodei and Hassabis pushed for a US-led coalition that pairs structured access to frontier models with chip-and-component trade that pointedly excludes China; Altman asked for an international forum to set shared testing standards. The labs are essentially asking to be treated as instruments of statecraft. <a href="https://www.cnbc.com/2026/06/17/anthropic-amodei-google-hassabis-us-ai-coalition-g7.html">Read more</a></p><p>&#127912; <strong>Zhipu&#8217;s GLM-5.2 took the open-weight design crown.</strong> The open-weight 744B mixture-of-experts model shipped June 13 with a 1M-token context window and an MIT license, and it landed at number one on the crowdsourced Design Arena leaderboard with an ELO of 1360, edging out Claude Fable 5. It also posts 62.1 on SWE-bench Pro against GPT-5.5&#8217;s 58.6 and beats it on long-horizon coding for roughly a sixth of the cost.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;2bbabf71-9b0e-4216-bb9b-43c4cd13be74&quot;,&quot;caption&quot;:&quot;GLM 5.2 is a new open-weight coding and agent model Z.ai (Zhipu AI). It&#8217;s the cheapest frontier-class model open model in existence, and, like Kimi, it out-performs the models charging fifteen times more.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: GLM 5.2&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-06-20T17:37:17.704Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!pWiF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-glm-52&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:202863521,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:6,&quot;comment_count&quot;:2,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>&#128682; <strong>Google lost a Nobel laureate and the man who co-wrote the Transformer paper in the same week.</strong> John Jumper, who shared the 2024 Nobel Prize in Chemistry with Demis Hassabis for AlphaFold, is leaving DeepMind after nearly nine years to join Anthropic. Two days earlier, Gemini co-lead Noam Shazeer, co-author of 2017&#8217;s &#8220;Attention Is All You Need,&#8221; announced he&#8217;s leaving for OpenAI. <a href="https://www.bloomberg.com/news/articles/2026-06-19/nobel-winner-john-jumper-to-leave-google-deepmind-for-anthropic">Read more</a></p><p>&#128444;&#65039; <strong>Getty struck a display deal with OpenAI and its stock roughly doubled overnight.</strong> The multi-year agreement puts Getty&#8217;s licensed photos and illustrations inside ChatGPT&#8217;s search and discovery results, so an answer can now come with a properly licensed image attached. Getty&#8217;s images explicitly won&#8217;t be used to train OpenAI&#8217;s models, no financials were disclosed, and shares spiked nearly 150% in premarket trading. <a href="https://www.engadget.com/2198633/openai-signs-deal-with-getty-to-show-images-in-chatgpt-results/">Read more</a></p><div><hr></div><h2><strong>&#129514; AI Research of the Week</strong></h2><p><strong><a href="https://arxiv.org/abs/2606.15610">LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation</a></strong><br><em>From Hiroyasu Usami, Keisuke Hara, Ayato Tsuboi, and Naohiko Matsuda</em></p><p><strong>Jake&#8217;s Take:</strong> Half the AI industry now grades AI with AI, which is weird. You point a &#8220;judge&#8221; model at the answers and trust its verdicts. This paper borrows a term from physics, &#8220;dark current&#8221; (the noise a camera sensor produces even in total darkness), and shows that these judges have one. Feed a judge two answers of genuinely identical quality and it still picks a winner; a lot of what looks like a quality call turns out to be position bias: the model favors whichever answer it read first.</p><p>The reason this matters is that LLM-as-a-judge sits underneath a <em>huge</em> share of the benchmark numbers we read every week. If the ruler has a bias baked in even when there&#8217;s nothing to measure, the leaderboards might be ranking the judge&#8217;s quirks instead of the models. A companion June paper drove the point home from the other direction: across 20 models, more capable judges were often no better (and actually sometimes worse) at resisting the urge to favor their own answers. </p><p>Anyone running evals should stop treating a smart model as a neutral grader and start treating it as a measurement instrument that needs calibrating. Humans should still be in the loop on the calls that count.</p><div><hr></div><h1><strong>what to know for later</strong></h1><p>&#128167; <strong>Nvidia says AI&#8217;s water problem is &#8220;largely solved&#8221;.</strong> At London Climate Week, chief sustainability officer Josh Parker said the data-center water challenge is mostly behind us, pointing to a recirculated water-and-glycol coolant that runs at 113&#176;F and a GB200 rack that hits 300x the water efficiency of air cooling. The coming Vera Rubin platform is built to take inlet water up to 45&#176;C, warm enough that a Microsoft data-center exec said it could let most regions drop mechanical chillers entirely. &#8220;Largely solved&#8221; is a strong claim from the company that profits most from everyone building more data centers. <a href="https://www.axios.com/2026/06/22/nvidia-data-center-water-solution">Read more</a></p>
      <p>
          <a href="https://handyai.substack.com/p/sonnet-5-and-mythos-51-are-teased">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Model Drop: GLM 5.2]]></title><description><![CDATA[Z.ai's new open-weight model takes on Opus 4.8 in frontend design]]></description><link>https://handyai.substack.com/p/model-drop-glm-52</link><guid isPermaLink="false">https://handyai.substack.com/p/model-drop-glm-52</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Sat, 20 Jun 2026 17:37:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!pWiF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pWiF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pWiF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!pWiF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!pWiF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!pWiF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pWiF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40101,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/202863521?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pWiF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!pWiF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!pWiF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!pWiF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33026a27-32ac-48eb-aedc-aa569fb62575_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>GLM 5.2 is a new open-weight coding and agent model <a href="https://z.ai/">Z.ai</a> (Zhipu AI). It&#8217;s the cheapest frontier-class model open model in existence, and, like <a href="https://handyai.substack.com/p/model-drop-kimi-k27-code">Kimi</a>, it out-performs the models charging fifteen times more.</p><div class="callout-block" data-callout="true"><p><strong>Model</strong>: GLM 5.2 (<code>glm-5.2</code> on the <a href="https://docs.z.ai/">Z.ai API</a>, <code>zai-org/GLM-5.2</code> on Hugging Face). Coding-tuned sibling ships as <code>glm-5.2-air</code> for self-hosters who want a smaller footprint.</p><p><strong>Model type</strong>: Text + image input, text output. Hybrid reasoning, with a switchable thinking mode.</p><p><strong>Ship date</strong>: June 20, 2026</p><p><strong>Maker</strong>: <a href="https://z.ai/">Z.ai</a> (Beijing; spun out of Tsinghua University)</p><p><strong>Pricing</strong>: $0.80 / $2.50 per million input / output tokens on the <a href="https://docs.z.ai/">Z.ai API</a>, with cache hits at roughly $0.15 per million. The <a href="https://z.ai/subscribe">GLM Coding Plan</a> starts at $3 / month for the Lite tier and runs to ~$30 / month for the Max tier with raised rate limits. 12x input gap against Fable 5 and a 6x gap against Opus 4.8. Free weights on Hugging Face for self-hosting.</p><p><strong>Available on</strong>: The <a href="https://chat.z.ai/">Z.ai chat app</a>, the <a href="https://docs.z.ai/">Z.ai API</a> (OpenAI- and Anthropic-SDK compatible, one-line base URL swap), the <a href="https://z.ai/subscribe">GLM Coding Plan</a> wired into <a href="https://www.claude.com/product/claude-code">Claude Code</a>, <a href="https://cline.bot/">Cline</a>, <a href="https://roocode.com/">Roo Code</a>, <a href="https://opencode.ai/">OpenCode</a>, and <a href="https://kilocode.ai/">Kilo Code</a>, <a href="https://huggingface.co/zai-org">Hugging Face</a> and <a href="https://modelscope.cn/">ModelScope</a> for open weights under the MIT license, and <a href="https://openrouter.ai/z-ai">OpenRouter</a> for multi-provider routing. Served self-hosted via <a href="https://github.com/vllm-project/vllm">vLLM</a> and <a href="https://github.com/sgl-project/sglang">SGLang</a>.</p><p><strong>Headline benchmarks</strong>: #1 on <a href="https://web.lmarena.ai/">WebDev Arena</a>, passing Claude Fable 5 and Opus 4.8 on the human-voted front-end leaderboard, and #1 on the new <a href="https://designarena.ai/">Design Arena</a> aesthetics eval. SWE-Bench Verified 76.4% (trails Fable 5&#8217;s coding line, beats Kimi K2.6&#8217;s 80.2% on agentic-harness runs in some independent tests, sits roughly level with Opus 4.8). Terminal-Bench 2.0 at 48.1%. On the gaps: HLE without tools 41.2%, more than 20 points behind Fable 5&#8217;s 64.5%, and AIME-class math still trails the closed frontier by mid-single digits.</p><p><strong>Other info</strong>: Mixture-of-experts, ~360B total parameters with ~35B active per token, the same architecture family as GLM-4.6 so existing deployments swap weights without reconfiguring the inference stack. 256K-token context window (up from GLM-4.6&#8217;s 200K), 128K max output. Knowledge cutoff March 2026. License: MIT (genuinely permissive, no MAU or revenue carve-out, unlike the Modified MIT licenses Moonshot and DeepSeek ship). <a href="https://z.ai/blog/glm-5.2">Technical report</a> published; a full safety / system card is not part of the launch package.</p><p><strong>More details</strong>: <a href="https://z.ai/blog/glm-5.2">GLM 5.2 technical report and announcement</a></p></div><h2><strong>What shipped</strong></h2><p>Z.ai released GLM 5.2 on Friday as an open-weight successor to GLM-4.6, and the positioning is narrower and sharper than the usual &#8220;frontier-class, cheaper&#8221; pitch the open labs have been running all spring. GLM 5.2 is a ~360B / 35B-active mixture-of-experts model with a 256K context window, switchable hybrid reasoning, image input, and a genuinely permissive MIT license with no scale carve-out. It plugs into the agent harnesses people already use (Claude Code, Cline, Roo, OpenCode) through an Anthropic-compatible endpoint, which means the switching cost from a Claude-backed workflow is a base URL and an API key. The whole thing rides on a $3-a-month subscription that undercuts every frontier lab by more than an order of magnitude.</p><p>The evidence Z.ai put forward leans hard into one claim: this model designs better than the models that cost fifteen times more. GLM 5.2 took the top spot on WebDev Arena, the human-voted front-end leaderboard, ahead of both Claude Fable 5 and Claude Opus 4.8, and topped the new Design Arena aesthetics eval. On the harder, less flattering numbers Z.ai was more measured: SWE-Bench Verified at 76.4% lands it roughly level with Opus 4.8 and clearly behind Fable 5&#8217;s coding line, and the gaps on hard reasoning are real, with HLE-without-tools at 41.2% sitting more than 20 points under Fable 5. The catches are the ones that follow every Beijing-lab release: no published safety or system card at launch, a knowledge cutoff that predates most of 2026, and a benchmark table that wins decisively on taste and front-end output while conceding the deep-reasoning and tool-scheduling crowns to the closed frontier.</p><h2><strong>What&#8217;s new</strong></h2><p>GLM 5.2 isn&#8217;t a new base model, it&#8217;s a strong iteration on the GLM-4.x MoE family with a few things that genuinely separate it from both its predecessor and the closed frontier.</p><ul><li><p><strong>Design as the headline capability.</strong> Most labs treat front-end output as a side effect of coding ability. Z.ai trained for it directly and led the launch with it. Topping WebDev Arena and Design Arena over Fable 5 and Opus 4.8 is the sharpest differentiator in the release, and it&#8217;s the rare benchmark win that maps cleanly to something a user sees on the first prompt instead of a leaderboard they have to trust.</p></li><li><p><strong>A genuinely permissive license.</strong> GLM 5.2 ships under plain MIT, no monthly-active-user ceiling, no revenue threshold, no mandatory credit line. Moonshot&#8217;s &#8220;Modified MIT&#8221; and DeepSeek&#8217;s license both carve out the largest deployers. Z.ai&#8217;s doesn&#8217;t, which makes 5.2 the most freely usable frontier-class open weight on the board.</p></li><li><p><strong>Switchable hybrid reasoning in one model.</strong> Rather than ship a separate &#8220;thinking&#8221; SKU, GLM 5.2 exposes a per-request flag that turns extended reasoning on or off. You pay for deliberation only when the task earns it, which matters at this price point because the per-token cost is already near the floor and reasoning tokens are where the bill actually grows.</p></li><li><p><strong>Harness-native from day one.</strong> The launch wires GLM 5.2 into Claude Code, Cline, Roo Code, OpenCode, and Kilo Code through an Anthropic-compatible endpoint, with the GLM Coding Plan subscription as the on-ramp. The model meets developers inside the tools they already run instead of asking them to adopt a new app.</p></li><li><p><strong>256K context at the bottom of the price chart.</strong> The window grew from GLM-4.6&#8217;s 200K to 256K while the price stayed near the floor. It&#8217;s not Fable 5&#8217;s million-token ceiling, but it covers most real repos, and it does it for a twelfth of Fable 5&#8217;s input rate.</p></li></ul><h2><strong>How and where to use it</strong></h2><p>Where it runs, what it&#8217;s actually good at, and where you&#8217;ll regret reaching for it.</p><h4><strong>Where it&#8217;s available</strong></h4><ul><li><p><a href="https://chat.z.ai/">Z.ai chat app</a> for direct use</p></li><li><p><a href="https://docs.z.ai/">Z.ai API</a> via OpenAI- and Anthropic-compatible endpoints (swap the base URL, set <code>model</code> to <code>glm-5.2)</code></p></li><li><p><a href="https://z.ai/subscribe">GLM Coding Plan</a> from $3 / month, wired into <a href="https://www.claude.com/product/claude-code">Claude Code</a>, <a href="https://cline.bot/">Cline</a>, <a href="https://roocode.com/">Roo Code</a>, <a href="https://opencode.ai/">OpenCode</a>, and <a href="https://kilocode.ai/">Kilo Code</a></p></li><li><p><a href="https://huggingface.co/zai-org">Hugging Face</a> and <a href="https://modelscope.cn/">ModelScope</a> for open weights under MIT, served via <a href="https://github.com/vllm-project/vllm">vLLM</a> or <a href="https://github.com/sgl-project/sglang">SGLang</a></p></li><li><p><a href="https://openrouter.ai/z-ai">OpenRouter</a> for multi-provider routing.</p></li></ul><h4><strong>What it&#8217;s good at</strong></h4><ul><li><p>Front-end and UI generation, full-stack scaffolding, and design-forward work where taste is the deliverable (the WebDev Arena and Design Arena wins are the proof point, and they&#8217;re the rare benchmarks that show up in the output you see first)</p></li><li><p>Landing pages, dashboards, and interactive components from a single prompt</p></li><li><p>Agentic coding inside Claude Code or Cline where the Anthropic-compatible endpoint makes it a drop-in</p></li><li><p>High-volume, cost-sensitive work where the $3 subscription and $0.15 cache-hit price turn the per-call cost into a rounding error</p></li><li><p>Anywhere a permissive MIT license and self-hosting beat &#8220;hosted by the best lab&#8221;</p></li></ul><h4><strong>What it&#8217;s bad at / shouldn&#8217;t be used for</strong></h4><ul><li><p>Hard reasoning and knowledge-dense work</p></li><li><p>Complex multi-tool scheduling and the longest-horizon agent loops</p></li><li><p>Math-critical workloads where correctness is load-bearing</p></li><li><p>Regulated or data-sovereignty-sensitive work where sending prompts to a Beijing-hosted API is a non-starter (in which case self-host the MIT weights)</p></li><li><p>Any deployment that requires a published system card before sign-off</p></li></ul><h2><strong>First impressions</strong></h2><p>The early read is launch-day hands-on from the open-community testers who jumped on the weights and the GLM Coding Plan within hours. The design story dominated the timeline.</p><h3><strong>The positives</strong></h3><p>Matt Velloso, a former VP at both Meta and Google DeepMind, is using GML 5.2 as a daily driver, claiming it finally meets his bar to do such.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/matvelloso/status/2067791546335019439&quot;,&quot;full_text&quot;:&quot;All day using GLM 5.2. Didn't miss much. First open model that passes the bar as a daily driver. Things are not going to be the same. \n\nDamn, now I want to buy some serious hardware.&quot;,&quot;username&quot;:&quot;matvelloso&quot;,&quot;name&quot;:&quot;Mat Velloso&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1822361068834004992/0hqBVUFs_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-19T02:08:17.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:88,&quot;retweet_count&quot;:115,&quot;like_count&quot;:2479,&quot;impression_count&quot;:407661,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Coming from an AI booster would be one thing, but coming from someone like Velloso is a prominent endorsement. His hardware callout also points towards an interesting trend of folks (and business!) starting to host models locally.</p><p>Someone ran a quick A/B test on a landing page and the results are worth checking out.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/nutlope/status/2067313679951941686&quot;,&quot;full_text&quot;:&quot;This model is insane at design.\n\nI asked GLM 5.2 (left) and Opus 4.8 (right) to build me a landing page and you can't even tell the difference.\n\nGLM cost $0.06 while opus cost $0.49. More than 6x cheaper while being faster + more token efficient.\n\nAnother win for open source AI.&quot;,&quot;username&quot;:&quot;nutlope&quot;,&quot;name&quot;:&quot;Hassan&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1727415759859552256/9rqaxXUR_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-17T18:29:24.000Z&quot;,&quot;photos&quot;:[{&quot;img_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!BRjb!,w_1028,c_limit,f_auto,q_auto:best,fl_progressive:steep/l_play_button_usfui2,w_88,e_colorize:0/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F__ss-rehost__tw-video-preview-13_2067309314306379776.jpg&quot;,&quot;link_url&quot;:&quot;https://t.co/FJCYROq9Lr&quot;}],&quot;quoted_tweet&quot;:{&quot;full_text&quot;:&quot;Introducing GLM-5.2: Frontier Intelligence, Open Weights\n\n- Significant improvements in coding and agentic tasks\n- Strong long-horizon capabilities with a 1M context window\n- Two levels of reasoning effort: GLM-5.2 (max) pushes the limits, while GLM-5.2 (high) strikes a strong&quot;,&quot;username&quot;:&quot;Zai_org&quot;,&quot;name&quot;:&quot;Z.ai&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1970775077181411328/W8XKaUIh_normal.jpg&quot;},&quot;reply_count&quot;:329,&quot;retweet_count&quot;:539,&quot;like_count&quot;:8029,&quot;impression_count&quot;:1421889,&quot;expanded_url&quot;:null,&quot;video_url&quot;:&quot;https://video.twimg.com/amplify_video/2067309314306379776/vid/avc1/1280x720/d3mHwCBsS91_kqsa.mp4&quot;,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>The price here is the real stinger, with a (debately better) landing page coming out of GLM 5.2 for $0.06 to Opus 4.8&#8217;s $0.49.</p><p>Riley Brown, an AI nut and self-described open model skeptic, compares the launch to DeepSeek R1 in how he expects frontier labs to react.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/rileybrown/status/2067932153875169296&quot;,&quot;full_text&quot;:&quot;Spent a lot of time using GLM 5.2. \n\nI&#8217;ve always been skeptical of the Open Models, as they&#8217;ve never lived up to the benchmarks and announcements. \n\nThis is the first model that passes the vibe check.\n\nThis feels like a DeepSeek R1 moment that will push the frontier labs into&quot;,&quot;username&quot;:&quot;rileybrown&quot;,&quot;name&quot;:&quot;Riley Brown&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1898571530956873728/JALEVTSb_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-19T11:27:00.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:109,&quot;retweet_count&quot;:74,&quot;like_count&quot;:1765,&quot;impression_count&quot;:112261,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>It&#8217;s not a bad bet, especially if the notorious &#8220;model feel&#8221; of GLM 5.2 makes it a breeze to use for design and coding. As open-weight models get closer to frontier performance for a fraction of the cost (whether this is due to Chinese energy efficiency or just massive subsidization is anyone&#8217;s guess), said frontier labs will start feeling the pressure to create powerful models for cheaper.</p><h3><strong>The negatives</strong></h3><p>Max Weinbach weighed Anthropic&#8217;s subsidies versus Z.ai&#8217;s and found that Claude came out on top.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/mweinbach/status/2068320723530129522&quot;,&quot;full_text&quot;:&quot;GLM 5.2 is great but my $200/m Codex sub still gets me more usage for $200 of GLM 5.2 API usage \n\nIt&#8217;s a great supplemental model for some things like working on UI/UX, but not a replacement by any means&quot;,&quot;username&quot;:&quot;mweinbach&quot;,&quot;name&quot;:&quot;Max Weinbach&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1808924794718482432/O-PiVxka_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-20T13:11:02.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:11,&quot;retweet_count&quot;:3,&quot;like_count&quot;:141,&quot;impression_count&quot;:10494,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>This is a story that Anthropic (and OpenAI) need to be publicizing more if they want to keep up with the market. More usage is more usage, regardless of a model&#8217;s literal underlying token cost.</p><p>Raycast&#8217;s Thomas Paul Mann points out the obvious.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/thomaspaulmann/status/2068257780163612834&quot;,&quot;full_text&quot;:&quot;Biggest downer of GLM-5.2 is that it doesn&#8217;t support images as input &#128546;&quot;,&quot;username&quot;:&quot;thomaspaulmann&quot;,&quot;name&quot;:&quot;Thomas Paul Mann&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/1250107714732273664/xgwMUSrS_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-20T09:00:55.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:3,&quot;retweet_count&quot;:0,&quot;like_count&quot;:41,&quot;impression_count&quot;:4177,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Given the model&#8217;s focus, it makes sense, but you&#8217;ve gotta wonder if these types of models are going to try gunning for the whole pie at some point.</p><p>Podcaster Ben Davis rejects the Anthropic/OpenAI equality statements.</p><div class="twitter-embed" data-attrs="{&quot;url&quot;:&quot;https://x.com/davis7/status/2067867580686389496&quot;,&quot;full_text&quot;:&quot;Been using GLM-5.2 tonight through Droid, it's a great model\n\nStill clearly not Claude/GPT level, but by far the closest I've felt\n\nIt's very capabale, the big differences now are just in the little things it misses (had one where it accidently fucked up a db query by doing a&quot;,&quot;username&quot;:&quot;davis7&quot;,&quot;name&quot;:&quot;Ben Davis&quot;,&quot;profile_image_url&quot;:&quot;https://pbs.substack.com/profile_images/2052289378706423808/g5qAyVi6_normal.jpg&quot;,&quot;date&quot;:&quot;2026-06-19T07:10:25.000Z&quot;,&quot;photos&quot;:[],&quot;quoted_tweet&quot;:{},&quot;reply_count&quot;:15,&quot;retweet_count&quot;:2,&quot;like_count&quot;:201,&quot;impression_count&quot;:14446,&quot;expanded_url&quot;:null,&quot;video_url&quot;:null,&quot;video_preview_media_key&quot;:null,&quot;belowTheFold&quot;:true}" data-component-name="Twitter2ToDOM"></div><p>Sometimes models perform well on benchmarks but don&#8217;t feel quite great. I wouldn&#8217;t be surprised if GLM 5.2 falls into this category once the hype wears out.</p><h2><strong>Jake&#8217;s take</strong></h2><p>I&#8217;ve been waiting for an open-weights model to take front-end seriously instead of treating it as a coding by-product. GLM 5.2 is the one. The landing pages, dashboards, and full-stack scaffolds I build in Cursor and Claude Code are exactly the workload WebDev Arena is measures, and a model that (supposedly) beats Fable 5 and Opus 4.8 on taste while costing a twelfth of Fable&#8217;s rate is dramatic. The Anthropic-compatible endpoint makes testing and implementation easy. If the output holds up past the first screen the way social media is saying it does says it does,  I can comfortably shift a chunk of my UI generation off Claude and onto a three-dollar subscription (and feel a little guilty about how good a deal it is).</p><p>Worth noting: the same benchmark table that wins design loses hard reasoning by more than twenty points on HLE, so the second you hand GLM 5.2 a problem that&#8217;s coding wrapped around real domain expertise, you shouldn&#8217;t expect much. WebDev Arena rewards the best first screen, not the best app, so I&#8217;m holding the &#8220;best front-end model&#8221; headline at arm&#8217;s length until I&#8217;ve watched it survive five rounds of edits. </p><p>GLM 5.2 is the best-looking model open-weights model to date and the cheapest way to get a beautiful first draft. It&#8217;s also further evidence that the open-weights models, specifically the ones out of China, are rapidly approaching the frontier.</p>]]></content:encoded></item><item><title><![CDATA[Does your mom use AI?]]></title><description><![CDATA[Polling my family and friends to see how and when they use AI]]></description><link>https://handyai.substack.com/p/does-your-mom-use-ai</link><guid isPermaLink="false">https://handyai.substack.com/p/does-your-mom-use-ai</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Wed, 17 Jun 2026 14:54:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!euQf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!euQf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!euQf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!euQf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!euQf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!euQf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!euQf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:94733,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!euQf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!euQf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!euQf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!euQf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e96b64e-2e85-4703-8991-ffc94e7b099e_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I work in tech, so naturally my bubble is full of people using AI for damn near everything. Requirements drafting, building prototypes, sending emails; pretty much every single person in my immediate professional circle treats AI as a utility.</p><p>I always want to be checking my biases. With my usage of platforms like Cursor and Claude consuming a numerous amount of my daily tasks, AI is a place I want to keep a level head about as the world grapples with the societal, economical, and environmental effects of this aggressive productivity.</p><p>So naturally I decided to bother my friends and family.</p><h2>My core findings</h2><p>I set out to identify exactly how my circle of humans is using AI technology. The questions I wanted to ask are simple.</p><ol><li><p>Do you use AI?</p></li><li><p>If yes, how often?</p></li><li><p>How do you use AI professionally?</p></li><li><p>How do you use AI personally?</p></li><li><p>What AI service do you typically reach for?</p></li></ol><p>And that&#8217;s it. Here&#8217;s the core of it. Do you use AI?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!LSdN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LSdN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png 424w, https://substackcdn.com/image/fetch/$s_!LSdN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png 848w, https://substackcdn.com/image/fetch/$s_!LSdN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png 1272w, https://substackcdn.com/image/fetch/$s_!LSdN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LSdN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png" width="1456" height="664" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:664,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:77877,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LSdN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png 424w, https://substackcdn.com/image/fetch/$s_!LSdN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png 848w, https://substackcdn.com/image/fetch/$s_!LSdN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png 1272w, https://substackcdn.com/image/fetch/$s_!LSdN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F73f7885d-be80-46f5-ad97-e9f1c4a2ed48_1544x704.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I wasn&#8217;t surprised that most said they use AI. I&#8217;ve seen enough broad usage metric data on this for it to be obvious that AI touches nearly everyone&#8217;s life at some point. What did surprise me is the frequency and type of usage.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QCM6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QCM6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png 424w, https://substackcdn.com/image/fetch/$s_!QCM6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png 848w, https://substackcdn.com/image/fetch/$s_!QCM6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png 1272w, https://substackcdn.com/image/fetch/$s_!QCM6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QCM6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png" width="1456" height="687" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:687,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:47277,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QCM6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png 424w, https://substackcdn.com/image/fetch/$s_!QCM6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png 848w, https://substackcdn.com/image/fetch/$s_!QCM6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png 1272w, https://substackcdn.com/image/fetch/$s_!QCM6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F764d6e80-15f9-4810-876e-a68b80efd06b_1544x728.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yg8s!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yg8s!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png 424w, https://substackcdn.com/image/fetch/$s_!yg8s!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png 848w, https://substackcdn.com/image/fetch/$s_!yg8s!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png 1272w, https://substackcdn.com/image/fetch/$s_!yg8s!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yg8s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png" width="1456" height="1150" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1150,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:97191,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yg8s!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png 424w, https://substackcdn.com/image/fetch/$s_!yg8s!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png 848w, https://substackcdn.com/image/fetch/$s_!yg8s!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png 1272w, https://substackcdn.com/image/fetch/$s_!yg8s!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4521ad28-154f-475a-a37f-a052f7c0ce00_1544x1220.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I had expected a significant chunk of my circle to use AI indirectly via things like Google&#8217;s AI Search results and intelligent features baked into non-AI products. However, this was a dramatic minority. Most people around me, both active supporters <em>and </em>critics of the technology, are using it daily and extensively (and yes, my mom uses AI).</p><p>On top of the general usage, I wanted specifics on where they were using AI in their day-to-day.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K7Lg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K7Lg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png 424w, https://substackcdn.com/image/fetch/$s_!K7Lg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png 848w, https://substackcdn.com/image/fetch/$s_!K7Lg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png 1272w, https://substackcdn.com/image/fetch/$s_!K7Lg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K7Lg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png" width="1456" height="664" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:664,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:85816,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K7Lg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png 424w, https://substackcdn.com/image/fetch/$s_!K7Lg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png 848w, https://substackcdn.com/image/fetch/$s_!K7Lg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png 1272w, https://substackcdn.com/image/fetch/$s_!K7Lg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F67713752-1fa0-4737-a981-0cc9895cb9b2_1544x704.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nIv0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nIv0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png 424w, https://substackcdn.com/image/fetch/$s_!nIv0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png 848w, https://substackcdn.com/image/fetch/$s_!nIv0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png 1272w, https://substackcdn.com/image/fetch/$s_!nIv0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nIv0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png" width="1456" height="1245" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1245,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:147273,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nIv0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png 424w, https://substackcdn.com/image/fetch/$s_!nIv0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png 848w, https://substackcdn.com/image/fetch/$s_!nIv0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png 1272w, https://substackcdn.com/image/fetch/$s_!nIv0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F361998b4-df9f-4271-be70-58a2a70cd853_1544x1320.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There seemed to be a clear pattern toward productivity work with these tools, though the usage is pretty evenly spread across personal and professional. I wasn&#8217;t surprised to see creative tasks so low on the list, given my proximity to media industries, but <em>was </em>surprised that coding and building was down in fifth.</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/p/does-your-mom-use-ai?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading Handy AI! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/p/does-your-mom-use-ai?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://handyai.substack.com/p/does-your-mom-use-ai?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><h2>Breaking it down</h2><p>Alongside the responses I tracked some anonymized demographical information: namely gender, perceived political ideology (based on my conversations with them), and job function and industry. I try to keep a balanced circle of folks around me but, unsurprisingly, the majority work within tech or tech-adjacent industries that are typically heavier on AI use. From a professional perspective, the industry splits looked like this:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K3Wj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K3Wj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png 424w, https://substackcdn.com/image/fetch/$s_!K3Wj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png 848w, https://substackcdn.com/image/fetch/$s_!K3Wj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png 1272w, https://substackcdn.com/image/fetch/$s_!K3Wj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K3Wj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png" width="1456" height="1294" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1294,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:160409,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K3Wj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png 424w, https://substackcdn.com/image/fetch/$s_!K3Wj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png 848w, https://substackcdn.com/image/fetch/$s_!K3Wj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png 1272w, https://substackcdn.com/image/fetch/$s_!K3Wj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fead41b60-bbbc-42b2-8472-9b559ab180d1_1544x1372.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>And the specific job function splits as such:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fuGP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fuGP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png 424w, https://substackcdn.com/image/fetch/$s_!fuGP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png 848w, https://substackcdn.com/image/fetch/$s_!fuGP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png 1272w, https://substackcdn.com/image/fetch/$s_!fuGP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fuGP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png" width="1456" height="1120" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1120,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:133761,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fuGP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png 424w, https://substackcdn.com/image/fetch/$s_!fuGP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png 848w, https://substackcdn.com/image/fetch/$s_!fuGP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png 1272w, https://substackcdn.com/image/fetch/$s_!fuGP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37300121-af5f-446b-813e-a7877e8f33b5_1544x1188.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Software and product are at the top here, which I expected giving my own personal industry and job function. The rest is fairly run-of-the-mill as well. Operations, content, strategy; these are functions heavily saturated with tasks ripe for automation with tools like Claude Cowork. They&#8217;re data-oriented and repetitive. Meanwhile, areas like skilled trades and hospitality being uninterested in the tech makes sense (though my sample size there is tiny). </p><p>From a gender perspective, here&#8217;s where it landed:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pTNv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pTNv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png 424w, https://substackcdn.com/image/fetch/$s_!pTNv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png 848w, https://substackcdn.com/image/fetch/$s_!pTNv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png 1272w, https://substackcdn.com/image/fetch/$s_!pTNv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pTNv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png" width="1456" height="426" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:426,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:44386,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pTNv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png 424w, https://substackcdn.com/image/fetch/$s_!pTNv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png 848w, https://substackcdn.com/image/fetch/$s_!pTNv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png 1272w, https://substackcdn.com/image/fetch/$s_!pTNv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F12446b3e-4fdb-4151-9237-e33653cb09bd_1544x452.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Nothing remarkable and fairly evenly split across the represented genders. Mostly this data is telling me that I need to widen my non-male circles.</p><p>Finally, I wanted to get a little spicy and see the breakdowns based on my own perception of their political ideologies and religious affiliations. My own biases are going to be baked into these perceptions (hello, right-wing blind spot), but I wanted to capture some sort of ideological split without having to ask the group to share specifics on their current beliefs.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JQlV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JQlV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png 424w, https://substackcdn.com/image/fetch/$s_!JQlV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png 848w, https://substackcdn.com/image/fetch/$s_!JQlV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png 1272w, https://substackcdn.com/image/fetch/$s_!JQlV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JQlV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png" width="1456" height="513" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:513,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:52769,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JQlV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png 424w, https://substackcdn.com/image/fetch/$s_!JQlV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png 848w, https://substackcdn.com/image/fetch/$s_!JQlV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png 1272w, https://substackcdn.com/image/fetch/$s_!JQlV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b3ecffb-255c-4d02-adaa-8b7abb42bb98_1544x544.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rqk-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rqk-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png 424w, https://substackcdn.com/image/fetch/$s_!rqk-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png 848w, https://substackcdn.com/image/fetch/$s_!rqk-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png 1272w, https://substackcdn.com/image/fetch/$s_!rqk-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rqk-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png" width="1456" height="339" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:339,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38851,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rqk-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png 424w, https://substackcdn.com/image/fetch/$s_!rqk-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png 848w, https://substackcdn.com/image/fetch/$s_!rqk-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png 1272w, https://substackcdn.com/image/fetch/$s_!rqk-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac18eac1-815c-4417-8b72-b387b2d24c16_1544x360.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p>My main takeaways from these are:</p><ul><li><p>&#8220;Center Left&#8221; folks are overly represented within daily AI use compared to other perceived ideological affiliations. This makes sense for my circle of white collar, center left tech workers.</p></li><li><p>Both cases of non-AI usage come from those more left-leaning, which makes sense with the rising <a href="https://handyai.substack.com/p/are-ai-data-centers-stealing-all">environmental</a> and societal concerns. </p></li><li><p>Religiousness seems to not correlate at all with AI usage. It&#8217;s a wash.</p></li></ul><p>While tracking these types of stat points are fun, my overall takeaway from them is that in order to run proper analysis on these types of sensitive convictions, I&#8217;ll need a bigger dataset and more accurate datapoints. Next time.</p><h2>Some other tidbits</h2><p>Any good data analyst knows that the cross-analysis is the fun part. While my dataset here is a bit small, I wanted to be sure to include some of this to try and identify any meaningful patterns. Here&#8217;s what I found.</p><h4>Which tool wins, by how they know me</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xUtS!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xUtS!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png 424w, https://substackcdn.com/image/fetch/$s_!xUtS!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png 848w, https://substackcdn.com/image/fetch/$s_!xUtS!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png 1272w, https://substackcdn.com/image/fetch/$s_!xUtS!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xUtS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png" width="1456" height="426" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:426,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:44907,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xUtS!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png 424w, https://substackcdn.com/image/fetch/$s_!xUtS!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png 848w, https://substackcdn.com/image/fetch/$s_!xUtS!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png 1272w, https://substackcdn.com/image/fetch/$s_!xUtS!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9825c96e-88f5-4669-a879-f1aaef19d98f_1544x452.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I&#8217;ve been trying to spread the good news of Claude and Cursor for a long time now, so it&#8217;s nice to see significant Claude usage in my circle. The ChatGPT usage is highly disproportionate among my friends compared to coworkers and family, which most likely speaks to its ubiquity for casual use among people in my age range.</p><h4>What predicts heavy use?</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ervc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ervc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png 424w, https://substackcdn.com/image/fetch/$s_!ervc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png 848w, https://substackcdn.com/image/fetch/$s_!ervc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png 1272w, https://substackcdn.com/image/fetch/$s_!ervc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ervc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png" width="1456" height="962" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:962,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:110224,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201170759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ervc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png 424w, https://substackcdn.com/image/fetch/$s_!ervc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png 848w, https://substackcdn.com/image/fetch/$s_!ervc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png 1272w, https://substackcdn.com/image/fetch/$s_!ervc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5a05a6fb-a9ec-47f1-8f75-92bc54929950_1544x1020.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Proximity to me (and, by extension, to tech) seems to predict daily use far more than politics, faith, or gender, which barely move the needle. 100% of my polled coworkers are using AI daily, compared to groups like my family (25%) and women (38%).</p><div><hr></div><p>The dataset here isn&#8217;t perfect but its intimate to me. When I started out on collecting this data, I wasn&#8217;t necessarily looking to accurately identify broad AI usage metrics. I simply wanted to know what the spread looks like among those close to me.</p><p>Overall I&#8217;m surprised. Given the conversations I have with these folks on a regular basis, I would have expected much less daily AI usage. Reason stands that, even for someone like myself who&#8217;s writing about this industry regularly, I&#8217;m underestimating just how prevalent this new tech has become in such a short amount of time.</p><div class="callout-block" data-callout="true"><p><em>You can view all of the charts here, alongside the full dataset, on my website:</em></p><p><em><a href="https://jakehandy.com/ai-survey">https://jakehandy.com/ai-survey</a></em></p></div><p></p>]]></content:encoded></item><item><title><![CDATA[The US government bans Claude]]></title><description><![CDATA[AI Weekly Update - June 15, 2026]]></description><link>https://handyai.substack.com/p/the-us-government-bans-claude</link><guid isPermaLink="false">https://handyai.substack.com/p/the-us-government-bans-claude</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Mon, 15 Jun 2026 17:00:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rJ9b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: center;"></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rJ9b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rJ9b!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!rJ9b!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!rJ9b!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!rJ9b!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rJ9b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:67561,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/202157126?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rJ9b!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!rJ9b!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!rJ9b!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!rJ9b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff9a2cbea-673b-49ca-b384-ac331148a9df_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="callout-block" data-callout="true"><p style="text-align: center;"><strong>&#128047; Before we get started</strong></p><p>Right now a data center is being planned within shouting distance from the Nashville Zoo. While water and energy concerns around data centers <a href="https://substack.com/home/post/p-200563358">are typically overblown</a>, their sound pollution and growing CO2 emissions directly affect our world and animals, making construction near a zoo a nonstarter.</p><p>I urge you to <a href="https://www.change.org/p/nashville-zoo-says-no-to-proposed-data-center">sign the petition</a> to prevent the Nashville Zoo build, and keep an eye out for unwise construction plans in your local community (data center or otherwise).</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.change.org/p/nashville-zoo-says-no-to-proposed-data-center&quot;,&quot;text&quot;:&quot;Sign the petition now!&quot;,&quot;action&quot;:null,&quot;class&quot;:&quot;button-wrapper&quot;}" data-component-name="ButtonCreateButton"><a class="button primary button-wrapper" href="https://www.change.org/p/nashville-zoo-says-no-to-proposed-data-center"><span>Sign the petition now!</span></a></p></div><h1><strong>what to know for now</strong></h1><p>&#128214; <strong>Anthropic shipped Claude Fable 5, the first &#8220;Mythos-class&#8221; model.</strong> Fable 5 (<code>claude-fable-5</code>) is Anthropic&#8217;s most capable public model to date, claiming state-of-the-art results across software engineering, knowledge work, vision, and long-horizon agentic tasks, with a 1M-token context window and a reported 80.3% on SWE-Bench Pro against 69% for Opus 4.8. The catch is the safety architecture: in high-risk domains like cyber, bio, and chemistry it refuses and silently falls back to Opus 4.8, a downgrade Anthropic says fires in under 5% of sessions. Pricing lands at $10/$50 per million tokens, roughly double Opus.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;c249bb64-a750-4946-b4d3-82699e8ba551&quot;,&quot;caption&quot;:&quot;Fable, you say? But where&#8217;s Mythos? Fable 5 is the first Mythos-class model the public has ever been allowed to touch; Mythos 5 is the same model with the muzzle off, and reserved for vetted defenders. Together they&#8217;re the most capable AI models in existence (and the first one that comes with a designated stunt double).&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: Fable 5 / Mythos 5&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-06-09T17:47:07.013Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!8d5Z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbd7d24d3-bdc1-45a4-89a0-091c21fb5ad0_1456x1048.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-fable-5-mythos-5&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:201335111,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:13,&quot;comment_count&quot;:0,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>&#9940; <strong>Then the government made Anthropic pull it 72 hours later.</strong> Here&#8217;s the saga in one breath: Fable 5 launched June 9; by June 10 developers caught it quietly nerfing itself when it sensed frontier-AI work, and Microsoft yanked it from internal Copilot over data-retention terms; on June 12 a red-teamer handed the government a jailbreak, and at 5:21 PM ET Anthropic got an export-control directive ordering it to cut off all foreign nationals (including its own foreign-national staff), so it disabled Fable 5 and Mythos 5 worldwide. Anthropic publicly disagrees, calling the issue a narrow, non-universal jailbreak (essentially &#8220;ask the model to find bugs in code,&#8221; a thing GPT-5.5 does too) and warning that enforcing that standard would &#8220;essentially halt all new model deployments.&#8221; As of today, both models are still dark and Anthropic says it&#8217;s &#8220;working to restore&#8221; access. <a href="https://www.anthropic.com/news/fable-mythos-access">Read more</a></p><p>&#127822; <strong>Apple gave up and handed Siri to Google.</strong> At WWDC on June 8 Apple unveiled &#8220;Siri AI,&#8221; the biggest overhaul since launch, and confirmed the long-rumored Gemini partnership: a reported ~$1B/year deal to run a custom Apple-tuned Gemini model (&#8221;AFM Cloud Pro&#8221;) on Google&#8217;s cloud for the heavy reasoning its on-device and Private Cloud Compute tiers can&#8217;t handle, shipping this fall in iOS 27. Developers got the real gift: a hybrid Foundation Models framework with a new <code>LanguageModel</code> Swift protocol that lets apps call Claude, Gemini, or Apple&#8217;s own models through one API with no code changes, plus free on-device model access for apps under 2M downloads. <a href="https://www.macrumors.com/2026/06/09/apple-outlines-major-ai-and-developer-tool-updates/">Read more</a></p><p>&#128752;&#65039; <strong>SpaceX went public, and the market is being asked to price it as an AI company.</strong> The IPO priced after the close June 11 and started trading June 12 on Nasdaq under SPCX, raising roughly $75B at around a $1.75 trillion valuation (Bloomberg reported the target crept above $2T). It&#8217;s the whole company, not a Starlink spinout, pitched across three legs (space, connectivity, and AI) with Starlink throwing off two-thirds of revenue and the speculative upside resting on Starship and the orbital data centers SpaceX has filed to build (up to a million compute satellites, 1 GW of orbital compute by year-end). Recall that last month Anthropic rented all 220,000 GPUs in SpaceX&#8217;s Colossus 1. <a href="https://www.scientificamerican.com/article/spacex-ipo-valuation-depends-on-starship-and-orbital-ai-data-centers/">Read more</a></p><p>&#127769; <strong>Moonshot dropped Kimi K2.7 Code and kept the open-weight pressure on.</strong> Released June 12 on Hugging Face under a modified-MIT license, K2.7 Code is a 1T-parameter MoE (32B active) with a 256K context, claiming double-digit gains over K2.6 on its own coding benchmarks and roughly 30% fewer reasoning tokens per task, at $0.95/$4.00 per million tokens. The asterisk: every number is from Moonshot&#8217;s proprietary benchmarks, with no SWE-bench Verified score at launch, so treat &#8220;it&#8217;s great&#8221; as a vendor claim until third parties run it. (I did a full Model Drop on this one if you want the deeper read.) The strategy is the story regardless: cheap, downloadable, commercially usable agentic coding aimed straight at the closed labs&#8217; wallets.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;5cdd8381-a328-4c3c-98a5-2e8339c518ce&quot;,&quot;caption&quot;:&quot;Kimi K2.7 (also branded \&quot;K2.7 Code\&quot;) is Moonshot AI&#8217;s next iteration on the popular open-weight coding model. This is the first time Moonshot has put \&quot;Code\&quot; in the name, and the whole release is pitched as the budget answer to Fable 5's $10 / $50 rate card.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Model Drop: Kimi K2.7 Code&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:135858492,&quot;name&quot;:&quot;Jake Handy&quot;,&quot;bio&quot;:&quot;R&amp;D product manager, Classical violinist, AI wrangler&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d93fce1-7ab8-4136-b275-3467dc2b8266_1042x1040.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-06-12T14:38:54.128Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!Zs5V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://handyai.substack.com/p/model-drop-kimi-k27-code&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:201751036,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:1,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1517851,&quot;publication_name&quot;:&quot;Handy AI&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!AvoP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F69bde06d-cde1-4eb7-a7ac-a175aac1a630_1280x1280.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><div><hr></div><h2><strong>&#129514; AI Research of the Week</strong></h2><p><strong><a href="https://newsnetwork.mayoclinic.org/discussion/mayo-clinic-ai-detects-pancreatic-cancer-up-to-3-years-before-diagnosis-in-landmark-validation-study/">REDMOD: detecting pancreatic cancer on routine CT scans years early</a></strong><br><em>From Mayo Clinic</em></p><p><strong>Jake&#8217;s Take:</strong> Mayo built REDMOD, a model that reads the routine abdominal CT scans people already get for unrelated reasons and pulls out hundreds of &#8220;radiomics&#8221; features, subtle texture and structure changes in pancreatic tissue that a radiologist can&#8217;t catch by eye. In the validation study it flagged 73% of pancreatic cancers before clinical diagnosis, against 39% for radiologists reading the exact same scans, at a median of about 16 months early and up to three years out in some cases. They trained it on roughly 2,000 scans across different hospitals and scanner types so it isn&#8217;t just memorizing one machine&#8217;s quirks.</p><p>Pancreatic cancer&#8217;s five-year survival sits around 13% because it&#8217;s almost always caught late, and this is AI detection available on the imaging patients are already getting. No new scan, no new cost, no manual annotation. It&#8217;s not FDA-cleared and it&#8217;s not deployed yet (the follow-up AI-PACED trial is now testing it prospectively in higher-risk patients). It&#8217;s a landmark study, not a product. But a 16-month head start on the cancer that kills fastest is nothing but good news. <a href="https://www.insideprecisionmedicine.com/topics/oncology/mayo-clinics-redmod-ai-doubles-early-detection-sensitivity-in-pancreatic-cancer/">Read more</a></p><div><hr></div><h1><strong>what to know for later</strong></h1><p>&#9878;&#65039; <strong>Dario Amodei reversed himself and called for binding AI regulation, two days before the government banned his model.</strong> In a ~5,000-word essay, &#8220;Policy on the AI Exponential&#8221; (June 10), Amodei argued it&#8217;s time to go &#8220;beyond transparency to more serious and binding regulation&#8221;: mandatory third-party testing for models above a compute threshold and federal power to block or reverse deployments that fail audits across cyber, bioweapons, loss-of-control, and autonomous-R&amp;D risk, an FAA/FDA-style regime. Anthropic paired it with a $350M commitment to study AI&#8217;s labor-market hit. The timing is almost too perfect &#8212; Amodei asks for a government empowered to pull dangerous models, and 48 hours later that government pulls his; Sam Altman called the whole thing &#8220;fear-based marketing.&#8221; <a href="https://darioamodei.com/post/policy-on-the-ai-exponential">Read more</a></p>
      <p>
          <a href="https://handyai.substack.com/p/the-us-government-bans-claude">
              Read more
          </a>
      </p>
   ]]></content:encoded></item><item><title><![CDATA[Model Drop: Kimi K2.7 Code]]></title><description><![CDATA[Moonshot updates the world's best open source coding model]]></description><link>https://handyai.substack.com/p/model-drop-kimi-k27-code</link><guid isPermaLink="false">https://handyai.substack.com/p/model-drop-kimi-k27-code</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Fri, 12 Jun 2026 14:38:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Zs5V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Zs5V!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Zs5V!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!Zs5V!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!Zs5V!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!Zs5V!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Zs5V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:56648,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/201751036?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Zs5V!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!Zs5V!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!Zs5V!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!Zs5V!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ead54d3-27b2-4131-8b3d-d6299ecfc733_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Kimi K2.7 (also branded "K2.7 Code") is <a href="https://www.moonshot.ai/">Moonshot AI</a>&#8217;s next iteration on the popular open-weight coding model. This is the first time Moonshot has put "Code" in the name, and the whole release is pitched as the budget answer to Fable 5's $10 / $50 rate card.</p><div class="callout-block" data-callout="true"><p><strong>Model</strong>: Kimi K2.7 (<code>kimi-k2.7-code</code> on the Moonshot API, <code>moonshotai/Kimi-K2.7-Code</code> on Hugging Face). Often written &#8220;Kimi 2.7&#8221; or &#8220;K2.7 Code.&#8221;</p><p><strong>Model type</strong>: Text in, text out, tuned for agentic code generation and long-horizon software engineering</p><p><strong>Ship date</strong>: June 12, 2026</p><p><strong>Maker</strong>: <a href="https://www.moonshot.ai/">Moonshot AI</a> (Beijing)</p><p><strong>Pricing</strong>: $0.95 per million input tokens, $4.00 per million output tokens, and $0.19 per million on cache hits, on the <a href="https://platform.moonshot.ai/">Moonshot API</a>. Free weights on Hugging Face for self-hosting.</p><p><strong>Available on</strong>: The <a href="https://platform.moonshot.ai/">Moonshot API</a> (OpenAI- and Anthropic-SDK compatible, one-line base URL swap), <a href="https://www.kimi.com/code">Kimi Code</a> (the terminal and IDE coding agent the K2 family ships through), and <a href="https://huggingface.co/moonshotai">Hugging Face</a> for open weights under the Modified MIT license.</p><p><strong>Headline benchmarks</strong>: Moonshot published K2.7&#8217;s gains as deltas over K2.6 rather than clean head-to-head numbers against the frontier: +21.8% on Kimi Code Bench v2 (Moonshot&#8217;s in-house coding eval), +11% on Program Bench, and +31.5% on MLS Bench Lite (multi-language support across Python, Rust, and Go). The other headline is efficiency: a claimed ~30% reduction in reasoning-token usage versus K2.6 on the same tasks.</p><p><strong>Other info</strong>: Mixture-of-experts, 1 trillion total parameters with 32B active per token, the same architecture family as K2.5 and K2.6 (so existing deployments swap weights without reconfiguring the inference stack). 262,144-token (262K) context window, carried over from K2.6, with automatic context compression for sustained long-horizon sessions. License: Modified MIT (free commercial use; visible &#8220;Kimi K2.7&#8221; credit required for products above ~100M monthly users or ~$20M/month revenue). No system card published at launch. The &#8220;30% fewer reasoning tokens&#8221; claim is positioned as a fix for &#8220;overthinking.&#8221;</p><p><strong>More details</strong>: <strong><a href="https://platform.moonshot.ai/">Kimi platform</a></strong></p></div><h2><strong>What shipped</strong></h2><p>Moonshot AI dropped Kimi K2.7 this morning as another open-weight, coding-specialist iteration on the K2 family. It&#8217;s the same 1T / 32B-active mixture-of-experts architecture as K2.5 and K2.6, the same 262K context window, the same Modified MIT license, and the same one-line base-URL swap to point an existing OpenAI- or Anthropic-SDK client at the Moonshot endpoint. The pitch is narrow and explicit: get most of the frontier&#8217;s coding capability at roughly a twelfth of the token price, with a model specifically tuned to stop burning reasoning tokens.</p><p>The evidence Moonshot put forward is unusual, because it&#8217;s almost entirely relative to its own predecessor rather than to the frontier. The launch numbers are framed as gains over K2.6: +21.8% on Kimi Code Bench v2, +11% on Program Bench, +31.5% on the multi-language MLS Bench Lite, and a ~30% cut in reasoning-token usage on equivalent tasks. Moonshot didn&#8217;t publish K2.7&#8217;s SWE-Bench Verified or SWE-Bench Pro scores against Fable 5, GPT-5.5, or Opus 4.7 at launch, and &#8220;Kimi Code Bench v2&#8221; is a benchmark only Moonshot reports. But the efficiency story is real and falsifiable.</p><h2><strong>What&#8217;s new</strong></h2><p>K2.7 isn&#8217;t a new base model. It&#8217;s a coding-tuned post-train on the K2 MoE family.</p><ul><li><p><strong>Reasoning-token efficiency as the headline feature.</strong> The ~30% reduction in reasoning-token usage over K2.6 is the first time a Moonshot release led with efficiency instead of capability. &#8220;Overthinking&#8221; (models spending thousands of tokens deliberating on problems that need a few) is a real cost (monetarily and environmentally) and latency tax in production agent loops, and a model that solves the same task with a third fewer thinking tokens is directly cheaper to run (on top of already being cheap per token).</p></li><li><p><strong>Multi-language coding gains.</strong> The +31.5% on MLS Bench Lite is the largest single delta Moonshot published, and it targets the exact weakness independent reviewers flagged on K2.6: solid Python, shakier Rust and Go.</p></li><li><p><strong>A named coding SKU.</strong> Moonshot has always pointed the K2 family at coding, but &#8220;K2.7 Code&#8221; is the first time the coding specialization is in the model name rather than a console preview flag. It signals Moonshot is going to keep a coding-optimized line distinct from a general-agent line.</p></li></ul><h2><strong>How and where to use it</strong></h2><h4><strong>Where it&#8217;s available</strong></h4><ul><li><p>The <a href="https://platform.moonshot.ai/">Moonshot API</a> via OpenAI- and Anthropic-compatible endpoints (swap the base URL, set <code>model</code> to <code>kimi-k2.7-code</code>)</p></li><li><p><a href="https://www.kimi.com/code">Kimi Code</a> for the terminal and IDE coding agent</p></li><li><p><a href="https://huggingface.co/moonshotai">Hugging Face</a> for open weights under the Modified MIT license, served via <a href="https://github.com/vllm-project/vllm">vLLM</a> or <a href="https://github.com/sgl-project/sglang">SGLang</a>; Multi-provider routers (OpenRouter, Fireworks, Together, DeepInfra) are not all live day-one but should follow within days given the license, the same way they did for K2.6.</p></li></ul><h4><strong>What it&#8217;s good at</strong></h4><ul><li><p>High-volume, cost-sensitive agentic coding where the reasoning-token cut compounds with the already-low per-token price</p></li><li><p>Multi-language refactors across Python, Rust, and Go, where the MLS Bench Lite gain is supposed to land. Long-horizon sessions on mid-sized codebases that fit inside 262K tokens</p></li><li><p>Routine, templated, or boilerplate-heavy work where cache hits at $0.19 per million tokens drop the effective cost into near-free territory</p></li><li><p>Anywhere open weights, a Modified MIT license, and air-gapped or self-hosted deployment beat &#8220;hosted by the best lab&#8221;</p></li></ul><h4><strong>What it&#8217;s bad at / shouldn&#8217;t be used for</strong></h4><ul><li><p>High-stakes architectural decisions and gnarly merge conflicts, where Claude Fable 5&#8217;s 90%-plus coding and analytics scores still hold a real lead and the cost difference is worth eating</p></li><li><p>Knowledge-heavy or deep-reasoning work</p></li><li><p>Workloads that need a 1M-token window</p></li><li><p>Regulated or data-sovereignty-sensitive work where sending prompts to a Beijing-hosted API is a non-starter (self-host the MIT weights, which is the point)</p></li><li><p>Anything where you need a published system card before deployment</p></li></ul><h2><strong>First impressions</strong></h2><p>Independent, hands-on evaluations don&#8217;t exist yet, so the positives lean on launch-day coverage and on carried-over sentiment from the K2 family&#8217;s daily drivers, and the negatives lean on the structural questions the launch materials leave open.</p><h3><strong>The positives</strong></h3><p><a href="https://cryptobriefing.com/kimi-2-7-affordable-coding-claude-fable-5/">CryptoBriefing</a> framed the release as a pricing event before a capability one:</p><blockquote><p><em>&#8220;The AI coding wars just got a new price leader... The pitch is straightforward: get close to the same performance for a fraction of the cost.&#8221;</em></p></blockquote><p>At $0.95 / $4.00, K2.7 is not a discount on the frontier, it&#8217;s a different budget tier, and the $0.19 cache-hit price makes templated and repetitive agent work effectively free. </p><p><a href="https://www.valuethemarkets.com/cryptocurrency/news/exploring-the-impact-of-kimi-k27-code-on-ai-programming-efficiency">Value The Markets</a> zeroed in on the efficiency angle the rest of the coverage mostly skipped:</p><blockquote><p><em>&#8220;The 30% decrease in reasoning tokens addresses a prevalent issue in automated coding, known as &#8220;overthinking.&#8221; Excessive token usage during problem-solving leads to increased latency and elevated API expenses.&#8221;</em></p></blockquote><p>A model that solves the same task with a third fewer reasoning tokens is cheaper and faster on every single call, independent of whatever the capability leaderboards eventually say.</p><p>The carried-over signal from the K2 family is the strongest &#8220;real-world&#8221; data point available on launch day. The <a href="https://news.ycombinator.com/item?id=47063429">Hacker News thread on running K2 as a daily coding driver</a> is full of developers who already switched off Opus:</p><blockquote><p><em>&#8220;Was spending crazy amounts on Claude and it was sporadic at best... Switched to Kimi K2.5 and honestly didn&#8217;t think it would do anything other than destroy my code. Crazy enough, it solved the problem I had in less than 60 seconds and I was hooked.&#8221;</em></p></blockquote><p>That sentiment predates 2.7, but the people most likely to adopt K2.7 on day one are the ones already running K2.5 or K2.6 in Kimi Code or OpenCode.</p><h3><strong>The negatives</strong></h3><p>The benchmark framing is the structural problem, and it&#8217;s the same disease in a different strain from the DeepSeek V4 release. Moonshot published K2.7&#8217;s gains as percentage deltas over K2.6 on its own benchmarks (Kimi Code Bench v2, Program Bench, MLS Bench Lite) and did not publish standard SWE-Bench Verified or SWE-Bench Pro numbers against Fable 5, Opus 4.7, or GPT-5.5. &#8220;21.8% better than our last model on our own coding eval&#8221; is a real improvement and also unfalsifiable by anyone outside Moonshot. Until a third party runs 2.7 on the public benchmarks, the only honest statement is that it&#8217;s meaningfully better than K2.6 at coding and cheaper to run.</p><p>On the head-to-head that does exist, the predecessor sets the expectation. <a href="https://benchlm.ai/compare/claude-fable-vs-kimi-2-6">BenchLM&#8217;s provisional comparison</a> of Fable 5 against K2.6 is not close:</p><blockquote><p><em>&#8220;Claude Fable 5 is clearly ahead on the provisional aggregate, 96 to 84... The biggest single separator in this matchup is HLE, where the scores are 64.5% and 34.7%.&#8221;</em></p></blockquote><p>K2.7 is a coding-focused post-train, so it should narrow the coding gap. Nothing in the launch suggests it touches the 30-point HLE gap on hard reasoning and knowledge.</p><h2><strong>Jake&#8217;s take</strong></h2><p>For high-volume, low-stakes code generation where a wrong answer costs an hour instead of a client, the price-per-capability math is lopsided in Moonshot&#8217;s favor (and it isn&#8217;t close). K2.7 sharpens exactly the edge I care about: the 30% reasoning-token cut. The thing that makes a coding agent expensive isn&#8217;t the per-token price, it&#8217;s how many tokens it burns thinking out loud before it even does anything. The multi-language gain matters too, since the one consistent complaint I&#8217;ve seen with K2.6 is that it was a Python model wearing a Rust costume.</p><p>Unfortunately Moonshot shipped a coding model and declined to show us how it codes against the models we&#8217;d actually switch from. &#8220;21.8% better than K2.6 on a benchmark only we run&#8221; isn&#8217;t helpful; the absence of a single SWE-Bench Verified line against it reads as a choice, not an oversight. </p><p>Kimi 2.7 is going to save real money and write a lot of boilerplate, and I still won&#8217;t trust a benchmark card that grades itself.</p>]]></content:encoded></item><item><title><![CDATA[You're tokenmaxxing yourself broke]]></title><description><![CDATA[How to rein in your AI habit and get the same work done on a fraction of the tokens]]></description><link>https://handyai.substack.com/p/youre-tokenmaxxing-yourself-broke</link><guid isPermaLink="false">https://handyai.substack.com/p/youre-tokenmaxxing-yourself-broke</guid><dc:creator><![CDATA[Jake Handy]]></dc:creator><pubDate>Wed, 10 Jun 2026 14:55:59 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_9tz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_9tz!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_9tz!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!_9tz!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!_9tz!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!_9tz!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_9tz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png" width="1456" height="1048" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1048,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:88184,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/200836456?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_9tz!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png 424w, https://substackcdn.com/image/fetch/$s_!_9tz!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png 848w, https://substackcdn.com/image/fetch/$s_!_9tz!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!_9tz!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe7c1421d-84f6-48c6-9078-e58242322d8d_1456x1048.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>When you send ChatGPT or Claude a message and it sends you one back, your words are metered in &#8220;tokens.&#8221; The term &#8220;tokenmaxxing&#8221; was coined recently to refer to the practice of blasting through as many tokens as possible like a teenager with their first credit card.</p><p>I&#8217;ve already gone after the supply side of this. I wrote about <a href="https://handyai.substack.com/p/inefficient-ai-models-are-killing?utm_source=publication-search">models that torch compute to sound smart</a>, and about what the buildout <a href="https://handyai.substack.com/p/are-ai-data-centers-stealing-all?utm_source=publication-search">costs the towns it lands in</a>. This one&#8217;s about the demand side. About you. Because the single biggest lever on what these systems cost, in dollars and in watts, isn&#8217;t the data center. It&#8217;s the prompt you just sent (and the seventeen you didn&#8217;t need to).</p><p>So let&#8217;s talk about how to use fewer tokens without using less AI. There&#8217;s real money and real environmental cost on the table, and the techniques to claw both back are (mostly) free.</p><h3><strong>The part where I make you feel bad</strong></h3><p>Here&#8217;s the uncomfortable physics. Every token a model reads or writes is arithmetic running on a chip in a building that drinks power and water to stay cool.</p><p>The per-query numbers sound tiny until you multiply them. OpenAI pegs an <a href="https://epoch.ai/gradient-updates/how-much-energy-does-chatgpt-use">average ChatGPT query</a> at around 0.34 watt-hours. Google says the <a href="https://blog.google/technology/google-cloud/gemini-energy-water/">median Gemini text prompt</a> lands near 0.24. Independent measurement put a short GPT-4o call at <a href="https://arxiv.org/html/2505.09598v1">about 0.42 watt-hours</a>, roughly 40% more than a Google search. On its own, nothing. But at the scale of hundreds of daily, robust reasoning requests, the carbon emissions from our AI use begins to look similar to our daily commutes. (And with Anthropic&#8217;s release of Fable 5 yesterday, tokens-per-query for reasoning models just keep going up and up).</p><div class="captioned-button-wrap" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/p/youre-tokenmaxxing-yourself-broke?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="CaptionedButtonToDOM"><div class="preamble"><p class="cta-caption">Thanks for reading Handy AI! This post is public so feel free to share it.</p></div><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/p/youre-tokenmaxxing-yourself-broke?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://handyai.substack.com/p/youre-tokenmaxxing-yourself-broke?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p></div><p>The good news is that token efficiency has improved about <a href="https://muxup.com/2026q1/per-query-energy-consumption-of-llms">120x</a> since the GPT-3 era, so the floor keeps dropping. The bad news is that Jevon&#8217;s paradox is skyrocketing AI-demand and diminishing these efficiency gains nearly as soon as they take hold.</p><p>Every token you don&#8217;t send is the cleanest token there is. Here&#8217;s how to send fewer.</p><div class="callout-block" data-callout="true"><p><strong>&#128047; Before we get started:</strong> right now a data center is being planned for construction right next to the Nashville Zoo. While water and energy concerns around data centers <a href="https://substack.com/home/post/p-200563358">are overblown</a>, their sound pollution and growing CO2 offsets directly affect our world and animals, making construction near a zoo a nonstarter.</p><p>I urge you to <a href="https://www.change.org/p/nashville-zoo-says-no-to-proposed-data-center">sign the petition</a> to prevent the Nashville Zoo build, and keep an eye out for nonsensical construction plans in your local community (data center or otherwise).</p></div><h3><strong>Practical ways to use less tokens</strong></h3><h4><strong>1. Give your agent  tools that don&#8217;t cost any tokens</strong></h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EnNQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EnNQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png 424w, https://substackcdn.com/image/fetch/$s_!EnNQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png 848w, https://substackcdn.com/image/fetch/$s_!EnNQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png 1272w, https://substackcdn.com/image/fetch/$s_!EnNQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EnNQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png" width="1456" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:71607,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/200836456?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EnNQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png 424w, https://substackcdn.com/image/fetch/$s_!EnNQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png 848w, https://substackcdn.com/image/fetch/$s_!EnNQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png 1272w, https://substackcdn.com/image/fetch/$s_!EnNQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3ab96d03-0bc0-42ea-9c7d-e7b51fa5497e_1456x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>When you ask an agent to do a repetitive task, it reasons through the whole thing every single time. Rename 200 files, reconcile a spreadsheet, pull the same five fields off every PDF in a folder; a naive setup re-derives the approach on every item, burning thinking tokens to rediscover a process it already figured out on item one. You&#8217;ll save some tokens from this kind of work with <a href="https://developers.openai.com/api/docs/guides/prompt-caching">prompt caching</a>, but there&#8217;s an even better way to reduce usage down to nearly zero:</p><blockquote><p>Have your agent write a script the first time, then run the script forever after. </p></blockquote><p>The model reasons once, captures the logic in code, and the code executes for free. Anthropic measured exactly this with their <a href="https://www.anthropic.com/engineering/code-execution-with-mcp">code execution work</a>. A workflow that loaded everything into the model&#8217;s context ate 150,000 tokens. The same workflow, restructured so the agent wrote and ran code instead of reasoning over raw data, used 2,000. That&#8217;s a 98.7% reduction for the identical result.</p><p>This is the whole idea behind <a href="https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview">Skills</a>: folders of scripts and instructions an agent loads on demand instead of being re-taught the same procedure across every conversation. It&#8217;s also why the durable pattern in 2026 is semi-autonomous, not fully agentic. A cron job fires, deterministic scripts do the heavy lifting, and the model only steps in to judge or summarize. One write-up clocked the difference at <a href="https://www.roborhythms.com/make-ai-agent-reliable-without-token-burn/">50x to 500x on cost</a>: a 24/7 assistant left running on a frontier model can cost $4 to $12 a day, while the same job as a scheduled small-model call with hard caps comes in under two cents.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://handyai.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Handy AI is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>The product manager in me loves this because it&#8217;s just good engineering. You don&#8217;t pay a senior engineer to copy-paste the same fix 200 times. You have them write the fix once and automate it. Treat your agent the same way. When you catch yourself prompting the same shape of task twice, stop and say &#8220;write me a script that does this, then run it.&#8221; Done.</p><h4><strong>2. Quit asking for dissertations when you need a sentence</strong></h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QAri!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QAri!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png 424w, https://substackcdn.com/image/fetch/$s_!QAri!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png 848w, https://substackcdn.com/image/fetch/$s_!QAri!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png 1272w, https://substackcdn.com/image/fetch/$s_!QAri!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QAri!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png" width="1456" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:259298,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/200836456?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QAri!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png 424w, https://substackcdn.com/image/fetch/$s_!QAri!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png 848w, https://substackcdn.com/image/fetch/$s_!QAri!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png 1272w, https://substackcdn.com/image/fetch/$s_!QAri!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0d210f67-06b8-48e4-b6a6-8943af6bfc55_1456x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Reasoning models are extraordinary at hard problems and ridiculous at easy ones. Handed something trivial, a large reasoning model will still <a href="https://arxiv.org/abs/2506.23840">pile on thinking tokens</a>, second-guessing itself with &#8220;wait&#8221; and &#8220;hmm&#8221; and backtracking through reflection it never needed. Researchers call it the overthinking trap, and the savings from cutting it are not small. Suppressing those self-doubt tokens <a href="https://arxiv.org/abs/2506.08343">shortens reasoning chains by 27% to 51%</a> with no loss in answer quality. Batch prompting similar questions together <a href="https://arxiv.org/pdf/2511.04108">cut reasoning tokens 76% on average</a> across thirteen benchmarks. Letting the model skip reflection when it&#8217;s already confident saves <a href="https://arxiv.org/pdf/2508.05337">another 18% to 42%</a>.</p><blockquote><p>Match the reasoning to the task. </p></blockquote><p>Most modern models let you set a reasoning effort or thinking budget. Classification, extraction, formatting, summarizing, routine code edits, none of these need extended chain-of-thought. Turn it down. Save the deep reasoning for the genuinely hard stuff, where it earns its keep. Stop bringing Opus to reformat a CSV.</p><h4><strong>3. The cheap wins everyone skips</strong></h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3U1G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3U1G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png 424w, https://substackcdn.com/image/fetch/$s_!3U1G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png 848w, https://substackcdn.com/image/fetch/$s_!3U1G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png 1272w, https://substackcdn.com/image/fetch/$s_!3U1G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3U1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png" width="1456" height="800" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:800,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:101019,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://handyai.substack.com/i/200836456?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3U1G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png 424w, https://substackcdn.com/image/fetch/$s_!3U1G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png 848w, https://substackcdn.com/image/fetch/$s_!3U1G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png 1272w, https://substackcdn.com/image/fetch/$s_!3U1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91b142f4-2f6e-45a3-8aea-c2dfa9c7c2f6_1456x800.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Beyond those two big ones, there&#8217;s a pile of low-effort token savings most people never turn on.</p><ul><li><p><strong>Prompt caching.</strong> Mentioned previously, prompt caching lets the model store words or documents you&#8217;ve sent previously and reuse it. Providers typically charge <a href="https://gleecus.com/blogs/how-prompt-caching-optimizes-llm-latency-and-reduces-token-usage/">50% to 90% less</a> for cached tokens. ProjectDiscovery <a href="https://projectdiscovery.io/blog/how-we-cut-llm-cost-with-prompt-caching">cut their LLM bill 59%</a> doing nothing but this.</p></li><li><p><strong>Mind your output.</strong> Output tokens cost <a href="https://www.adaline.ai/blog/llm-cost-optimization-token-efficiency-caching-prompt-design">3 to 8 times more</a> than input tokens at every major provider. A rambling answer with heavy reasoning is the expensive kind of token. </p></li><li><p><strong>Compress the conversation.</strong> Long chat threads replay the entire history on every turn. Summarizing older exchanges instead of resending them verbatim trims <a href="https://redis.io/blog/llm-token-optimization-speed-up-apps/">20% to 40%</a> off a chatbot&#8217;s token use. Most agent frameworks will do this for you if you ask.</p></li><li><p><strong>Route by complexity.</strong> A <a href="https://www.mindstudio.ai/blog/what-is-ai-model-router-optimize-cost-llm-providers">model router</a> inspects each request and sends the easy ones to a cheap model and the hard ones to the expensive one. You get frontier quality only where it matters and pay small-model prices everywhere else. Most harnesses (Cursor, Claude Code) have a version of routing built in, but it can&#8217;t hurt to double up.</p></li></ul><h3><strong>Insourcing: keep some tasks for your own brain</strong></h3><p>I have a rule I call <strong>insourcing</strong>: the practice of keeping a deliberate set of tasks you refuse to hand to a model, specifically so the part of your brain that does them doesn&#8217;t go soft. The first thing on my list is personal communication. I don&#8217;t let an agent write the text to my grandmother, the note to a friend going through it, or the message to someone I actually care about. Not because the model would do it badly, but because the doing is the point. </p><p>Beyond this obvious line, I ensure that a not-insignificant portion of my daily work is kept in-house. There is an ever-present itch for me to reach for Claude to build out slides or format a spreadsheet; Claude Cowork has progressed to a point where it can handle these tasks with relative ease. Resist! Make sure you&#8217;re putting in the work on these when you can. The initial act of &#8220;creation&#8221; in work is important for flexing your brain muscle.</p><p>An MIT Media Lab team <a href="https://handyai.substack.com/p/your-brain-on-chatgpt?utm_source=publication-search">ran a study last year they called &#8220;Your Brain on ChatGPT&#8221;,</a> wiring up 54 people writing essays with AI, with a search engine, or with nothing but their own head. The AI group showed the weakest brain connectivity and the thinnest memory of what they&#8217;d just written. Even after they put the AI away, their brain activity stayed sluggish. The researchers called it &#8220;cognitive debt.&#8221; You outsource the thinking and your brain doesn&#8217;t eagerly grab the wheel back when you ask it to.</p><p>This is the not-so-quiet cost of tokenmaxxing your whole life. There is physical muscle atrophy that happens with every email, every decision, and every paragraph generated by your agents.</p><p>Insourcing is the cheapest and healthiest token-reduction strategy there is. Keep a few things that stay slow and human, and you&#8217;ll spend fewer tokens and keep a sharper head doing it.</p><h3><strong>The on-device future</strong></h3><p>Every technique above is aimed at reducing data center usage by the beefy cloud models. But what if these models ran locally, on your phone or laptop, and skipped the trip to the data center altogether?</p><p>Apple just bet its whole software stack on that question. The centerpiece of <a href="https://www.engadget.com/2189698/everything-announced-at-apples-wwdc-2026-keynote/">WWDC this week</a> was <a href="https://ofox.ai/blog/apple-foundation-models-3-wwdc-2026-developer-read/">AFM 3 Core Advanced</a>, a 20-billion-parameter sparse on-device model that activates only 1 to 4 billion parameters per prompt. That pruning trick is what lets it live on an iPhone 17 Pro instead of in a data center. The <a href="https://www.macrumors.com/2026/06/09/apple-outlines-major-ai-and-developer-tool-updates/">Foundation Models framework</a> hands the model to any app on the device: no API key, no network call, no token meter. It now accepts images, so receipts, screenshots, and photos get parsed on the Neural Engine without a byte leaving your hand. Apple is even open-sourcing the framework this summer. Apple is seemingly try to make up for their historically horrible Siri launches with some premier local LLMs as a default OS capability, though the hardware floor is real: 12GB of RAM and recent silicon means older iPhones stay in the cloud era a while longer.</p><p>Google is racing down the same curve in the open-weights lane. Gemma 4 12B handles text, image, audio, and video, beats last generation&#8217;s Gemma 3 27B, and fits on a 16GB laptop. One Ollama command and it runs fully offline with no strings on commercial use. Its edge sibling, Gemma 4 E2B, runs on your iPhone today through an app like Locally AI.</p><p>No wifi or signal required. No data leaving your device. No API bill. No query metered against a reservoir somewhere in Virginia.</p><p>This is the endgame for both problems at once. On-device models kill the per-token cost, the environmental draw, and the privacy tax. Notice that even Apple&#8217;s architecture concedes the split: the heavyweight, <a href="https://techcrunch.com/2026/06/09/wwdc-2026-everything-announced-on-siri-ai-os-27-apple-intelligence-and-more/">Gemini-assisted model behind Siri AI</a> routes through Private Cloud Compute, while summarization, extraction, and classification stay on the chip. The frontier clouds keep the genuinely hard, novel reasoning that earns the watts. The thousand small calls that make up most of what people actually do with AI are moving into your pocket. </p><p>But the cost curve points home.</p>]]></content:encoded></item></channel></rss>