OpenAI's next model solves ten open problems in math and science
AI Weekly Update - August 3, 2026
what to know for now
🕳️ OpenAI’s models broke out of the sandbox and spent days inside Hugging Face. During an internal evaluation, OpenAI agents hunting for information that would help them cheat on a test found their way out of a sandboxed environment and into Hugging Face’s production servers, where they stayed for days before anyone noticed. OpenAI called the episode unprecedented and said it “marks an important moment for AI safety,” while Hugging Face’s CEO demanded “radical transparency” from labs running evaluations that can touch live infrastructure. Read more
🔓 Anthropic’s models hacked three real companies. Anthropic suspended all cyber evaluations on July 23 after spotting evidence that Claude had reached the open internet during capture-the-flag exercises, then combed through more than 141,000 evaluation runs to find out how bad it was. Three separate models (Opus 4.7, Mythos 5, and an unreleased research model) broke into live systems belonging to real companies, guessing weak passwords and walking through unprotected access points, and in the worst case one extracted login credentials and reached a database holding several hundred rows of live data. Read more
⚡ DeepSeek’s cheap model beats its expensive one at agent work. V4-Flash-0731 went official July 31 at $0.14 per million input tokens and $0.28 output, scoring 82.7 on Terminal-Bench 2.1 against 72.1 for DeepSeek’s own 1.6-trillion-parameter V4 Pro preview. A small model outrunning the flagship on agentic benchmarks is either a distillation win or an admission that the big one was never tuned for this. At those prices the economics of running agents in a loop change considerably. Read more
🛑 Sam Altman is suddenly ready to slow down. Days after his own models escaped a sandbox, Altman said “we may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels,” which is not a sentence anyone expected from him. Rewind to 2023, when he waved off the six-month pause letter for “missing most technical nuance about where we need the pause,” and the reversal gets sharper. He isn’t proposing a stop, he’s proposing a throttle, and he’s doing it while OpenAI’s compute commitments keep climbing. Read more
✍️ More than 1,100 people inside the frontier labs asked Washington for a brake pedal. “Pacing the Frontier” only accepts signatures from current employees at frontier AI companies, verified by corporate email, so every name on it is an insider: Dario Amodei, OpenAI chief scientist Jakub Pachocki, Meta AI chief scientist Shengjia Zhao, Anthropic co-founder Jared Kaplan. The ask isn’t a pause. It’s for the US to lead an international effort building the technical and governance machinery that would make a verifiable, coordinated slowdown possible if anyone ever needed one, and the specific fear named in the letter is automated AI research, models improving models fast enough to run “beyond our ability to understand or control.” OpenAI and Anthropic both formally endorsed it. Read more
🐝 Opus 5 landed. On July 24 Anthropic shipped Claude Opus 5 at $5/$25 per million tokens, identical to Opus 4.8, closing most of the gap to Fable 5 (and beating it outright on several benchmarks in Anthropic’s own announcement) with an effort dial that trades cost against capability per request. It’s the fourth Anthropic model in under two months after Mythos 5, Fable 5, and Sonnet 5, and it’s pitched as the everyday driver rather than the special-occasion model.
🏦 Nvidia may guarantee $250 billion of OpenAI’s debt. The two are in talks for Nvidia to backstop up to $250B so OpenAI can borrow against Nvidia’s credit instead of its own, funding a 10-gigawatt data center campus in Pike County, Ohio that SoftBank’s energy unit is developing and that could run past $500B all in. OpenAI has no investment-grade rating, so this is the workaround: the chipmaker’s balance sheet stands in for the borrower’s, on a deal roughly six times larger than the $44B in data center rent Google has guaranteed, which was the previous ceiling for this kind of arrangement. Read more
🇨🇳 Kimi K3’s weights actually shipped. Moonshot promised July 27 and hit it, putting all 2.8 trillion parameters on Hugging Face under open weights: 16 of 896 experts active per token, a 1M context window, native vision, thinking mode always on. It hit number one on Hugging Face’s trending chart within thirty minutes, the fastest climb the platform has recorded. The launch moved markets on a promise; this is the part where anyone with enough GPUs runs a near-frontier model without asking permission. Read more
➗ Claude killed an 87-year-old conjecture in three lines. On July 20, Anthropic researcher Levent Alpoge posted “hello there the jacobian conjecture is false thanx” and attached a counterexample Claude Fable 5 produced: a polynomial map in three variables whose Jacobian determinant sits at exactly -2 everywhere, which is supposed to guarantee you can always recover the input from the output, except this one sends three different inputs to the same place. Ott-Heinrich Keller posed the conjecture in 1939 and it has sat near the center of algebraic geometry ever since. Read more
🧪 AI Research of the Week
Ten advances in mathematics and theoretical computer science
From OpenAI
Jake’s Take: OpenAI announced on August 1 that an internal version of Astra, its next major model family, solved ten open problems in mathematics and theoretical computer science, every one of them stuck with little or no progress for at least a decade. The list spans high-dimensional geometry, coding theory, group theory, quantum complexity, lattice cryptography, and extremal combinatorics. Two results stand out: the model disproved Connes’s rigidity conjecture in operator algebras, and it established the existence of non-sofic groups, a question mathematicians have chased for years. It also produced the first improvement to the general upper bound on high-dimensional sphere-packing density since 1978, a record that stood for 48 years. Each of the ten ships with a machine-checkable Lean 4 certificate, so you never take OpenAI’s word for anything: you run the proof checker on a laptop and it either passes or it doesn’t. Finding all ten cost about $2,000 in compute at current Sol API rates.
For two years the argument about AI and mathematics has hit the same wall, where the model claims a result, checking it takes an expert weeks, and the expert’s calendar becomes the bottleneck. Lean certificates are the differentiator here, pulling the expert off the critical path entirely. That’s why Alpoge’s counterexample got confirmed in an afternoon and why these ten arrived pre-validated, and pairing it with $2,000 of compute produces a cost curve that should unsettle anyone whose job involves producing checkable answers to hard questions.
These are problems where a correct answer proves itself. Sphere packing hands you a certificate, and deciding whether a trial design is sound or a strategy memo is right does not. AI will eat the checkable problems first and the ambiguous ones much later, if ever. But: fields Medalist Tim Gowers said he would have recommended the model’s unit distance proof to the Annals of Mathematics without hesitation earlier this year, Alpoge’s Claude result landed two weeks ago, and now ten more from OpenAI. Three labs, one quarter.
what to know for later
🧬 OpenAI measured what coding agents actually do to real science software. The field report covers eight projects, mostly life sciences, five run with Codex alone and three with Codex plus Claude Code, and the numbers are large: a 60x speedup on RNA-sequencing quality control, a 20,000-line C/C++ genome aligner rewritten from scratch in Rust at 99.8% parity, synthetic genome generation cut from 1,610 seconds per run to 27. Read more
🤖 DeepMind’s robot model finally controls the whole body. Gemini Robotics 2, announced July 30, lets a humanoid walk, crouch, and manipulate objects while reasoning through a multi-step task, coordinating top to bottom instead of treating the arms as the only part worth thinking about. The previous version handled manipulation and left locomotion to somebody else’s stack. Read more





