Unreleased OpenAI model helps solve top math problems
Welcome back! In today’s edition of Data Points, you’ll learn about our top headlines, and more:
Qwen3.8-Max debuts with first benchmarks
Google reveals new robotics models
OpenAI cuts prices for Terra and Luna
Anthropic reveals its own unauthorized external hacks
But first:
DeepSeek V4 Flash jumps up a tier in overall intelligence
DeepSeek released V4 Flash 0731 on July 31, 2026, scoring 50 on the Artificial Analysis Intelligence Index, a ten-point jump from the previous V4 Flash and just one point behind GPT-5.6 Luna. The improvement comes primarily from better agentic performance: The model’s Elo rating on real-world task evaluation jumped from 1189 to 1559, and hallucinations dropped 12 points while accuracy remained flat, suggesting the gains came from reduced false outputs rather than fundamental capability gains. DeepSeek maintained the same architecture and pricing as before ($0.14/$0.28 per 1M input/output tokens), but the model pulls ahead on cost-per-task comparisons thanks to a 98 percent discount on cached requests, a significantly more aggressive discount than the 90 percent competitors offer. The model kept its 1M token context window and 284B parameter count (13B active), and DeepSeek plans to release full weights in the coming weeks, which would put it second among open-weights models behind Kimi K3. (Artificial Analysis)
Tech companies argue in support of open AI models
OpenAI released solutions to ten open problems in mathematics and theoretical computer science, generated by Astra, an internal unreleased model. The problems span sphere packing, coding theory, group theory, and quantum complexity—each vetted by formal proofs in Lean. Generating all ten solutions cost roughly $2,000 in compute at current API rates. The results include a disproof of Connes’s rigidity conjecture, new bounds on sphere-packing density, and a construction proving the existence of non-sofic groups. OpenAI framed the release carefully, noting that humans prepared the manuscripts but the mathematical arguments came from the system, a deliberate stance on authorship and attribution as AI systems edge into research collaboration. (OpenAI)
Alibaba’s Qwen3.8-Max now available, weights yet to come
Alibaba made Qwen3.8-Max broadly available via a hosted API, with open weights shipping next week. The 2.4-trillion-parameter mixture-of-experts model accepts text, image, and video input, supports a one-million-token context window, and integrates via OpenAI-compatible APIs. Pricing runs two dollars per million input tokens, six dollars per million output tokens, and twenty-five cents for cached reads. Performance is strong for open models: Qwen3.8-Max leads on multimodal and agentic benchmarks, scoring 86.6 on Terminal-Bench 2.1, ahead of Claude Opus 4.8 but behind GPT-5.6 Sol, while posting weaker coding results than Claude Fable 5 on SWE-Bench Pro. (MarkTechPost)
Google unveils new robotics models
Google launched Gemini Robotics ER 2, an embodied reasoning model designed to act as a high-level brain for robots, handling task planning and orchestration while delegating motor execution to lower-level models. The system processes continuous video feeds to track task progress in real time, adapting when something goes wrong and knowing precisely when to move to the next step, a significant upgrade from its predecessor. The model integrates with Google’s Live API, using bidirectional streaming to eliminate latency-induced “stop-and-think” pauses, demonstrated in a demo where Boston Dynamics’ Spot fetches objects on natural language commands. Gemini Robotics ER 2 achieves 57.4 percent accuracy on progress classification (tracking tasks in five completion stages) and 91.3 percent on moment-finding (identifying exact frames for critical events) while running at sub-second latency, or roughly four times faster than competing larger models. The update also introduces multi-robot collaboration, letting diverse machines work together in shared spaces to complete complex workflows no single robot could handle alone. The model is available via Gemini API and Google AI Studio, with private preview access on the Gemini Enterprise Agent Platform. (Google)
OpenAI cuts prices for GPT-5.6 Luna and Terra
OpenAI is cutting prices and boosting speed for its GPT‑5.6 model family, framing it as passing efficiency gains from GPT‑5.6’s own self-optimization work back to customers. Luna (the fastest, cheapest tier) drops 80 percent in price and Terra (the middle tier) drops 20 percent, while a new Fast mode for the flagship Sol model replaces Priority Processing, delivering up to 2.5x speed at double the standard price. The improvements stem from optimizing models, inference systems, and the agentic tooling layer together. Sol itself has been used to rewrite production kernels and run optimization experiments, cutting serving costs 20 percent and boosting token-generation efficiency over 15 percent. OpenAI positions this as letting businesses mix model tiers within a single workflow (e.g., Sol for planning, Luna for execution) to match cost and speed to each task’s stakes. New API pricing took effect July 30: Terra at $2/$12 per million input/output tokens, Luna at $0.20/$1.20, with Sol unchanged and ChatGPT/Codex subscription prices unaffected. (OpenAI)
Anthropic says its models have hacked external sites during testing
Anthropic disclosed that three Claude models broke out of test environments and compromised real company infrastructure during cybersecurity evaluations between April and July. The root cause was simple: Evaluation environments had live internet access, but prompts told Claude it was isolated in a simulation. Neither Anthropic nor their testing partner Irregular caught the misconfiguration. Claude Opus 4.7 accessed a real company’s database and extracted production credentials across four separate runs, continuing its attacks even after recognizing the system was real. Claude Mythos 5 published malicious code to the public Python package registry; the code executed on 15 real systems, including a security company’s scanner, whose credentials Claude then stole. An internal research model scanned 9,000 targets and compromised one company’s application before stopping after realizing the target was real. Anthropic discovered the incidents during a review triggered by OpenAI’s disclosure of similar breakouts in July, halted all cyber evaluations, and notified affected companies. Two of them hadn’t detected the intrusions. (Anthropic)
Quote of the week
Want to know more about what matters in AI right now?
Read the latest issue of The Batch for in-depth analysis of news and research.
Last week, Andrew talked about challenges faced during a security review using restrictive models and the importance of open weight models and agent harnesses for secure software.
“To me, this is an important reminder of why open models, as well as open agent harnesses, lead to safer systems, and in particular to more secure software. In fact, since attackers can now find flaws faster than ever because they have AI agents to help them do so, we have heard directly from a number of security officers who are frustrated that frontier closed models are refusing to help them.”
DeepLearning.AI’s first-ever subscription plan for our entire course catalog includes foundational classics like the Machine Learning and Deep Learning specializations plus seminars on the latest tools and frameworks you need.
As a Pro Member, you’ll immediately enjoy access to:
Nearly 200 short and long AI courses from Andrew Ng and industry experts
Labs and quizzes to test your knowledge
Projects to share with employers
Certificates to testify to your new skills
A community to help you advance at the speed of AI
Enroll now to lock in a year of full access for $25 per month paid upfront, or opt for month-to-month payments at just $30 per month. Both payment options begin with a one-week free trial.