|
Dear friends,
Anthropic recently released an encouraging analysis of the cyber capabilities of open weight model GLM-5.3. The study shows that GLM-5.3 approaches the cyber capabilities of Claude's closed weight Mythos. For example, the diagram below shows that, using a comparable number of tokens, GLM-5.3 succeeded in 12% of attempts to exploit vulnerabilities on a subset of ExploitBench tasks, compared to Mythos’ 14%. (Interestingly, the gap reported by Anthropic is smaller in GLM-5.3’s creator's report on the full benchmark: 54.4% success compared to Mythos' 78.0%.) This paves the way for cyberdefenders, including ones that do not have access to Mythos, to use a highly capable model to defend themselves. Given attackers’ access to similar capabilities, the urgency of ramping up defenses grows.
Further, the cost of finding vulnerabilities and using it to create an attack (that is, to create a payload that targets the identified vulnerability) is falling. Anthropic found that spending $20.40 on GLM-5.3-Flash tokens was sufficient to find a recently disclosed flaw in Google Chrome. The results reported used an unmodified version of GLM-5.3. It is also easy to obtain frontier open weight models that have had their guardrails weakened. Such models would be even more effective for cyberdefense (or cyberoffense).
A lot of AI risks, including cyber vulnerabilities and how to contain agents, are engineering problems to be solved. To be clear, these are hard engineering problems! But I find that a lot of the fear-mongering scenarios implicitly assume that we will make little or no progress on them.
For an analogy, consider aviation. According to the National Postal Museum, in 1919, one person died for about every 115,000 miles flown. Today, airlines fly about 6 trillion passenger miles. So, airplanes kill about 6 trillion /115,000 = 52 million persons a year, right? Because we have engineered airplanes to be much safer, we now see one death approximately every 30 billion passenger miles.
Similarly, even though mass media describes the OpenAI-Hugging Face hack as dangerous agents “going rogue,” this incident is leading teams everywhere to engineer better sandboxing and monitoring – important changes that are making AI safer. (On OpenWorker, our agent harness that supports security workflows, Rohit Prsad, Devika Verma and I are building on Nvidia's new safety framework to improve sandboxing.)
In aviation and in software, we improve safety by repeatedly finding problems, fixing them, and further identifying and addressing root causes that reduces the odds of future problems. Frontier models are speeding this process up.
The advanced cyber capabilities of open weight models do give attackers a window to do damage. Long term, defenders have the advantage, because they have more information with which to discover and mitigate vulnerabilities. The key is to carry out this important and difficult engineering work as quickly as we can, so we can all come out the other side with safer, more robust software.
Keep building! Andrew
A MESSAGE FROM DEEPLEARNING.AIThe Data Engineering Professional Certificate is now available on DeepLearning.AI. Across four courses, design and build systems that generate, ingest, store, transform, and serve data, including batch and streaming pipelines on AWS and open-source tools. Taught by Joe Reis, co-author of Fundamentals of Data Engineering. Enroll for free
News
An Unexpected Open Weights Leader
Xiaomi, best known for its smartphones and electric vehicles, released the highest-scoring open weights model on Artificial Analysis’ Intelligence Index. It performs similarly to GPT-6 Sol, but costs less per task.
What’s new: Along with MiMo-V2.6-Pro-RL and its smaller sibling MiMo-V2.6-Flash, the company also released MiMo-V2.6-Distill-Qwen-9B (Alibaba’s Qwen3.5-9B fine-tuned on MiMo outputs). In addition to the models, Xiaomi also published more than 7,000 reinforcement learning (RL) task environments (software workspaces in which a model attempts tasks and receives feedback) and the code to train on them.
How it works: MiMo-V2.6-Pro processes images, video, and audio through separate encoders that feed its language model. According to Xiaomi’s technical report, the training recipe combines methods from earlier work, much of it Xiaomi’s own. The main innovation rewards coding attempts for quality rather than only for passing tasks.
Performance: MiMo-V2.6-Pro topped open weights models on two independent composite evaluations, making a large gain over its predecessor. It matched some proprietary models at a small fraction of their cost per task, but it trailed the leaders and generated output slowly.
Behind the news: Recently, Anthropic alleged that, over 20 days in March and April 2026, Xiaomi routed MiMo users’ conversations and coding sessions to Claude through the coding tools OpenClaw and OpenCode, producing more than 400,000 exchanges to use as training data. To date, Xiaomi has not publicly responded to Anthropic’s accusation.
Why it matters: Code that passes tests isn’t necessarily good code. Xiaomi found that a version of its model that was trained without the code quality grader picked up habits that make software harder to maintain: The model added code the task didn’t call for, let errors pass silently, and were lax in checking on incoming data until tests passed. Trained with the code quality grader, the model made smaller, more focused changes. Rewarding code that a human reviewer would accept is as important as rewarding code that runs.
We’re thinking: Open weights are great, but open recipes are even better. By releasing its reinforcement learning tasks, the software that runs them, and its training code, Xiaomi lets others repeat the steps that produced much of its models’ improvements. It also disclosed what those steps cost: $2.6 million for MiMo-V2.6-Pro and $0.9 million for MiMo-V2.6-Flash. Teams planning to train agents get a set of tasks they don’t have to build and a rough price for reinforcement learning at this scale. We hope other labs can reproduce and extend on what Xiaomi learned training these models.
Voice, Video, and Reasoning in One
Google released two new speech-to-speech voice models, joining a suddenly crowded field of models built to power voice agents.
What’s new: Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech-to-speech models that can serve as real-time voice agents. The Gemini 3.8 Live model is built for scale and cost-efficiency while 3.8 Live Extended Thinking is intended for high-complexity tasks. Both models are designed to reduce response latency during live dialogue.
How it works: Google released few details about the model’s architecture and training data, but did say that the models are based on Gemini 3 Pro. Gemini 3.8 Live can take in image and video input, supports 97 languages, and can execute tools and API calls in the background while continuing a conversation. Live Extended Thinking can also reason and think simultaneously. The models’ knowledge cutoff is January 2025.
Behind the news: Traditionally, voice agents have relied on a pipeline that converts speech to text, sends it to a reasoning model, and converts the response back to speech. This approach has historically been seen as more accurate and easier for developers to control because LLMs generally reason in text. However, each step can add latency, which detracts from the user experience. OpenAI recently released a speech-to-speech model called GPT-Live-1, a voice model developers can use to build voice-enabled apps. This model can listen and speak at the same time, allowing it to handle interruptions and respond more naturally, while a separate reasoning model runs in the background.
Why it matters: It’s correct to call Gemini 3.8 Live a voice model since its primary language input is voice rather than text, but its ability to understand and reason over image and video input makes it more versatile. Putting voice and video input together is particularly powerful for helping users use AI in real-time to deal with anything on screen: web interfaces, games, multi-application workflows, etc. This differentiates it from GPT-Live-1 and other models that work strictly with voice and audio.
We’re thinking: Voice-driven agents are particularly exciting for developers because they offer a new real-time interface for building applications. Speech is a natural input format that most people can use regardless of their experience using a computer. It’s particularly useful for mobile and automotive interfaces where text input is less convenient. It’s up to developers to create applications that democratize users’ access to AI.
Learn More About AI With Data Points!
AI is moving faster than ever. Data Points helps you make sense of it just as fast. Data Points arrives in your inbox twice a week with six brief news stories. This week, we covered OpenAI halting model training and the biggest announcements from OpenAI’s Dev Day. Subscribe today!
DeepSeek’s Flash Leapfrogs Pro Again
When agents make a tool call, they send the same context back through the model every time. DeepSeek says storing and moving all that context has become a bigger obstacle to keeping serving costs down than the computation a model performs to generate output. Its new architecture sidesteps that bottleneck.
What’s new: DeepSeek released DeepSeek-V4.1-Flash, its first model with a redesigned architecture that uses far less memory per token of context. It also cut its API prices. This is the first model in DeepSeek’s V4 series to accept images as input, outside of an experimental preview. It’s also very fast, second only to Gemini 3.8 Flash in tokens/second.
How it works: A paper details how DeepSeek redesigned its V4 architecture. The authors reduced the key-value cache (the keys and values a transformer stores for every token it has read so that it avoids recomputing them) and the computation spent reading input.
Performance: Independent evaluators found that DeepSeek-V4.1-Flash outperforms its predecessor and the larger DeepSeek-V4-Pro-0813. It costs less than half as much per task and generates output more than twice as fast. It briefly led open weights models on Vals.ai’s index (prior to the release of MiMo v2.6 Pro) and tops all models on one independent test of business-software agents.
Behind the news: Shrinking the key-value cache has been a DeepSeek theme since its second-generation model. Rivals have borrowed DeepSeek’s techniques and developed their own to likewise compete on cache size.
Why it matters: Agents usually read more than they write. Because each tool call sends a growing transcript back through the model, storing and re-reading context can cost more than generating output. DeepSeek’s price cut is steepest where long-running agents use the most tokens: The cost for previously cached input fell 57 percent, while the cost of output fell 9 percent.
We’re thinking: A year ago, model makers strove to lengthen context windows. DeepSeek-V4.1-Flash keeps its predecessor’s 1 million-token context and instead makes it cheaper to keep: 890 bytes of cache per token, 437 times smaller than DeepSeek-V1’s and 4 times smaller than DeepSeek-V4-Flash’s. If other model makers follow DeepSeek’s lead, launches may soon tout cache size per token alongside context windows and cost per token.
Agents That Can Check Their Own Work Piece by Piece
When asked questions requiring extensive reasoning and tool use to answer, agents may search without making progress, repeat unsuccessful strategies, or settle for partially verified answers. Building answers iteratively can help to address these shortcomings.
What’s new: Researchers at the Beijing Academy of Artificial Intelligence developed AREX, a research agent that hones its work through repeated cycles of gathering evidence, reflecting on provisional answers, and launching targeted follow-up searches. The combination of an iterative harness and a fine-tuned model outperformed competing approaches.
Key insight: Some earlier agents check a candidate answer against a question’s requirements and accept or reject it. However, the check can do more work. When a tentative answer fails to satisfy every requirement, checking it requirement by requirement reveals which ones remain unsupported and where the available evidence conflicts. Instead of discarding the answer, the agent can keep the parts it has confirmed and use unmet requirements as the basis for a further question to research next. This turns a pass/fail judgment into a loop that drives improvement.
How it works: The authors built an agentic harness to answer difficult research questions. They used it to build a dataset and fine-tuned two models to work with it more effectively: Qwen3.5-4B, which is fairly small, and Qwen3.5-122B-A10B, a much larger mixture of experts.
Results: Across six agentic benchmarks, the AREX systems based on the fine-tuned Qwen3.5 models outperformed systems that used the same harness with similarly sized models as well as the Qwen3.5 models without fine-tuning.
Why it matters: Research agents can perform better with both fine-tuning and repeated efforts to improve their output. The training strategy of identifying and amplifying the most informative decision points offers a practical lesson for improving agent behavior: Not all steps in a long trajectory are equally important. Focusing training on critical decision points yielded better results than treating every action the same way.
We're thinking: The authors used hand-written rules, not a learned model, to detect and target training examples. These let the model learn how to compress long traces into more valuable decision points. The combination shows that cheap but clear heuristics can outperform elaborate ones when the thing being detected is well-defined.
Work With Andrew Ng
Join the teams that are bringing AI to the world! Check out job openings at DeepLearning.AI, AI Fund, Landing AI, and LearnVector.
Subscribe and view previous issues here.
Thoughts, suggestions, feedback? Please send to thebatch@deeplearning.ai. Avoid our newsletter ending up in your spam folder by adding our email address to your contacts list.
|