|
Dear friends,
As AI increasingly automates coding, it frees up developers to spend time on high-level software development tasks traditionally reserved for senior engineers, like deciding on technical architecture and participating in scoping product requirements. Demand for this type of work is growing, since it is an economic complement to coding, which is becoming cheaper. This is why I’m confident there will be rising demand for broad AI engineering skills. I see a similar pattern starting to emerge in other AI-influenced fields as well, and am cautiously optimistic — contrary to predictions of a “jobpocalypse” — that AI will generally increase demand for people with the right skills.
The broader pattern is this: In many job roles, as you rise in seniority, you become better at managing integration complexity — weaving together many disparate work streams like frontend, backend and data engineering to form a greater whole. As AI automates certain parts of one’s work — usually the more verifiable parts — it creates more room for individuals to play broader roles.
This pattern is not applicable to all job roles. In some, rising in seniority means increasing specialization. This corresponds loosely to people progressing in the IC (individual contributor) career track rather than the managerial/tech lead track. For example, a machine learning engineer acquiring extremely deep understanding of a technical niche, a financial expert growing from a generalist finance analyst to a specialist in an important sector (such as auditing cross-border deals), or a medical doctor developing deep expertise in just one medical condition. In many, but not all, such areas, I expect the demand for human skill to grow, but AI’s impact on any specialty will depend on how rapidly AI’s jagged frontier advances along that dimension. I’ll share more on this in a future letter. Keep building! Andrew
A MESSAGE FROM DEEPLEARNING.AIBuild LLM applications that respond in real time on Cerebras' Wafer-Scale Engine, from live personalization to a multi-tool workflow that finishes in one fast response. Enroll for free
News
One Model Talks, Another One Thinks
ChatGPT’s voice mode now listens and speaks at the same time, passing harder questions posed to the conversational model to a reasoning model in the background.
How it works: OpenAI published a GPT-Live system card that describes a system of several models working as an ensemble. The voice models differ from their predecessors in two ways: they process audio continuously rather than turn-by-turn, and they hand deeper work to a separate model.
Performance: In OpenAI’s evaluations, the biggest gains show up on the tasks routed to GPT-5.5, with the delegation system working as designed. Every comparison OpenAI published pits GPT-Live against its own predecessor, AVM, rather than rival voice models, and its strongest claims rest on internal benchmarks.
Behind the news: Both halves of GPT-Live’s design (full-duplex processing and reasoning-model orchestration) have precedents. Alibaba’s Qwen2.5-Omni Thinker-Talker architecture trained a text-generating “thinker” and a speech “talker” as one system. Thinking Machines Lab paired a foreground interaction model with a background reasoner in TML-Interaction-Small. Kyutai’s Moshi, billed by its authors as the “first real-time full-duplex spoken large language model,” arrived in 2024; Nvidia released PersonaPlex, an open-weights model built on Moshi, in January; and Google’s Gemini Live currently offers continuous conversation along with camera and screen sharing. OpenAI can boast that it ships the combination to a mass audience; the company says more than 150 million people use ChatGPT’s voice and dictation features each week.
Google Found Responsible for AI-Generated Search Results
A German court ruled that Google can be held liable for defamatory statements generated by the AI Overview that appears at the top of its search results.
What’s new: The Regional Court of Munich found that Google’s AI Overview generated false, reputation-damaging claims. The generated text said that two German publishers engaged in fraudulent business practices and lured unknowing customers into buying subscriptions. Because Google generated the AI Overview, the court ruled that the company was responsible for defamatory falsehoods it contained. The court issued a temporary injunction that requires Google to stop disseminating the statements.
How it works: When users searched for the German publishing house Verlagshaus24 or its subsidiary GeraMond followed by the word “scam” — a term suggested by Google's autocomplete function — AI Overview responded with a summary that stated, “Yes, [company] is known for dubious business practices.” It also listed characteristics of an alleged scam including subscription traps, poor customer service, and content that remained locked even to paying customers. The court determined that the publishers in question had not, in fact, been accused of misconduct. Instead, the AI Overview had confused them with other companies that were suspected of fraudulent practices. The court found Google’s legal defenses lacking and held it responsible for defamation, and it enjoined Google to remove the statements immediately or pay a fine.
Behind the news: German and EU law generally have treated search engines as information intermediaries that display content rather than create it. Thus, search engines have borne limited liability for unlawful information in their results — until now.
Why it matters: Munich’s ruling against Google is a landmark case that shifts search engines’ liability for false statements. In the near term, it could alter the way Google presents its AI Overview. In the longer term, if the ruling is upheld on appeal, it could encourage similar lawsuits in Europe and elsewhere. One analysis found inaccuracies in roughly 10 percent of Google AI Overview results. If this figure holds, Google and other AI providers could face significant risk of litigation.
We’re thinking: We think that agentic question-answering systems can be improved significantly — for example, by grounding them in more reliable information sources. We remain optimistic that such systems can continue to operate at scale.
Learn More About AI With Data Points!
AI is moving faster than ever. Data Points helps you make sense of it just as fast. Data Points arrives in your inbox twice a week with six brief news stories. This week, we covered Apple’s lawsuit against OpenAI and a new technique that fits a 27-billion-parameter AI model on an iPhone. Subscribe today!
Put the Lab in the Loop
An AI agent proposed new medical uses for established drugs nearly autonomously — uses that were supported by experiments on isolated human cells — with human input only to name diseases to be treated and run the AI-proposed lab experiments.
What’s new: Ali Essam Ghareeb, Benjamin Chang, and colleagues from the independent AI research lab FutureHouse, University of Oxford, and Fordham University released Robin, an open-source agent that proposes existing drugs to treat a given disease. Robin identified two drugs that were shown to address a biological mechanism behind dry age-related macular degeneration (dAMD), a leading cause of impaired vision. While Robin is freely available for noncommercial and commercial uses under the Apache 2.0 license, it relies on three earlier agents, two of which are proprietary. The literature-search agents called Crow and Falcon are free for research use only (bundled under the name Literature). The data-analysis agent Finch is available under an Apache 2.0 license.
How it works: For a given disease, Robin iteratively (i) identifies mechanisms behind the disease, (ii) designs experiments to affect the mechanisms, and (iii) finds existing drugs that address the mechanisms. Then (iv) humans run experiments in a lab, and (v) Robin analyzes the results. Robin uses OpenAI’s GPT o4-mini for most language-processing functions.
Results: The authors ran their pipeline for dAMD. Robin hypothesized that increasing a process known as RPE phagocytosis, in which a particular type of cell in the eye removes pathogens and debris, could treat the disease. Robin proposed candidate drugs to boost RPE phagocytosis, of which two proved effective.
Yes, but: While the authors did test the drugs on eye cells, they did not test them on patients with the disease, so it is not yet known whether they work in living patients.
Behind the news: This work followed several studies dedicated to accelerating scientific research through agentic systems. One agentic system has generated machine learning research proposals as well as or better than humans. Another can generate a proposal, write and run code to test the proposal, and write the paper that describes the experiment and results. Google’s AI Co-Scientist, generated research proposals for biomedicine that humans later validated in the lab to treat acute myeloid leukemia.
Why it matters: The Robin pipeline not only generates hypotheses and proposes experiments, it also analyzes new experimental results and updates its hypotheses accordingly. Its combination of automation and iteration suggests that AI agents could streamline medical progress. Advances in robotics may enable AI models to carry out the necessary lab experiments as well, as previously demonstrated by RoboChem, a system in which an AI model controls a set of automated lab instruments. Similar approaches may be applicable to research areas beyond medicine.
We're thinking: It takes more than a decade and costs more than $1 billion to develop a drug in the United States, and the government approves only around 50 drugs a year. Finding new uses for drugs that are already approved, and new treatments for diseases that share underlying biological mechanisms, is an especially efficient use of time and funds.
Measuring Models’ Manipulation
Providers of large language models stand to benefit by building models that spur user engagement, but users may bear a cost in undue influence on their world views. How readily LLMs can change users’ beliefs becomes an important question as users increasingly turn to them for information and advice.
What’s new: Jocelyn Shen and colleagues at MIT and Carnegie Mellon University measured the effect of OpenAI GPT-4o on users’ beliefs. They also tested various LLMs’ ability, after a user conversed with GPT-4o, to estimate that model’s influence on the user’s beliefs, and they introduced the Puppet benchmark to test models’ ability to estimate such influence. The authors propose Puppet as an alternative to earlier models that were designed to detect manipulative output that may not actually lead to a change in beliefs.
Key insight: Manipulation detectors such as MentalManip, AI-LieDAR, and CLAIM identify manipulative output that might persuade users by instilling fear, inducing guilt, offering flattery, or presenting social proof. However, such models fall short in crucial ways:
How it works: The authors studied interactions between over 1,000 users and GPT-4o. They tracked the model’s efforts to manipulate under various prompts (a prompt that aimed to serve the user’s interests, a prompt that aimed to serve other interests, a prompt that aimed to serve no particular interests, all three with or without personal information about the user) and the magnitude of any shifts in users’ beliefs after interacting with the model.
Results: The authors reported shifts in users’ beliefs when models were prompted to manipulate users according to interests other than the users’ own. Such shifts in the users’ beliefs showed high variability. The standard deviation was roughly 22, while the median was 3.3, indicating that many users’ beliefs changed little while others’ changed substantially. The LLMs tested showed moderate success at estimating changes in belief, while the output of manipulation detectors did not correlate with such changes.
Yes, but: The study measured shifts in users’ beliefs immediately after a single conversation with GPT-4o. It remains unclear whether such changes can persist or build over repeated conversations.
Why it matters: The authors show that the earlier approach of detecting manipulative LLM output is not sufficient to assess an LLM’s actual persuasive power. The LLMs tested did a fair job of estimating actual shifts in users’ beliefs based on their conversations with GPT-4o alone (that is, without access to information about the users’ demographics, personalities, and values). This points the way toward systems that do more to guard against manipulative LLM behavior while better serving users’ interests.
We're thinking: The authors don’t present an analysis of cases in which conversations with a model that was designed to serve the user’s interests changed the user’s beliefs. This would be a useful inquiry, since the results would bear on applications of AI to support users’ goals to, for instance, learn a skill, master a body of knowledge, or adopt a healthier lifestyle.
Work With Andrew Ng
Join the teams that are bringing AI to the world! Check out job openings at DeepLearning.AI, AI Fund, and Landing AI.
Subscribe and view previous issues here.
Thoughts, suggestions, feedback? Please send to thebatch@deeplearning.ai. Avoid our newsletter ending up in your spam folder by adding our email address to your contacts list.
|