In today’s edition of Data Points, you’ll learn about our top headlines, and more:
- New rules for Android AI agents in EU
- Nemotron 3 Embed takes the embedding model crown
- NotebookLM is now Gemini Notebook
- How Hugging Face used AI to fight off an AI attack
But first:
Kimi K3, a giant open-weights model to rival Fable and Sol
Kimi released K3, a 2.8-trillion-parameter model that includes native vision, a one-million-token context window, and a sparse mixture-of-experts design that activates only 16 of 896 experts. Two architectural innovations, Kimi Delta Attention and Attention Residuals, improve information flow across sequence length and depth, translating to roughly 2.5 times better scaling efficiency than K2. Full weights will be released by July 27, 2026; the API is live now at $3 per million input tokens and $15 per million output tokens, though caching drops effective input costs to $0.30 per million tokens for coding tasks. Kimi demonstrated K3 on autonomous GPU compiler development, chip design completed in 48 hours, and automated video editing, but the company concedes it still lags behind Claude Fable 5 and GPT 5.6 Sol in head-to-head performance. (Kimi)
Open-weights Inkling optimizes for versatility and fine-tuning
Thinking Machines Lab released Inkling, an open-weights mixture-of-experts model with 975 billion total parameters and 41 billion active ones. It was trained on 45 trillion tokens across text, images, audio, and video, opting for a deliberately broad curriculum rather than optimization for any single benchmark or domain. The model’s key differentiator is controllable thinking effort: Developers can dial compute spend up or down depending on task complexity. Inkling hits the same performance as Nvidia’s Nemotron 3 Ultra on agentic coding tasks while using one-third the tokens. Full weights are available for fine-tuning on Thinking Machines’ Tinker platform. A smaller variant, Inkling-Small, runs 12 billion active parameters for lower-cost deployments. The company’s pitch is foundation-model flexibility: strong enough across enough domains that builders can customize it for their specific needs. (Thinking Machines)
Google must give AI agents more access to Android in EU
The European Commission issued two new antitrust rules forcing Google to share anonymized search data with competitors and open Android to rival AI services. The EU found that non-Google AI agents couldn’t function on Android at the same level as Gemini, so Google must now enable voice-activation and background tasks for alternatives. Data sharing begins by January 2027. Google’s chief of global affairs argued the rules could backfire by exposing European search queries to unfamiliar companies without adequate anonymization, threatening user privacy and business secrets. The new rules could open competition for AI agents on mobile devices, but also are an unusually broad intervention in the market. (Associated Press)
Nvidia’s lightweight, quantized embedding models top benchmarks
Nvidia released Nemotron 3 Embed, a collection of three open embedding models targeting different deployment scales. The flagship 8B model leads the RTEB leaderboard at 78.5 percent, while two 1B variants handle production deployments—one in BF16 for general use, another in NVFP4 quantization tuned for Blackwell GPUs. All three support 32,000-token context windows and retrieve across multilingual text and code. Nvidia built the 1B models by pruning a 3B base through two rounds of structured compression and distillation from the 8B, keeping 99 percent of BF16 accuracy in the quantized version while halving memory use and doubling throughput on Blackwell. Automation Anywhere, Boomi, and IBM report early gains in agentic retrieval and question answering. The models are live on Hugging Face and deployable as Nvidia NIM microservices. (Hugging Face)
Google rebrands its smart notebook
Google renamed NotebookLM as Gemini Notebook, formalizing the tool’s role within its broader AI ecosystem. The rebranding reflects three years of expansion from a basic note-taking experiment into a multi-modal research platform that now supports audio, video, and interactive content. Gemini Notebook is already available in the Gemini app and will soon appear directly in Google Search. The core functionality remains unchanged for now, but the team hinted at additional features coming soon, including folder organization. (Google via X)
Hugging Face turns to GLM 5.2 to fend off AI agent attack
An autonomous AI agent breached Hugging Face’s production infrastructure through a malicious dataset, moving laterally across systems for an entire weekend without detection. When the incident response team tried to analyze the attack using commercial frontier models, safety guardrails blocked every forensic query—treating the defenders’ real exploit data the same way they would treat a live attack. The agent executed thousands of actions through short-lived sandboxes, harvesting cloud credentials and reaching multiple internal clusters, all without human guidance. Hugging Face ultimately completed its forensic analysis using GLM 5.2, an open-weight model running on its own infrastructure, because it was the only option that wouldn’t refuse to process attacker artifacts. The incident exposes a basic asymmetry: Defenders operating under enterprise governance hit safety controls that don’t constrain attackers running uncensored models, turning AI tooling into a potential single point of failure during the exact moment security teams need it most. (VentureBeat)
Want to know more about what matters in AI right now?
Read the latest issue of The Batch for in-depth analysis of news and research.
Last week, Andrew discussed how AI was automating coding and other tasks, allowing professionals to focus on higher-level responsibilities, and how this shift was leading to the emergence of “full-stack” roles in wider fields like marketing and recruiting, as well as engineering.
“As AI increasingly automates coding, it frees up developers to spend time on high-level software development tasks traditionally reserved for senior engineers, like deciding on technical architecture and participating in scoping product requirements… This is why I’m confident there will be rising demand for broad AI engineering skills.”
Read Andrew’s letter here.
Other top AI news and research covered in depth:
- GPT-Live pairs full-duplex voice models with a reasoning model (GPT-5.5) to enable seamless conversation and complex problem-solving.
- A German court ruled that Google's AI Overviews defamed two businesses, holding the tech giant responsible for AI-generated search results.
- Robin is automating drug discovery by matching chemical processes of known drugs to research on diseases, potentially accelerating medical breakthroughs.
- Researchers at MIT and Carnegie Mellon built Puppet to measure how AI models can manipulate users’ beliefs, shedding light on ethical AI usage.
Start today with DeepLearning.AI Pro!
DeepLearning.AI’s first-ever subscription plan for our entire course catalog includes foundational classics like the Machine Learning and Deep Learning specialization plus seminars on the latest tools and frameworks you need.
As a Pro Member, you’ll immediately enjoy access to:
- Nearly 200 AI short and long courses from Andrew Ng and industry experts
- Labs and quizzes to test your knowledge
- Projects to share with employers
- Certificates to testify to your new skills
- A community to help you advance at the speed of AI
Enroll now to lock in a year of full access for $25 per month paid upfront, or opt for month-to-month payments at just $30 per month. Both payment options begin with a one-week free trial.
Explore Pro’s benefits and start building today!
Data Points is produced by human editors with AI assistance.