Machine Learning Research

613 Posts

Image illustrates two scenarios of AI chat responses, affecting user reliance on AI vs. professional help.
Machine Learning Research

Measuring Models’ Manipulation: Researchers at MIT and Carnegie Mellon built Puppet to gauge models' influence on users' beliefs

Providers of large language models stand to benefit by building models that spur user engagement, but users may bear a cost in undue influence on their world views.
Diagram shows Robin system generating hypotheses, analyzing data, and producing therapeutic suggestions.
Machine Learning Research

Put the Lab in the Loop: Robin automates drug discovery, matching chemical processes of known drugs to research on diseases

An AI agent proposed new medical uses for established drugs nearly autonomously — uses that were supported by experiments on isolated human cells — with human input only to name diseases to be treated and run the AI-proposed lab experiments. 
Diagram illustrates real-time audio processing with async checks and turn-based message checks in models.
Machine Learning Research

One Model Talks, Another One Thinks: GPT-Live pairs full-duplex voice models with a reasoning model (GPT-5.5) on the backend

ChatGPT’s voice mode now listens and speaks at the same time, passing harder questions posed to the conversational model to a reasoning model in the background.
Architecture of the character, word, and sentence level of Brain2Qwerty (encoder, aligner, and LLM)
Machine Learning Research

Text Without Typing: Researchers at Meta and other institutions built Brain2Qwerty v2 to generate sentences from brain waves

Imagine you are typing a sentence. But instead of a keyboard, joystick, or eye tracker, a device surrounding your head reads your intentions and generates that sentence on screen.
Flowchart illustrates DSpark's multi-stage process, highlighting hardware-aware prefix scheduler and token flow.
Machine Learning Research

DeepSeek’s DSpark Gains Velocity: DeepSeek open sourced a speculative decoding module that speeds up text generation without losing accuracy

DeepSeek built a speculative decoding module that speeds up its production models’ text generation by more than 50 percent without sacrificing accuracy, then made its technique open source.
Graph compares image generation, editing, latency, and price for Nano Banana models against competitors.
Machine Learning Research

Google Pairs Nano Banana Update With Video API: Gemini's image (Nano Banana 2 Lite) and video (Gemini Omni Flash) models, built to work together

Google built a low-cost, high-throughput pipeline for developers working with media, combining Gemini’s fastest image model yet with a similarly speedy multimodal model that can turn images into video with synchronized sound.
Robot arm successfully places pot on cloth, demonstrating reward verification and scoring effectiveness.
Machine Learning Research

Better Reward Models for Robots: Inside RoboReward, a family of vision-language reward models that train robots to take action

When you’re training a robot via reinforcement learning, a handcrafted reward function is labor-intensive to build but often dispenses rewards more effectively than a general-purpose reward model based on a vision-language model. Researchers built reward models that narrowed the gap.
The table shows MAI-Thinking-1 leading in several benchmarks, compared to other AI models.
Machine Learning Research

Microsoft Strikes Out on Its Own: Microsoft revealed MAI-Thinking-1, a Claude Sonnet 4.6-sized reasoning model developed without distillation

Microsoft, once OpenAI’s exclusive partner and still a major reseller of other companies’ AI models, built its own reasoning model from scratch.
Six charts show Fugu and Fugu Ultra scoring highest, marked by red bars, on various tasks and benchmarks.
Machine Learning Research

Fugu Blends Models Task by Task: Sakana debuted dedicated orchestrator models, Fugu and Fugu-Ultra, that spawn Claude, Gemini, and GPT agents

Models that orchestrate the activities of other models and agents achieved state-of-the-art performance on a variety of benchmarks, outperforming the best individual models working alone.
Detailed eagle talons grip a metallic logo amid a clear blue background, symbolizing control and power.
Machine Learning Research

GPT-5.6 Lands in Limbo: OpenAI previewed three GPT-5.6 Models (Sol, Terra, and Luna), wider release coming soon

OpenAI announced a preview of its GPT-5.6 family, including a top-tier model comparable to Claude 5 Mythos — but so far it’s available only to users that are selected by the U.S. government.
Anthropic Opus 4.8 Leaps Forward: Claude Opus 4.8 won back the high-performance crown for Anthropic, pending wider availability of its Mythos-class models
Machine Learning Research

Anthropic Opus 4.8 Leaps Forward: Claude Opus 4.8 won back the high-performance crown for Anthropic, pending wider availability of its Mythos-class models

Anthropic mid-2026 update of Opus held the the top of a leading intelligence ranking for about a week, only to be overtaken by Claude Fable 5.
Flowchart of an ESMC-6B model with sequence encoding layers, language model, and diffusion transformer output.
Machine Learning Research

Biological Molecules as Language: ESMFold2 approaches AlphaFold 3 performance but with an open, Transformer-based architecture

Google’s AlphaFold models pioneered the task of finding the shapes of biologically active molecules, opening new pathways for drug development.
AFM 3 Core model architecture visualizes DRAM and NAND processes in AI with focus on sparsely-activated LLM operations.
Machine Learning Research

Large-Model AI for Apple Devices: 2026's Apple Foundation Models bring AI to MacBooks, iPhones, and the cloud

The third generation of Apple Foundation Models — fruit of Apple’s collaboration with Google — introduces a variation on the mixture-of-experts architecture that runs on local devices. 
AI performance chart shows GLM, GPT models competing in reasoning, coding benchmarks. Models highlight performance.
Machine Learning Research

Top Agentic Performance, Low Cost: GLM-5.2, designed for coding and long-running agentic jobs, now the top open model

Z.ai released an open-weights model that rivals proprietary leaders for autonomous agentic tasks.
Flowchart illustrates the POPE method, transitioning from guided to unguided problem-solving in reinforcement learning.
Machine Learning Research

Reinforcement Learning With Hints: Privileged On-Policy Exploration (POPE) trains models to expand on partial solutions

Reinforcement learning can’t train a model to solve a difficult problem if the model doesn’t discover all the right steps.
Load More

Subscribe to The Batch

Stay updated with weekly AI News and Insights delivered to your inbox