Generalist’s GEN-1.5 Lets Robots Try a Task After Watching Once
Generalist’s new model treats a demonstration as temporary instruction, potentially reducing task-specific training—but its reported one-shot results remain far from dependable execution.
Listen to this story
The audio brief
Story brief
3 key pointsGeneralist AI’s GEN-1.5 turns a short robot demonstration into a temporary behavioral prompt: after watching 3–12 seconds of video plus language and sensor context, the model can attempt a new manipulation task without parameter updates. In a 10-task evaluation, that produced 59% average success, versus 83% after 10 training steps using roughly five minutes of task data. The result points to a lower-friction...
- 01
GEN-1.5 retains about 30 seconds of video, language, sensor, and robot-position context while acting.
- 02
Success averaged 59% across 10 tasks, rising to 83% after 10 training steps and roughly 50 demonstrations per task.
- 03
A separate held-out task reached 66.5% after one training step; that result is not comparable with the 10-task averages.
Generalist AI has released GEN-1.5, a robot foundation model that can use one brief demonstration to attempt some physical tasks without changing its underlying parameters. The setup offers a different way to teach a robot: show it an action, then let it try rather than first building a task-specific training run.
The demonstration acts as a prompt
For some tasks, the demonstration lasts three to 12 seconds. GEN-1.5 processes video, language, sensor readings and robot-position information, retaining about 30 seconds of context. Generalist calls this physical prompting: the demonstration supplies sensory and movement information that the model uses to infer its actions.
What the model holds in context
- Video of the demonstrated action.
- Language, sensor data and the robot’s position information.
- About 30 seconds of context while the robot acts.
In this one-shot mode, the model’s underlying parameters stay fixed. That separates an immediate attempt based on temporary context from the company’s few-shot training mode, which uses training steps and task-specific data.
Across 10 short manipulation tasks, Generalist reported 59% average success from one demonstration without additional training.
Generalist reported 83% average success after 10 training steps using about five minutes of data, or roughly 50 demonstrations, per task.
The reliability gap remains substantial
The 59% and 83% figures come from the same 10-task evaluation. Separately, Generalist reported 66.5% success after one training step and one minute of data on a held-out task; that result is not directly comparable with the 10-task averages. The listed tasks included opening jars, unzipping a pencil pouch, retrieving money from a purse and sweeping with a brush.
Generalist describes the tasks as relatively simple and short-horizon, and says one-shot skills remain less reliable than fine-tuned models. The 59% result shows the model can make one-shot attempts on that limited evaluation; it does not match the fine-tuned result reported for the same task set.
Pretraining is meant to make the brief example count
Generalist says it pretrained GEN-1.5 for more than eight months on physical-interaction data from homes, warehouses, factories and other settings. Its stated goal is to cut the task-specific data and training needed before a robot can attempt a new skill.
The company also says the model can combine separate demonstrations into longer behaviors, imitate some actions shown by human hands and improvise with unfamiliar objects or tools. It reported a simulation-to-real transfer in which a demonstration recorded in simulation prompted a physical robot, despite no simulation data in GEN-1.5’s pretraining. Those claims broaden the proposed interface, but the launch’s quantitative evidence remains the short manipulation-task evaluation.
Sources
- theaiinsider.techGeneralist AI Releases GEN-1.5 Robot Foundation Model That Learns From a Single Demonstration