Modelspublished

Generalist’s GEN-1.5 Lets Robots Try a Task After Watching Once

Generalist’s new model treats a demonstration as temporary instruction, potentially reducing task-specific training—but its reported one-shot results remain far from dependable execution.

By 3 min read
Generalist’s GEN-1.5 Lets Robots Try a Task After Watching Once

Listen to this story

The audio brief

About 1:27
0:001:27
Read transcript
Generalist AI’s GEN-1.5 lets a robot watch a short demonstration, then attempt a new physical task without changing the model’s underlying parameters. The example can last just three to twelve seconds. While acting, GEN-1.5 combines the video with language, sensor readings, and the robot’s position, keeping about thirty seconds of that context in view. Generalist calls this “physical prompting”: the demonstration temporarily serves as an instruction, rather than becoming a new training run. The appeal is a simpler teaching workflow. Instead of collecting task-specific data and updating the model first, an operator could show the robot what to do and let it try immediately. But the initial results show a meaningful reliability gap. Across ten short manipulation tasks—including opening jars, unzipping a pencil pouch, retrieving money from a purse, and sweeping with a brush—one-shot prompting succeeded 59 percent of the time. After ten training steps, using roughly five minutes of data, or about fifty demonstrations per task, average success rose to 83 percent. Generalist also reported 66.5 percent after one training step on a separate held-out task, but that result isn’t directly comparable with the ten-task averages. The company says GEN-1.5 was pretrained for more than eight months on physical-interaction data from homes, warehouses, factories, and other settings. The key constraint is clear: one-shot attempts are possible, but the evidence so far covers a limited set of short-horizon tasks, not dependable general-purpose execution.

Story brief

3 key points

Generalist AI’s GEN-1.5 turns a short robot demonstration into a temporary behavioral prompt: after watching 3–12 seconds of video plus language and sensor context, the model can attempt a new manipulation task without parameter updates. In a 10-task evaluation, that produced 59% average success, versus 83% after 10 training steps using roughly five minutes of task data. The result points to a lower-friction...

  1. 01

    GEN-1.5 retains about 30 seconds of video, language, sensor, and robot-position context while acting.

  2. 02

    Success averaged 59% across 10 tasks, rising to 83% after 10 training steps and roughly 50 demonstrations per task.

  3. 03

    A separate held-out task reached 66.5% after one training step; that result is not comparable with the 10-task averages.

Generalist AI has released GEN-1.5, a robot foundation model that can use one brief demonstration to attempt some physical tasks without changing its underlying parameters. The setup offers a different way to teach a robot: show it an action, then let it try rather than first building a task-specific training run.

The demonstration acts as a prompt

For some tasks, the demonstration lasts three to 12 seconds. GEN-1.5 processes video, language, sensor readings and robot-position information, retaining about 30 seconds of context. Generalist calls this physical prompting: the demonstration supplies sensory and movement information that the model uses to infer its actions.

What the model holds in context

  • Video of the demonstrated action.
  • Language, sensor data and the robot’s position information.
  • About 30 seconds of context while the robot acts.

In this one-shot mode, the model’s underlying parameters stay fixed. That separates an immediate attempt based on temporary context from the company’s few-shot training mode, which uses training steps and task-specific data.

A first attempt versus added task data
59%One-shot prompting

Across 10 short manipulation tasks, Generalist reported 59% average success from one demonstration without additional training.

83%After 10 training steps

Generalist reported 83% average success after 10 training steps using about five minutes of data, or roughly 50 demonstrations, per task.

The reliability gap remains substantial

The 59% and 83% figures come from the same 10-task evaluation. Separately, Generalist reported 66.5% success after one training step and one minute of data on a held-out task; that result is not directly comparable with the 10-task averages. The listed tasks included opening jars, unzipping a pencil pouch, retrieving money from a purse and sweeping with a brush.

Generalist describes the tasks as relatively simple and short-horizon, and says one-shot skills remain less reliable than fine-tuned models. The 59% result shows the model can make one-shot attempts on that limited evaluation; it does not match the fine-tuned result reported for the same task set.

Pretraining is meant to make the brief example count

Generalist says it pretrained GEN-1.5 for more than eight months on physical-interaction data from homes, warehouses, factories and other settings. Its stated goal is to cut the task-specific data and training needed before a robot can attempt a new skill.

The company also says the model can combine separate demonstrations into longer behaviors, imitate some actions shown by human hands and improvise with unfamiliar objects or tools. It reported a simulation-to-real transfer in which a demonstration recorded in simulation prompted a physical robot, despite no simulation data in GEN-1.5’s pretraining. Those claims broaden the proposed interface, but the launch’s quantitative evidence remains the short manipulation-task evaluation.

Sources

  1. theaiinsider.techGeneralist AI Releases GEN-1.5 Robot Foundation Model That Learns From a Single Demonstration