HeyGen Releases Code2Video to Judge AI-Made Motion Graphics

The benchmark aims to distinguish video code that renders from motion graphics that feel designed, using a company-trained judge across five visual criteria.

By 2 min read
HeyGen Releases Code2Video to Judge AI-Made Motion Graphics
HeyGen Releases Code2Video to Judge AI-Made Motion Graphics

Listen to this story

The audio brief

About 1:40
0:001:40
Read transcript
HeyGen has released Code2Video, a benchmark designed to test whether coding agents can make motion graphics that feel intentionally designed—not merely videos that render without errors. The benchmark starts with 168 human-written briefs built around product-launch beats: a hook, a product introduction, a feature demonstration, social proof, and a call to action. An agent turns each brief into structured code, and HyperFrames, HeyGen’s open-source H-T-M-L-to-video framework, renders the result deterministically. Instead of giving every clip one overall quality score, HeyGen’s company-trained judge compares two videos from the same brief across five dimensions: engagement, prompt intent, composition, temporal quality, and craft. Those head-to-head preferences are then used to create Elo rankings, showing which outputs are preferred more often. Across 1,464 evaluation samples, HeyGen reports that its judge matched human preferences 82 percent of the time, compared with 75 percent for the general-purpose vision-language models it tested. When the judge expressed stronger confidence, HeyGen says agreement with human raters rose above 95 percent. The benchmark is aimed at failures that code execution alone misses: elements colliding or drifting outside the frame, animation that happens all at once, and videos that satisfy the brief while still looking flat or poorly balanced. The important constraint is that these results apply to videos rendered through HyperFrames and judged against HeyGen’s own five criteria. The open question is how well that preference system transfers to other rendering stacks and creative tasks.

Story brief

3 key points

HeyGen has released Code2Video, a benchmark for testing coding agents that create motion-graphics videos, with evaluation focused on visual preference rather than executable code alone. The benchmark covers 168 human-written briefs and compares outputs through head-to-head Elo rankings across engagement, intent, composition, temporal quality, and craft. HeyGen reports its specialized judge matched human preferences...

  1. 01

    Code2Video contains 168 briefs covering launch-video structures such as hooks, demos, social proof, and calls to action.

  2. 02

    HeyGen evaluated 1,464 samples and reports its judge exceeded 95% human agreement when expressing stronger confidence.

  3. 03

    The judge produces separate preference probabilities for five dimensions instead of one aggregate quality score.

HeyGen wants Code2Video to catch the gap between an animation that runs and one that looks designed. The new benchmark evaluates coding agents that generate motion-graphics video from code across timing, composition and craft. HeyGen says its specialized judge agreed with human preferences more often than the general-purpose vision-language models it tested, though those are company-reported benchmark results.

From a written brief to a rendered clip

Code2Video contains 168 human-created motion-graphics briefs spanning product-launch beats, including hooks, product introductions, feature demonstrations, social proof and calls to action. An agent turns a brief into structured code, then HyperFrames, HeyGen’s open-source HTML-to-video framework, renders the result deterministically into video.

HeyGen’s reported judge validation
82%HeyGen’s judge

HeyGen reports 82% agreement with human preferences across 1,464 evaluation samples.

75%Tested general-purpose vision-language models

HeyGen reports 75% agreement for the general-purpose vision-language models it tested.

Five separate verdicts

The judge compares two videos made from the same prompt across engagement, prompt intent, composition, temporal quality and craft. It returns a separate preference probability for each axis, rather than one blended quality score. HeyGen says its judge’s agreement with human raters rose above 95% when it expressed stronger confidence.

The motion-design problems it seeks to expose

  • Text or interface elements that collide, clip, or drift beyond the frame.
  • Animation that moves at once instead of building a readable sequence.
  • Videos that satisfy a brief but remain visually flat or unbalanced.

A score for a defined creative task

Code2Video uses head-to-head comparisons to produce Elo rankings, a system that reflects how often one output is preferred over another. That puts visual preference at the center of the test, not simply whether an agent produced valid code. Its results measure videos rendered through HyperFrames and judged against HeyGen’s five defined criteria for this code-to-video task.

Sources

  1. heygen.comCode2Video Benchmark: Evaluating AI Agents on Motion Design | HeyGen Research

Loading discussion...

YOUR READING SPACE

Notifications

HeyGen Releases Code2Video to Judge AI-Made Motion Graphics | Superpower Daily