HeyGen Releases Code2Video to Judge AI-Made Motion Graphics
The benchmark aims to distinguish video code that renders from motion graphics that feel designed, using a company-trained judge across five visual criteria.
Listen to this story
The audio brief
Story brief
3 key pointsHeyGen has released Code2Video, a benchmark for testing coding agents that create motion-graphics videos, with evaluation focused on visual preference rather than executable code alone. The benchmark covers 168 human-written briefs and compares outputs through head-to-head Elo rankings across engagement, intent, composition, temporal quality, and craft. HeyGen reports its specialized judge matched human preferences...
- 01
Code2Video contains 168 briefs covering launch-video structures such as hooks, demos, social proof, and calls to action.
- 02
HeyGen evaluated 1,464 samples and reports its judge exceeded 95% human agreement when expressing stronger confidence.
- 03
The judge produces separate preference probabilities for five dimensions instead of one aggregate quality score.
HeyGen wants Code2Video to catch the gap between an animation that runs and one that looks designed. The new benchmark evaluates coding agents that generate motion-graphics video from code across timing, composition and craft. HeyGen says its specialized judge agreed with human preferences more often than the general-purpose vision-language models it tested, though those are company-reported benchmark results.
From a written brief to a rendered clip
Code2Video contains 168 human-created motion-graphics briefs spanning product-launch beats, including hooks, product introductions, feature demonstrations, social proof and calls to action. An agent turns a brief into structured code, then HyperFrames, HeyGen’s open-source HTML-to-video framework, renders the result deterministically into video.
HeyGen reports 82% agreement with human preferences across 1,464 evaluation samples.
HeyGen reports 75% agreement for the general-purpose vision-language models it tested.
Five separate verdicts
The judge compares two videos made from the same prompt across engagement, prompt intent, composition, temporal quality and craft. It returns a separate preference probability for each axis, rather than one blended quality score. HeyGen says its judge’s agreement with human raters rose above 95% when it expressed stronger confidence.
The motion-design problems it seeks to expose
- Text or interface elements that collide, clip, or drift beyond the frame.
- Animation that moves at once instead of building a readable sequence.
- Videos that satisfy a brief but remain visually flat or unbalanced.
A score for a defined creative task
Code2Video uses head-to-head comparisons to produce Elo rankings, a system that reflects how often one output is preferred over another. That puts visual preference at the center of the test, not simply whether an agent produced valid code. Its results measure videos rendered through HyperFrames and judged against HeyGen’s five defined criteria for this code-to-video task.
Sources
- heygen.comCode2Video Benchmark: Evaluating AI Agents on Motion Design | HeyGen Research
Reader comments
Newest comments first. Replies stay oldest first.