- Which option wins? Compare named candidates: models, prompts, checkpoints, or anything else you can produce outputs from. This is the shape of a benchmark, and also of a release check: put the new version against the one you ship and see if it actually wins.
- Which settings win? Explore generation parameters, such as temperature or guidance scale, and let HeyBee recommend the best configuration and show which parameters actually matter.
Basics
Pick a name, an optional note on what you are comparing, and the output type: text, image, audio, or video.Video must be MP4 with H.264 video and AAC audio, the format nearly every editor and ffmpeg produce by default. HeyBee checks files on upload and tells you exactly what to re-export, but it does not convert them for you.
Group your uploads
The shared input, by defaultprompt, is how HeyBee knows which outputs compete against each other. Two outputs are only ever compared when they share the same input value.
For example, if you upload variant_a and variant_b outputs for 50 prompts, voters always compare A against B on the same prompt, never across different prompts.
Feedback
The feedback step separates two decisions:- Compare outputs asks voters to choose A, B, tie, both bad, or cannot tell. Score outputs gives each output a numeric score and is available for benchmarks, optimizations, and RLHF training loops.
- Overall quality collects one combined judgment. Multiple criteria collects separate judgments for dimensions such as Motion, Detail, and Prompt adherence.
- a name voters see
- a weight used for the combined ranking
- optional guidance for voters, such as “Which one moves more smoothly and naturally?”
Voting requirements
Control how carefully voters must look before they can submit:- Minimum viewing seconds holds the submit button until enough time has passed.
- Minimum playback percent (audio) requires listening to both clips.
- Allow playback speed control lets voters speed up long clips.
- Require reviewer notes asks for a short written justification with every vote.
Readiness and lifecycle
An evaluation moves through Draft, Ready, Collecting, Paused, Complete, and Archived. Collection can start once three checks pass:- at least two outputs exist
- every output has a shared input value
- at least one input has outputs from two different candidates

