Skip to main content
An evaluation is one question you want answered with human judgments. HeyBee supports two kinds:
  • Which option wins? Compare named candidates: models, prompts, checkpoints, or anything else you can produce outputs from. This is the shape of a benchmark, and also of a release check: put the new version against the one you ship and see if it actually wins.
  • Which settings win? Explore generation parameters, such as temperature or guidance scale, and let HeyBee recommend the best configuration and show which parameters actually matter.
Create one from Evaluations with New evaluation on a tablet or desktop. The wizard walks through four steps: basics, grouping, feedback and settings, and review.

Basics

Pick a name, an optional note on what you are comparing, and the output type: text, image, audio, or video.
Video must be MP4 with H.264 video and AAC audio, the format nearly every editor and ffmpeg produce by default. HeyBee checks files on upload and tells you exactly what to re-export, but it does not convert them for you.
The evaluation is created in the workspace named in the page URL. Use the workspace switcher before opening the wizard if you need a different destination. Personal evaluations stay private; Shared evaluations follow the workspace roles.

Group your uploads

The shared input, by default prompt, is how HeyBee knows which outputs compete against each other. Two outputs are only ever compared when they share the same input value. For example, if you upload variant_a and variant_b outputs for 50 prompts, voters always compare A against B on the same prompt, never across different prompts.
The shared input field, output type, feedback method, and criteria cannot be changed after creation. The name, notes, and voting requirements stay editable until the evaluation is completed or archived.

Feedback

The feedback step separates two decisions:
  • Compare outputs asks voters to choose A, B, tie, both bad, or cannot tell. Score outputs gives each output a numeric score and is available for benchmarks, optimizations, and RLHF training loops.
  • Overall quality collects one combined judgment. Multiple criteria collects separate judgments for dimensions such as Motion, Detail, and Prompt adherence.
These choices produce regular A/B, multi-criteria A/B, graded, multi-criteria graded, and mixed evaluations for every objective. Add up to five criteria when using Multiple criteria. Each criterion has:
  • a name voters see
  • a weight used for the combined ranking
  • optional guidance for voters, such as “Which one moves more smoothly and naturally?”
Results separates every criterion, showing a leaderboard for comparisons and score summaries for graded feedback. Voters judge all criteria in one pass over each pair, so three criteria still cost one credit per comparison.
Two or three criteria is the sweet spot. Beyond that, each vote takes noticeably longer and voter attention drops.

Voting requirements

Control how carefully voters must look before they can submit:
  • Minimum viewing seconds holds the submit button until enough time has passed.
  • Minimum playback percent (audio) requires listening to both clips.
  • Allow playback speed control lets voters speed up long clips.
  • Require reviewer notes asks for a short written justification with every vote.

Readiness and lifecycle

An evaluation moves through Draft, Ready, Collecting, Paused, Complete, and Archived. Collection can start once three checks pass:
  • at least two outputs exist
  • every output has a shared input value
  • at least one input has outputs from two different candidates
Setup & collect lists anything missing, each with its fix. You can pause and resume collection at any time. Completed and archived evaluations become read-only but stay fully readable and exportable.