Benchmark Lab
Marketing pages all claim the same things. The only way to compare tools honestly is to give them identical work and publish what comes back. Below are the five standard briefs we run, the exact measurements we record, and — importantly — how far through the programme we actually are.
Current status
- Tools catalogued
- 206
- Hands-on tested
- 0
- Standard briefs
- 5
We are at the start of this. No tool has completed a full brief yet, so no tool on this site carries a numeric score. We would rather show an empty scoreboard than a full one we invented — every listing says plainly whether it has been tested, price-checked, or simply catalogued.
Upcoming Tests
We test every tool on a standard 24-point benchmark brief. Schedule is tentative based on community interest.
Want a tool tested sooner? Submit it for editorial review
The five standard briefs
Every tool in a category receives byte-identical input. No tool gets a second chance the others did not get.
- Fixed input
- One fixed 60-minute two-person podcast episode, 1080p, lavalier audio.
- Task
- Produce the best five vertical short clips the tool can find, with captions burned in.
- What we record
- Wall-clock time from upload to downloadable output
- Caption word error rate against a human transcript
- Speaker framing accuracy across cuts
- How many of the five clips are genuinely publishable without edits
- Cost in credits, converted to cost per finished clip
- Fixed input
- One fixed prompt: a specific camera move, subject and lighting condition.
- Task
- Generate the shot. Five attempts allowed, best result counts.
- What we record
- Prompt adherence, scored against a written rubric
- Temporal coherence — when artefacts first appear
- Maximum usable clip length before the model drifts
- Whether audio is generated natively
- Attempts needed before one usable result, and total cost of those attempts
- Fixed input
- One fixed 150-word educational script, English.
- Task
- Generate a full-screen avatar speaking the script.
- What we record
- Lip-sync accuracy at normal playback speed
- Natural eye movement and blink cadence
- Render time from prompt submission to finished MP4
- Cost per minute of finished video at the starter tier
- Fixed input
- One 60-second voice recording with air-conditioning hum and room echo.
- Task
- Clean the audio and level the voice.
- What we record
- Noise reduction without audible gating artefacts
- Preservation of natural voice timbre
- Processing time
- Fixed input
- One 60-second English talking-head clip.
- Task
- Dub into Spanish with voice cloning.
- What we record
- Voice likeness to the original English speaker
- Natural Spanish prosody and pacing
- Whether commercial rights are included at the tier tested
- Cost per finished audio minute
How scores are calculated
Each tested tool gets five sub-scores from 0 to 10. The overall figure is a weighted average, computed rather than hand-adjusted, so we cannot quietly nudge a favourite upward.
- Output quality 35% — How good is the result you can actually publish
- Speed 20% — Wall-clock time from input to usable output
- Value for money 20% — Output quality per dollar at the entry paid tier
- Ease of use 15% — Time to a first good result without reading documentation
- Export freedom 10% — Watermarks, resolution caps and commercial rights
We use the full range. A 5 out of 10 is a normal, useful score — if nothing ever scored below 7, the numbers would carry no information at all.
Independence
- We pay for our own subscriptions at the tier we test.
- Vendors cannot pay for a score, a ranking position, or a re-test with a better result.
- If we ever accept paid placement it will be labelled “Sponsored” and excluded from scoring entirely.
- Where a tool has not been tested we say so, rather than filling the gap with a plausible-looking number.
See also our affiliate disclosure and editorial policy.