Skip to main content
Benchmark Lab

Benchmark Lab

Marketing pages all claim the same things. The only way to compare tools honestly is to give them identical work and publish what comes back. Below are the five standard briefs we run, the exact measurements we record, and — importantly — how far through the programme we actually are.

Current status

Tools catalogued
206
Hands-on tested
0
Standard briefs
5

We are at the start of this. No tool has completed a full brief yet, so no tool on this site carries a numeric score. We would rather show an empty scoreboard than a full one we invented — every listing says plainly whether it has been tested, price-checked, or simply catalogued.

Editorial Lab Queue

Upcoming Tests

We test every tool on a standard 24-point benchmark brief. Schedule is tentative based on community interest.

Methodology
#1
SynthesiaAI Avatars
Scheduled: Sep 19, 2026 Awaiting test
#2
InVideo AIVideo Generation
Scheduled: Sep 26, 2026 Awaiting test
#3
MunchVideo Repurposing
Scheduled: Oct 03, 2026 Awaiting test
#4
Luma Dream MachineVideo Generation
Scheduled: Oct 10, 2026 Awaiting test

Want a tool tested sooner? Submit it for editorial review

The five standard briefs

Every tool in a category receives byte-identical input. No tool gets a second chance the others did not get.

  1. B1

    The podcast cut

    Video Repurposing
    Fixed input
    One fixed 60-minute two-person podcast episode, 1080p, lavalier audio.
    Task
    Produce the best five vertical short clips the tool can find, with captions burned in.
    What we record
    • Wall-clock time from upload to downloadable output
    • Caption word error rate against a human transcript
    • Speaker framing accuracy across cuts
    • How many of the five clips are genuinely publishable without edits
    • Cost in credits, converted to cost per finished clip
  2. B2

    The cinematic shot

    Video Generation
    Fixed input
    One fixed prompt: a specific camera move, subject and lighting condition.
    Task
    Generate the shot. Five attempts allowed, best result counts.
    What we record
    • Prompt adherence, scored against a written rubric
    • Temporal coherence — when artefacts first appear
    • Maximum usable clip length before the model drifts
    • Whether audio is generated natively
    • Attempts needed before one usable result, and total cost of those attempts
  3. B3

    The talking head

    AI Avatars
    Fixed input
    One fixed 150-word educational script, English.
    Task
    Generate a full-screen avatar speaking the script.
    What we record
    • Lip-sync accuracy at normal playback speed
    • Natural eye movement and blink cadence
    • Render time from prompt submission to finished MP4
    • Cost per minute of finished video at the starter tier
  4. B4

    The clean-up

    Voice & Audio
    Fixed input
    One 60-second voice recording with air-conditioning hum and room echo.
    Task
    Clean the audio and level the voice.
    What we record
    • Noise reduction without audible gating artefacts
    • Preservation of natural voice timbre
    • Processing time
  5. B5

    The dub

    Voice & Audio
    Fixed input
    One 60-second English talking-head clip.
    Task
    Dub into Spanish with voice cloning.
    What we record
    • Voice likeness to the original English speaker
    • Natural Spanish prosody and pacing
    • Whether commercial rights are included at the tier tested
    • Cost per finished audio minute

How scores are calculated

Each tested tool gets five sub-scores from 0 to 10. The overall figure is a weighted average, computed rather than hand-adjusted, so we cannot quietly nudge a favourite upward.

  • Output quality 35% — How good is the result you can actually publish
  • Speed 20% — Wall-clock time from input to usable output
  • Value for money 20% — Output quality per dollar at the entry paid tier
  • Ease of use 15% — Time to a first good result without reading documentation
  • Export freedom 10% — Watermarks, resolution caps and commercial rights

We use the full range. A 5 out of 10 is a normal, useful score — if nothing ever scored below 7, the numbers would carry no information at all.

Independence

  • We pay for our own subscriptions at the tier we test.
  • Vendors cannot pay for a score, a ranking position, or a re-test with a better result.
  • If we ever accept paid placement it will be labelled “Sponsored” and excluded from scoring entirely.
  • Where a tool has not been tested we say so, rather than filling the gap with a plausible-looking number.

See also our affiliate disclosure and editorial policy.

Benchmark Lab | Noxifera