---
title: "How to Run Your Own AI Benchmark (Step by Step)"
description: "A reproducible, step-by-step guide to benchmarking any AI model yourself: pick tasks, fix seeds and versions, compute scores, and share results you can trust."
updated: 2026-05-10
url: https://lforla.org/blog/how-to-run-your-own-ai-benchmark
---
Vendor comparisons are full of fine print. The most convincing evaluation is the one you run yourself — because you control the tasks, the settings, and the scoring. Here is a step-by-step guide.

## Step 1 — Write down the question

Every benchmark starts with a question. Not "is model A better?" but **"is model A better than model B at writing SQL for our Postgres schema?"**. A specific question makes every later step obvious.

## Step 2 — Choose between 5 and 20 tasks

Pick tasks that mirror your real workload, not the model's niche strengths. Use a mix: a few easy (baseline), a few hard (separating signal). More than 20 tasks usually adds noise, not insight.

## Step 3 — Pin everything that can be pinned

- Pin the **model versions**.
- Pin the **environment and library versions**.
- Fix the **random seed** and **temperature**.
- Use the **same prompt template** everywhere.

Without this, you are measuring temperature roulette, not model quality.

## Step 4 — Run it

Use a benchmark harness so the runs are recorded. With the [LFORLA CLI](/cli), a run looks roughly like:

```bash
lforla-eval benchmark sql-gen   --model your-model:v1   --seed 42   --save-run ./runs/sql-gen-2026-08-13/
```

The harness saves the command, versions, and per-task outputs alongside the score.

## Step 5 — Normalize, then average

Map each task score so that **0 = a random/naive baseline** and **100 = an expert target**. Averaging normalized scores (instead of raw totals) is what lets you compare across very different tasks.

## Step 6 — Publish the run record

Share the score *and* the recipe: command, versions, seed, per-task breakdown. If someone can replay your run and get the same number, your benchmark is trustworthy.

## Step 7 — Keep it small and iterative

Benchmarks rot. Re-run monthly, add new tasks, retire saturated ones. A benchmark you maintain is worth more than a headline you don't.

Ready to go deeper? Read [how we benchmark RL agents](/blog/how-we-benchmark-rl-agents) for the methodology LFORLA uses at scale.
