---
title: "How to Create an AI Benchmark"
description: "Building your own AI benchmark: from writing a clean task definition to publishing a reproducible prompt and scoring rule. A practical walkthrough with the pitfalls to avoid."
updated: 2026-03-10
url: https://lforla.org/blog/how-to-create-an-ai-benchmark
---
Sometimes the benchmark you need does not exist. Building your own is a project — but a well-scoped one, if you follow the same discipline as a good public benchmark.

## 1. Start from a real task

The best benchmarks come from real work, not from a desire to rank models. Collect tasks that you or your users actually perform, and write them down as test cases. A benchmark of invented puzzles measures nothing you care about.

## 2. Write a task definition

For each task, record:

- **Goal** — the desired end state, stated so a program can check it.
- **Input** — exactly what the model receives.
- **Output** — the artifact to be scored.
- **Scoring rule** — how a program turns the output into a number.

Keep the definition tight enough that two people reading it would score the same output the same way.

## 3. Publish the prompt

If the prompt is secret, the benchmark is a black box. Publish the exact system prompt so others can reproduce your measurement. LFORLA benchmarks include the prompt inline on each benchmark page.

## 4. Choose a scoring scheme

Prefer automatic scoring: exact match, a reference metric, or a deterministic rubric. Where judgment is needed, use multiple graders and publish inter-grader agreement. Normalize the final score if you want your benchmark to sit alongside others on a [leaderboard](/leaderboard).

## 5. Split public and held-out data

Publish enough examples that people can run the benchmark, but hold back a set that is never released. This is what keeps the benchmark honest as models train on public data — see [what benchmarks don't tell you](/blog/what-ai-benchmarks-dont-tell-you).

## 6. Test the benchmark on yourself

Run your own benchmark against at least three models before publishing. If they all score the same, the task is too easy. If none can make progress, it may be under-specified. The existing [benchmarks catalogue](/benchmarks) is a good reference for the level of documentation to aim for.

## 7. Publish everything

Publish the task set, the prompt, the scoring code and the reference results. A benchmark nobody can reproduce is just an opinion with a decimal point.

## Common mistakes

- **Too small.** A dozen examples cannot separate close models.
- **Leaked tasks.** Public examples that models have already seen.
- **Ambiguous scoring.** Two reasonable people get different numbers.
- **No baseline.** Without a reference point, a score means nothing.

## Getting it onto LFORLA

LFORLA is a platform for exactly this. Package your benchmark, run it locally with the [lforla-eval CLI](/cli), and publish it so its results appear on the [AI agent leaderboard](/leaderboard). Before you start, read [what makes a good AI benchmark](/blog/what-makes-a-good-ai-benchmark) — it will save you a rewrite.
