---
title: "How We Benchmark Reinforcement Learning Agents"
description: "Reproducibility starts with methodology. Here is how LFORLA computes, normalizes and publishes agent scores."
updated: 2026-05-28
url: https://lforla.org/blog/how-we-benchmark-rl-agents
---
Reproducible evaluation is the hardest part of comparing reinforcement learning agents. This post explains the methodology behind every LFORLA score.

## Environments and episodes

Each benchmark defines a fixed set of environments, initial states and episode budgets. Runs execute a **deterministic number of episodes** with seeded initial states, so two runs of the same agent produce comparable trajectories.

## Score normalization

Raw returns are normalized per environment to a **0–100 scale**, where 0 is a random baseline and 100 is the documented expert maximum. This lets us aggregate across environments with very different reward scales.

## Aggregating across benchmarks

Benchmark-level rankings use the mean of normalized environment scores:

1. Normalize each environment return.
2. Average across environments within a benchmark.
3. Only agents that completed the full episode budget are ranked.

## Guardrails against overfitting

- **Held-out environments.** Some environments are not disclosed until submission time.
- **Computational honesty.** We record the inference budget and enabled flags alongside every run.
- **Version pinning.** Every run records exact library and environment versions.

## Reproducing a score

Every leaderboard row links back to the run record, which includes the exact command, versions and raw per-episode returns. If you can't reproduce a score, that's a bug — [tell us](/company).
