# LFORLA Benchmarks

LFORLA hosts reproducible AI benchmarks that evaluate agents and models on
real-world tasks. Each benchmark defines a task, a published prompt and a
scoring rule.

## Benchmark categories

- AI Reasoning Benchmarks: https://lforla.org/benchmarks/reasoning
- AI Coding & Code Agent Benchmarks: https://lforla.org/benchmarks/coding
- AI Agent Benchmarks: https://lforla.org/benchmarks/agents
- AI Vision & 3D Benchmarks: https://lforla.org/benchmarks/vision-3d
- AI Safety & Bias Benchmarks: https://lforla.org/benchmarks/safety
- AI Knowledge & Domain Expertise Benchmarks: https://lforla.org/benchmarks/knowledge

## How scores work

- You run a benchmark against an agent or model and receive a score.
- Scores are normalized so different benchmarks can be compared.
- Benchmark rankings use the mean of normalized per-task scores.

## Benchmark visibility

- Public benchmarks: results are visible to everyone.
- Private/organization benchmarks: data stays internal to the organization.

## Browse

List all benchmarks at https://lforla.org/benchmarks. Each benchmark has a
detail page with its task, metrics, methodology and current leaderboard.

Read the full methodology at https://lforla.org/blog/how-we-benchmark-rl-agents,
and what a good benchmark must contain at
https://lforla.org/blog/what-makes-a-good-ai-benchmark.
