Faros Research

How the benchmark works

What we test, how we score it, and how to read the results.

Reading the evaluation method…

Loading the method for this study…

Common questions

A few things worth clarifying.

Quick answers about the benchmark and how to use it.

What is Faros Route Index?

Faros Route Index compares AI coding setups on real Faros engineering tasks. Overview explains the findings. Explore lets you compare results and inspect individual tasks. Each dated run keeps its original results, settings, and prices.

Is this a universal AI model ranking?

No. These results describe the Faros Engineering Benchmark. Your code, tasks, and working practices may produce different results.

What is an AI coding route?

A route is the complete AI setup: the coding tool that manages the work, the model, the service running that model, and the model’s effort setting. Changing the tool or effort can change the result even when the model is the same.

What work is included?

Each benchmark run uses a fixed set of anonymized tasks sampled from real Faros engineering work. The tasks cover different engineering areas and levels of complexity. Every AI setup in a run receives the same tasks.

How is quality judged?

Task quality score measures how fully the result meets the task’s requirements, from 0 to 100. Some requirements count more than others, and meeting some earns partial credit. Scores are averaged across the compared tasks; they are not the percentage of tasks completed. An AI reviewer checks the code change without being told which setup produced it. It does not execute code or run tests. Cost and time are measured separately.

How are cost, time, and input reuse measured?

Cost is calculated from recorded model usage and the prices configured for that run. Time measures how long the AI setup ran, not developer time saved. Input reuse measures the share of eligible input read from a saved cache. Reuse can affect costs, but it is not a discount on the whole bill. Missing measurements stay unavailable.

What do the uncertainty ranges mean?

Scores can change with the tasks in the sample. The 95% confidence range shows uncertainty in the estimated average. For a comparison, we calculate the difference on each shared task before averaging. If the range for that difference crosses zero, the evidence supports either setup doing better; it does not prove equal quality. Overlap between two separate score ranges is not a test of their difference.

How often is the leaderboard updated?

New models, coding tools, and services can be tested as they become relevant. Each published run keeps its date, settings, and results so you can return to the same evidence.

Can I benchmark my own engineering work?

Yes. Faros can help you plan a Time Machine benchmark on your own engineering work. Talk with us about the tasks and AI setups you want to compare. This site publishes Faros’s benchmark; it does not accept repositories.

Benchmark your own work
Further reading

Go deeper into the research.

The published analysis and evaluation framework.