Published by Faros Research
The next AI release.
Your next decision.
Compare complete AI coding setups on real engineering work. Find the tradeoff worth investigating.
Which setup deserves a place in your workflow?
Time Machine evaluates complete AI coding setups—coding tool, model, provider and reasoning effort—on the same engineering tasks. This Index shows the results from Faros’s own work. Compare quality, estimated AI cost and runtime, then see where the answer changes by workload.
Dated study · 2026-06-25 · 197 common tasks
Where can we lower AI costs?
Codex + GPT-5.5 (High) has 35.4% lower estimated AI cost.
Compared with Claude Code + Opus 4.8 (High) on the same 197 tasks, the average quality score changes by -0.45 points out of 100. The 95% uncertainty range for that change is -3.20 to +2.30 points. The range includes zero, so the quality advantage remains unclear. That does not mean the options are equally good.
Typical cost: $1.51 for Claude Code + Opus 4.8; $1.43 for Codex + GPT-5.5. About 95% of tasks cost no more than $6.84 for the starting option and $3.21 for the alternative. Typical cost is the middle task cost when ordered from least to most expensive. These estimates use the prices recorded for this study; your costs may differ.
Explore this comparisonPreparing the evidence chart…
Comparable studies · 2026-07-27 · 187 common tasks
Which AI configurations produce the strongest work?
Claude Code + Opus 5 (Default) has the highest average quality score in these results.
Every configuration below is compared on the same tasks, across studies that use comparable measurements. The lines show uncertainty in each average. A small gap does not prove one option is better.
Explore this comparisonPreparing the evidence chart…
Dated study · 2026-07-27 · 205 common tasks
Does asking the model to think harder produce better work?
Claude Code + Opus 5 (Default) has 33.6% lower estimated AI cost.
Compared with Claude Code + Opus 5 (Max) on the same 205 tasks, the average quality score changes by +1.63 points out of 100. The 95% uncertainty range for that change is -0.72 to +3.98 points. The range includes zero, so the quality advantage remains unclear. That does not mean the options are equally good.
Explore this comparisonPreparing the evidence chart…
Dated study · 2026-06-25 · 197 common tasks
How often does each configuration produce better work?
Codex + GPT-5.5 scores higher on 79 tasks; Claude Code + Opus 4.8 on 71.
The alternative’s average score is 0.45 points lower; how often it scores higher and the size of each change are different questions. The two configurations receive equal scores on 47 of 197 tasks. These counts show how often the score changes; they do not show the size of each change or how many tasks passed tests. Equal scores do not mean identical solutions.
Explore this comparisonPreparing the evidence chart…
Dated study · 2026-06-25 · 197 common tasks
How long do we wait for a result?
Typical wait: 7.2 minutes for Claude Code + Opus 4.8; 3.3 for Codex + GPT-5.5.
About 95% of recorded AI runs finished within 37.3 minutes for the starting option and 6.0 for the alternative. Typical means the middle recorded time when ordered from shortest to longest. These are measured AI times, not developer time saved or promises about future tasks.
Explore this comparisonPreparing the evidence chart…
Dated study · 2026-06-25 · 197 common tasks
Does the choice depend on the work?
The higher-scoring configuration changes by work type.
For Data systems work (65 tasks), Claude Code + Opus 4.8 averages 51.7 out of 100 and Codex + GPT-5.5 averages 46.3. The 4 groups range from 7 to 86 tasks. Groups with fewer tasks provide less evidence. No uncertainty range for these differences is available; use the breakdown to investigate, not to assume the same pattern will hold on your work.
Explore this comparisonPreparing the evidence chart…
Dated study · 2026-06-25 · 197 common tasks
What evidence is this comparison based on?
197 of 211 tasks have recorded results from every configuration.
The run has 1,457 of 1,477 planned results; 20 are unavailable. Comparisons use the 197 tasks with results from every configuration. Availability describes the evidence collected, not whether the proposed code works.
Explore this comparisonPreparing the evidence chart…