Model benchmarks
Loading…
Scenarios in this run
| Kind |
Scenario |
What it tests |
Agent |
LLM judge |
Leaderboard — this run
| Rank |
Model |
Pass |
Rate |
Implement |
Chat |
Collab |
Arena |
CAD |
Tokens |
TTFT |
tok/s |
LLM judge |
Total time |
Leaderboard — all time
| Rank |
Model |
Runs |
Wins |
Avg rate |
Best rate |
Latest |
Per-scenario matrix
Structural pass/fail and timing per model. When a scenario uses an LLM deliverable judge, the second line shows that verdict.
Run history
| Run |
Suite |
Date (UTC) |
Models |
Top model |