Model benchmarks

Loading…

Leaderboard — this run

Rank Model Pass Rate Implement Chat Collab Arena CAD Tokens TTFT tok/s LLM judge Total time

Per-scenario matrix

Structural pass/fail and timing per model. When a scenario uses an LLM deliverable judge, the second line shows that verdict.