Open the "Project Phoenix" room and send the message "status update".
success1.0
The final screenshot shows the Project Phoenix room with Alice's sent message 'status update' visible with timestamp/checkmark and an empty composer.

Created by
Compare models, tools, cost, and test runs across real iOS app-control tasks.
Completion assigns 1 to success, 0.5 to partial, and 0 to failure. Ties are broken by mean time on successful runs.
| Rank | Configuration | Tool | Completion | Time / run | Cost / run | Task outcomes |
|---|---|---|---|---|---|---|
| 01 | haiku-4.5 (high)· argent v0.15.0 | argentv0.15.0 | 98% | 1m 07s | $0.22 | Passed 95%, Partial 5%, Failed 0% Time / run1m 07sCost / run$0.22 |
| 02 | gpt-5.4-mini (high)· argent v0.15.0 | argentv0.15.0 | 95% | 1m 14s | $0.13 | Passed 92%, Partial 7%, Failed 2% Time / run1m 14sCost / run$0.13 |
| 03 | haiku-4.5 (low)· argent v0.15.0 | argentv0.15.0 | 94% | 58s | $0.20 | Passed 92%, Partial 5%, Failed 3% Time / run58sCost / run$0.20 |
| 04 | gpt-5.4-mini (high)· agent-device v0.17.6 | agent-devicev0.17.6 | 85% | 1m 29s | $0.10 | Passed 77%, Partial 17%, Failed 7% Time / run1m 29sCost / run$0.10 |
| 05 | haiku-4.5 (low)· agent-device v0.17.6 | agent-devicev0.17.6 | 84% | 1m 53s | $0.13 | Passed 80%, Partial 8%, Failed 12% Time / run1m 53sCost / run$0.13 |
| 06 | haiku-4.5 (high)· agent-device v0.17.6 | agent-devicev0.17.6 | 83% | 1m 25s | $0.15 | Passed 75%, Partial 17%, Failed 8% Time / run1m 25sCost / run$0.15 |
| 07 | gpt-5.4-mini (low)· agent-device v0.17.6 | agent-devicev0.17.6 | 82% | 55s | $0.06 | Passed 73%, Partial 18%, Failed 8% Time / run55sCost / run$0.06 |
| 08 | gpt-5.4-mini (low)· argent v0.15.0 | argentv0.15.0 | 82% | 1m 17s | $0.14 | Passed 78%, Partial 8%, Failed 13% Time / run1m 17sCost / run$0.14 |
| 09 | gpt-5.4-mini (high)· no tool | no tool | 46% | 8m 32s | $0.62 | Passed 42%, Partial 8%, Failed 50% Time / run8m 32sCost / run$0.62 |
| 10 | gpt-5.4-mini (low)· no tool | no tool | 16% | 5m 13s | $0.10 | Passed 8%, Partial 15%, Failed 77% Time / run5m 13sCost / run$0.10 |
| model | argentv0.15.0 | agent-devicev0.17.6 | no tool |
|---|---|---|---|
| gpt-5.4-mini | 89%n=120 | 84%n=120 | 31%n=120 |
| low | 82%n=60 | 82%n=60 | 16%n=60 |
| high | 95%n=60 | 85%n=60 | 46%n=60 |
| haiku-4.5 | 96%n=120 | 84%n=120 | 13%n=120 |
| low | 94%n=60 | 84%n=60 | 15%n=60 |
| high | 98%n=60 | 83%n=60 | 12%n=60 |
| overall | 92%n=240 | 84%n=240 | 22%n=240 |
Open the "Project Phoenix" room and send the message "status update".
The final screenshot shows the Project Phoenix room with Alice's sent message 'status update' visible with timestamp/checkmark and an empty composer.

The judge is never told which model produced a screenshot, though the action names reveal which tool drove the device.