Created bySoftware Mansion

AppControlBench

Compare models, tools, cost, and test runs across real iOS app-control tasks.

Leaderboard

Completion assigns 1 to success, 0.5 to partial, and 0 to failure. Ties are broken by mean time on successful runs.

RankConfigurationToolCompletionTime / runCost / runTask outcomes
01haiku-4.5 (high)· argent v0.15.0argentv0.15.098%1m 07s$0.22
Passed 95%, Partial 5%, Failed 0%
Time / run1m 07sCost / run$0.22
02gpt-5.4-mini (high)· argent v0.15.0argentv0.15.095%1m 14s$0.13
Passed 92%, Partial 7%, Failed 2%
Time / run1m 14sCost / run$0.13
03haiku-4.5 (low)· argent v0.15.0argentv0.15.094%58s$0.20
Passed 92%, Partial 5%, Failed 3%
Time / run58sCost / run$0.20
04gpt-5.4-mini (high)· agent-device v0.17.6agent-devicev0.17.685%1m 29s$0.10
Passed 77%, Partial 17%, Failed 7%
Time / run1m 29sCost / run$0.10
05haiku-4.5 (low)· agent-device v0.17.6agent-devicev0.17.684%1m 53s$0.13
Passed 80%, Partial 8%, Failed 12%
Time / run1m 53sCost / run$0.13
06haiku-4.5 (high)· agent-device v0.17.6agent-devicev0.17.683%1m 25s$0.15
Passed 75%, Partial 17%, Failed 8%
Time / run1m 25sCost / run$0.15
07gpt-5.4-mini (low)· agent-device v0.17.6agent-devicev0.17.682%55s$0.06
Passed 73%, Partial 18%, Failed 8%
Time / run55sCost / run$0.06
08gpt-5.4-mini (low)· argent v0.15.0argentv0.15.082%1m 17s$0.14
Passed 78%, Partial 8%, Failed 13%
Time / run1m 17sCost / run$0.14
09gpt-5.4-mini (high)· no toolno tool46%8m 32s$0.62
Passed 42%, Partial 8%, Failed 50%
Time / run8m 32sCost / run$0.62
10gpt-5.4-mini (low)· no toolno tool16%5m 13s$0.10
Passed 8%, Partial 15%, Failed 77%
Time / run5m 13sCost / run$0.10

Tool comparison

Completion score by model and tool.

Average completion score
84%n=240
modelargentv0.15.0agent-devicev0.17.6no tool
gpt-5.4-mini89%n=12084%n=12031%n=120
low82%n=6082%n=6016%n=60
high95%n=6085%n=6046%n=60
haiku-4.596%n=12084%n=12013%n=120
low94%n=6084%n=6015%n=60
high98%n=6083%n=6012%n=60
overall92%n=24084%n=24022%n=240

Breakdowns

Time uses successful runs only.

Cost by model

gpt-5.4-mini (low)$0.10
gpt-5.4-mini (high)$0.28
haiku-4.5 (low)$0.30
haiku-4.5 (high)$0.33

Cost by tool

agent-device$0.11
argent$0.17
no tool$0.48

Time by model

haiku-4.5 (high)1m 17s
gpt-5.4-mini (low)1m 19s
haiku-4.5 (low)1m 41s
gpt-5.4-mini (high)2m 46s

Time by tool

argent1m 09s
agent-device1m 26s
no tool7m 52s

Cost efficiency

Average price per task vs completion score. Colour identifies the model, shape the tool; higher and cheaper is better.

Model:
Tool:
0255075100free$0.20$0.40$0.60avg price per taskcompletion score (%)

Methodology

1ArrangeA fresh simulator clone and re-seeded app state; the agent gets only the task text and its tool
2ActIt drives the device until it decides it is done or the clock runs out
3CaptureThe final screenshot
4JudgeA vision model matches that against the judge screen description
element-28model gpt-5.4-mini (high)tool agent-device

Open the "Project Phoenix" room and send the message "status update".

success1.0confidence .99

The final screenshot shows the Project Phoenix room with Alice's sent message 'status update' visible with timestamp/checkmark and an empty composer.

Final screenshot for element-28, graded success

The judge is never told which model produced a screenshot, though the action names reveal which tool drove the device.