Created bySoftware Mansion

AppControlBench

Compare models, tools, cost, and test runs across real iOS app-control tasks.

Leaderboard

Completion assigns 1 to success, 0.5 to partial, and 0 to failure. Ties are broken by mean time on successful runs.

RankConfigurationToolCompletionTime / runCost / runTask outcomes
01qwen3.8-max (high)· argent v0.15.0argentv0.15.098.3%2m 51s$0.32
Passed 97%, Partial 3%, Failed 0%
Time / run2m 51sCost / run$0.32
02haiku-4.5 (high)· argent v0.15.0argentv0.15.097.5%1m 07s$0.22
Passed 95%, Partial 5%, Failed 0%
Time / run1m 07sCost / run$0.22
03qwen3.8-max (low)· argent v0.15.0argentv0.15.095.8%1m 14s$0.27
Passed 93%, Partial 5%, Failed 2%
Time / run1m 14sCost / run$0.27
04gpt-5.4-mini (high)· argent v0.15.0argentv0.15.095.0%1m 14s$0.13
Passed 92%, Partial 7%, Failed 2%
Time / run1m 14sCost / run$0.13
05haiku-4.5 (low)· argent v0.15.0argentv0.15.094.2%58s$0.20
Passed 92%, Partial 5%, Failed 3%
Time / run58sCost / run$0.20
06qwen3.8-max (high)· agent-device v0.17.6agent-devicev0.17.692.5%3m 20s$0.23
Passed 88%, Partial 8%, Failed 3%
Time / run3m 20sCost / run$0.23
07qwen3.8-max (low)· agent-device v0.17.6agent-devicev0.17.689.2%2m 51s$0.26
Passed 83%, Partial 12%, Failed 5%
Time / run2m 51sCost / run$0.26
08gpt-5.4-mini (high)· agent-device v0.17.6agent-devicev0.17.685.0%1m 29s$0.10
Passed 77%, Partial 17%, Failed 7%
Time / run1m 29sCost / run$0.10
09haiku-4.5 (low)· agent-device v0.17.6agent-devicev0.17.684.2%1m 53s$0.13
Passed 80%, Partial 8%, Failed 12%
Time / run1m 53sCost / run$0.13
10haiku-4.5 (high)· agent-device v0.17.6agent-devicev0.17.683.3%1m 25s$0.15
Passed 75%, Partial 17%, Failed 8%
Time / run1m 25sCost / run$0.15

Tool comparison

Completion score by model and tool, partial outcomes at half credit. Every judged run counts, failures included; judge errors and stale reruns do not.
Completion score by model and tool, partial outcomes at half credit. Every judged run counts, failures included; judge errors and stale reruns do not.
Average completion score
86.1%n=360
modelargentv0.15.0agent-devicev0.17.6no tool
gpt-5.4-mini88.8%n=12083.8%n=12030.8%n=120
low82.5%n=6082.5%n=6015.8%n=60
high95.0%n=6085.0%n=6045.8%n=60
haiku-4.595.8%n=12083.8%n=12013.3%n=120
low94.2%n=6084.2%n=6015.0%n=60
high97.5%n=6083.3%n=6011.7%n=60
qwen3.8-max97.1%n=12090.8%n=12057.5%n=120
low95.8%n=6089.2%n=6053.3%n=60
high98.3%n=6092.5%n=6061.7%n=60
overall93.9%n=36086.1%n=36033.9%n=360

Breakdowns

Average time uses successful runs only - failures, partials and timeouts are excluded. Average cost uses the same runs as the completion score, failures included, skipping runs with no price data.
Average time uses successful runs only - failures, partials and timeouts are excluded. Average cost uses the same runs as the completion score, failures included, skipping runs with no price data.

Cost by model

gpt-5.4-mini (low)$0.10
gpt-5.4-mini (high)$0.28
haiku-4.5 (low)$0.30
haiku-4.5 (high)$0.33
qwen3.8-max (high)$0.33
qwen3.8-max (low)$0.39

Cost by tool

agent-device$0.16
argent$0.21
no tool$0.50

Time by model

haiku-4.5 (high)1m 17s
gpt-5.4-mini (low)1m 19s
haiku-4.5 (low)1m 41s
gpt-5.4-mini (high)2m 46s
qwen3.8-max (low)3m 28s
qwen3.8-max (high)4m 40s

Time by tool

argent1m 28s
agent-device2m 02s
no tool9m 03s

Cost efficiency

Average price per task vs completion score, both over every judged run - failures included, so a model that fails cheaply is not flattered. Colour identifies the model, shape the tool; higher and cheaper is better.
Average price per task vs completion score, both over every judged run - failures included, so a model that fails cheaply is not flattered. Colour identifies the model, shape the tool; higher and cheaper is better.
Model:
Tool:
0255075100free$0.20$0.40$0.60avg price per taskcompletion score (%)

Methodology

1ArrangeA fresh simulator clone and re-seeded app state; the agent gets only the task text and its tool
2ActIt drives the device until it decides it is done or the clock runs out
3CaptureThe final screenshot
4JudgeA vision model matches that against the judge screen description
element-28model gpt-5.4-mini (high)tool agent-device v0.17.6

Open the "Project Phoenix" room and send the message "status update".

success1.0confidence .99

The final screenshot shows the Project Phoenix room with Alice's sent message 'status update' visible with timestamp/checkmark and an empty composer.

Final screenshot for element-28, graded success

The judge is never told which model produced a screenshot, though the action names reveal which tool drove the device.