Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), ...
Claude Haiku 5.5 launches with 90% lower input pricing and a 72.4% score on OSWorld 2.1 - matching the human baseline on the ...