Routine Work Becomes the New Stress Test
Simular announced that its AI agent, Sai, achieved a 73 percent success rate on the OSWorld 2.0 benchmark. This result was revealed on Thursday alongside the release of the updated evaluation suite. The new standard focuses heavily on complex, real-world professional workflows. It moves beyond simple commands to test sustained autonomy.
Breaking news
Eufy Unveils Local AI Home Security Ecosystem at IFA
The Rapid Evolution of Data Center Security in the AI Era
The High-Voltage Risks Facing Modern AI Data Centers
Apple’s New CEO Renames Lake Ontario To Lake America In Maps AppThe benchmark consists of 108 distinct tasks designed to mimic daily computer usage. These scenarios require agents to navigate operating systems for extended periods. Most tasks take a skilled human over one hour to finish. The difficulty level reflects genuine workplace demands rather than abstract puzzles. Simular claims this performance represents the current state of the art.
The shift to OSWorld 2.0 highlights a growing need for agents that handle mundane but critical jobs. Previous benchmarks often tested isolated actions or short sequences. The new standard demands consistency across long, multi-step processes. Agents must manage file systems, applications, and user interfaces simultaneously. They cannot rely on shortcuts or partial completion.
Can AI Finally Handle the Boring Stuff?
Simular emphasized that Sai’s score places it at the top of current rankings. The company noted that such high-level performance is rare in this specific domain. The agent demonstrates robustness when dealing with unexpected interface changes. It maintains focus without human intervention for significant durations. This capability suggests a maturation in general-purpose computer operation.
The phrase posterity will find it ludicrouscaptures the sentiment behind this progress. Early AI struggles with basic tasks seem almost comical now. As models improve, the baseline for acceptable performance rises quickly. What once required hours of manual clicking now happens automatically. The gap between human effort and machine execution narrows rapidly.
Experts observe that 73 percent is a strong indicator of reliability. It implies that three out of four attempts succeed fully. For enterprise adoption, this level of accuracy reduces oversight requirements. Companies can deploy agents for repetitive administrative duties with greater confidence. The technology moves closer to practical utility in office environments.
Frequently Asked Questions
How many tasks are included in the OSWorld 2.0 benchmark? The updated benchmark contains 108 specific tasks. Each task simulates a lengthy professional workflow. These scenarios typically exceed one hour in duration for human users.
What does a 73 percent success rate indicate about Sai? It indicates that Sai successfully completes nearly three-quarters of all attempted tasks. This performance marks it as the current leader in this category. The result reflects high reliability in complex, multi-step operations.
When was this benchmark update released? Simular released the OSWorld 2.0 update on Thursday. The announcement coincided with the publication of Sai’s performance metrics. This timing allowed for immediate comparison against previous standards.


