
New benchmark shows AI agents struggle with real-world office tasks
Recent tests suggest that AI agents may have trouble handling real-world office tasks. In a large benchmark, humans scored about 80.7% accuracy, while the best AI setup reached only 68.7%, with many agents performing much lower, especially on hard tasks. Another study found leading AI models answered fewer than one in four real workplace questions correctly, showing a gap between lab results and actual job performance. Researchers say this may be due to something called the 'Data Association Gap,' where AI struggles to connect information from messy and changing files. Some new methods may help improve results, but so far, evidence suggests AI agents still have a lot of work to do before they can reliably help with office workflows.













