AI Agents Tested on Real Office Tasks: Strong at Coding, Stumble on Human Judgment
Summary
- • AI agents given full laptop access excelled at coding tasks but struggled with nuanced human language.
- • Agents handled ambiguity well: misspelled names, non-responding colleagues, and multi-document budget analysis.
- • Critical failure: agents confused employees on temporary leave with candidates for permanent layoffs.
- • Over 120,000 tech jobs cut in 2026 as companies cite AI as the driving force behind workforce reductions.
Details
Empirical office agent test setup
Used Claude Cowork app on a laptop with Slack, spreadsheets, and other pre-configured apps; tasks adapted from CMU and OpenAI published benchmark frameworks.
Agents excel at coding/programming tasks
Tasks solvable by writing computer programs were completed reliably; structured logic and code generation play to current agents' core strengths.
Agents struggle with UI navigation and language nuance
Navigating Chrome and interpreting nuanced human communication were identified as persistent weak points in real-world deployment.
Graceful handling of structured ambiguity
Correctly messaged 'Sarah Johnson' despite a 'Sara Johnson' typo; added non-responding colleague to spreadsheet without stalling — showing improved robustness.
Critical failure: leave vs. layoff confusion
Agent classified employees on temporary leave as potential permanent cuts without querying return dates — a consequential HR error a human would avoid.
Tacit knowledge gap (Stanford / NBER)
Stanford and National Bureau of Economic Research: AI is least capable of replacing idiosyncratic expertise accumulated through experience that is never digitized.
120,000+ tech jobs cut in 2026 citing AI
Per Layoffs.fyi: 200+ tech companies cut ~120,000 jobs in 2026, with Meta, Oracle, and Cloudflare citing AI as the driving force.
Cloudflare CEO: AI to replace middle management
After cutting ~1,100 employees, Cloudflare CEO predicted AI replacement of workers in middle management, finance, and marketing.
Open-source CMU + OpenAI benchmark materials
Fabricated spreadsheets, memos, staff lists, and prompts published openly to enable standardized real-world agent evaluation across models.
AI office agent capability assessment: what agents can and cannot do in real workplace scenarios as of mid-2026.
What This Means
This investigation adds crucial empirical nuance to the AI-replacing-workers narrative: agents are genuinely capable in structured, programmable contexts but still make consequential errors when tasks require the kind of contextual judgment humans develop through experience. The leave-vs-layoff confusion is a microcosm of a broader risk — agents produce plausible-sounding outputs that can be badly wrong in HR, legal, or financial contexts. As 120,000+ tech jobs have already been cut citing AI in 2026, this testing suggests organizations deploying agents without human oversight in judgment-heavy roles may be moving faster than current capabilities justify. The clearest takeaway: AI is a reliable co-pilot for coding tasks and a risky autonomous decision-maker wherever tacit knowledge is required.
