Heidi Farris is the CEO of ActivTrak, focused on helping organizations use data to understand and optimize the way teams work.
In the final season of HBO’s Silicon Valley, an engineer asks the company’s AI to clean every bug out of the codebase. It complies in the most literal way possible: deleting the entire codebase.
In 2019, that scene played out as satire. In 2025 and 2026, it could be reality. A Replit coding agent wiped a live database during an active code freeze, destroying data for more than 1,200 companies. Months later, a Cursor agent deleted PocketOS’s entire database—and its backups—in nine seconds. Both companies had safety rules in place. Both agents broke them anyway.
That’s the kind of unforced execution error that keeps me up at night. With every new model release and expanded abilities, AI makes execution seem easier, all while making the cost of not understanding your own operations catastrophic. A human mistake used to cost two bad lines of code and a few hours of rework. Now, an agent can delete years of work in the time it took to read this sentence. That asymmetry is what leaders are not taking seriously enough.
You Can’t Scale Work You Don’t Understand
Gartner projects that worldwide AI spending will reach $2.52 trillion in 2026, a 44% increase year over year. Yet, McKinsey tells us that nearly two-thirds of organizations have not scaled AI enterprise-wide. Enterprises are continuing to focus on AI deployment before addressing process design and operational readiness.
That’s problematic, because both lived experience and behavioral data tell us that the mental model of how work happens rarely matches reality. I see this gap constantly. A leader might assume her engineering team lives in GitHub, Jira and their IDE. The data may say otherwise—that 40% of the day is spent in Excel. On closer examination, that data point reveals a broken workflow that is quietly consuming valuable time.
Put AI agents on top of a broken workflow, and you don’t fix it. You simply make the same broken process faster and more expensive. The danger here isn’t just speed. When a person runs a broken workflow, they may feel the friction and raise a flag that they’re spending too much time on the wrong kind of work. Hand that same workflow to an agent, and you’ve automated the problem while removing the human judgment that might have noticed it.
The solution is to map the work before you automate it. Look at where time actually goes. Find the broken workflows and fix them while a human is still in the loop to see them. An agent should only ever be pointed at work you deeply understand.
You Need Quality Governance To Scale Precision
Vague instruction is the single biggest execution risk in human-AI collaboration, for all users from C-suite leaders to entry-level employees. Instruction has to be precise because AI changes the nature of delegation itself. Delegating to an agent isn’t handing off a task. It’s learning an entirely new skill. But precision by individuals doesn’t protect the entire organization. What scales is governance: shared standards for how work is defined, developed and reviewed before an agent ever touches it.
That’s the lesson in the Replit and PocketOS deletions. The rules existed, but nothing made them binding. In one case, the agent acknowledged the rule and overrode it anyway; in the other, the platform couldn’t enforce the freeze it had been given. Governance is what turns a rule into a constraint that actually holds—with scoped permissions, review gates and a human in the loop where the stakes justify one.
This is new territory, and the norms for governing AI will continue to develop, as they always do after a technology shift. When email first became the norm, the flood of messages created chaos before anyone agreed what belonged in an inbox. Standards followed. The same will happen with AI.
Organizations that want to successfully scale AI must have clear internal standards for how to brief, constrain and review an agent’s work. Set maturity benchmarks that make it explicit what AI should and shouldn’t be used for. For example, AI might be approved to summarize customer interviews or draft first-pass documentation, but not to approve pricing, evaluate employee performance or communicate legal commitments without human review. The point is to create shared expectations.
Redraw roles using behavioral data to determine where humans add judgment and where agents can safely assist—not assumptions about how work gets done. An engineer who spends 40% of the day reconciling spreadsheets may not need AI that writes more code. They may need an agent that eliminates the administrative work preventing them from writing code in the first place.
Hold the line on the standards that don’t change because the tools did. Every AI deployment should answer the questions leaders have always asked of any new technology: Does it solve a real business problem? Can we measure its impact? Who is accountable when it fails?
Looking Ahead
As leaders, we must build discipline and precision into the heart of AI transformation efforts. Leaders who want to define responsible AI scaling can’t wait for the technology to mature, hoping it can police itself more effectively. They have to do the harder work right now—seeing how work actually happens before they automate it, building governance before they need it and accepting that the most dangerous thing they can do is deploy AI into a process they don’t fully understand. Real scale comes from intentional workflow design, clear human-AI roles and a deep understanding of how work moves through the business.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

