In Part One of this series, I proposed thinking about enterprise AI strategy as a bell curve. Some employees need relatively simple AI tools. A much larger group can benefit from customized assistants and workflows, and a smaller group of advanced users can work with agents and more sophisticated automation.
In short, not everyone needs the same level of AI to benefit.
This approach also creates natural boundaries around where more advanced AI enters the organization. We don't need autonomous agents operating everywhere simply because the technology exists. We can introduce greater capability where the business case, the people, and the underlying processes are ready.
As those capabilities improve, we need to answer another question: How do we build an AI strategy today when we don't really know what the technology will be capable of tomorrow?
This has become a very practical problem. AI systems are getting better at completing longer and more complicated tasks. Research summarized in the 2026 International AI Safety Report1 found that frontier AI agents evaluated on software tasks achieved roughly 50% success on software tasks that would take a human a little over two hours. To reach approximately 80% reliability, however, the tasks had to be much shorter, around 25 minutes. In the same report, research from METR also suggests that the length of tasks frontier models can complete autonomously has historically been increasing quickly, although estimates of the precise rate vary.
For business leaders, the important point isn't the technical measurement. It is the gap between capability and reliability. An AI system may become capable of doing something well before an organization is comfortable allowing it to do that thing independently.
We also don't know which models or vendors we will be using several years from now, how quickly agents will improve, or exactly how the legal and regulatory environment will develop. Trying to build a strategy around precise predictions about any of those things is probably not particularly useful.
There is, however, something much closer to home that organizations can control: how much authority they give AI.
From Capability to Authority
An AI agent is simply an AI system that can take multiple steps or use other tools to complete a task, rather than only giving you an answer. As these systems become more capable, I think one of the most useful distinctions for enterprise leaders is the difference between capability and authority.
Consider contract review. An AI tool might start by identifying unusual clauses or comparing a contract against a playbook. It could then draft fallback language or recommend a negotiating position. So far, a lawyer is still deciding what happens next. But connect that same AI to other systems, and it might be able to update the contract record, send proposed language to the counterparty, or, eventually, accept terms. The underlying technology may be similar, but the organization's risk has changed because the AI has been given increasing authority to affect the business.
This is why I don't think it is particularly useful to describe an AI application simply as "approved." Approved for research? Approved to draft? Approved to recommend a course of action? Approved to change a company record or communicate externally? Those are very different decisions.
A simple way to think about that progression is Inform, Prepare, Recommend, Act and Commit.
At the Inform stage, AI finds and explains information. At Prepare, it creates work product for someone else to review. At Recommend, it proposes a decision or next action. At Act, it can take a defined action within a company system. At Commit, it can create a legal, financial, external or otherwise difficult-to-reverse business consequence.
The point is not that companies should avoid the higher levels. There may be tremendous value in allowing AI to act. The point is that moving up the authority scale should be a deliberate business decision rather than an automatic consequence of the technology becoming more capable.
Capability and authority are two different decisions.
“Capability and authority are two different decisions.”
This framework also gives us a relatively simple way to think about controls. As AI receives more authority, four questions become increasingly important: What can it access? How much can it do before something stops it? Where does a person need to review or approve? And can we stop or undo what it did?
These aren't particularly technical questions. Organizations already ask versions of them when delegating authority to people or deploying important systems. AI changes some of the mechanics, but the underlying management concepts should be familiar.
Access: What Does the AI Actually Need?
Organizations have been managing technology permissions for decades. What changes with AI is the ability of one system to combine information and actions across multiple applications as part of the same task.
Imagine an AI that can read an employee's email, search the document management system, access the CRM and send external messages. Each permission might make sense individually. Together, they create something much more powerful.
There is also a good chance AI will expose permission problems that already exist. Microsoft has warned2 that Copilot can surface information a user is technically permitted to access, including information available because of stale or overly broad permissions.
For an AI agent, then, the relevant question isn't necessarily, "What can this employee access?" It is, "What does this AI need to access to perform this particular job?"
A contract-review agent might need the contract, an approved playbook, and a clause library. It probably doesn't need the lawyer's entire mailbox, every matter in the document management system and unrestricted ability to communicate externally.
The nature of the access matters too. Reading a record is different from changing it. Preparing an email is different from sending it. Identifying an error is different from correcting the official record.
This becomes more important because AI introduces a security problem that looks a little different from the attacks most business leaders are accustomed to. Researchers and AI developers have demonstrated that AI agents can be manipulated by malicious instructions hidden inside the webpages, documents, emails, or other information they are asked to review. This form of indirect prompt injection3 can cause an AI system to reveal information or take actions the user or organization never intended.
In controlled testing by NIST's Center for AI Standards and Innovation4, agent-hijacking attacks succeeded an average of 57% of the time across five test scenarios on a single attempt. When researchers repeated each attack 25 times, the average success rate rose to 80%. These were deliberately adversarial experiments designed to identify security weaknesses, not estimates of how often enterprise AI systems will be compromised in practice.
The business implication is more important than the technical attack. An AI can have completely legitimate access to company systems and still be manipulated into using that access in an unintended way.
Microsoft's EchoLeak vulnerability5 provided a real product example of the same general problem. Researchers demonstrated that, under certain conditions, a malicious email could influence Microsoft 365 Copilot and cause limited internal information available to the user to be exfiltrated. Microsoft fixed the vulnerability, and the research cited here does not establish that it was exploited against customers in the wild.
What is interesting isn't simply that there was a security vulnerability. Every major technology platform has vulnerabilities. The intriguing part is that information the AI was supposed to consume could also influence what the AI did.
That is why access should be based on what the AI needs for the job, not simply everything the employee using it can reach. And permission to read should not automatically include permission to change, send or execute.
Scale: How Far Can a Mistake Travel?
People make mistakes. Software makes mistakes. AI will make mistakes. What changes with AI is how quickly the same mistake can potentially be repeated.
An AI system doesn't need to make a spectacular error to create a significant problem. A fairly ordinary error repeated across hundreds or thousands of records, transactions or communications can become significant very quickly.
We already know how to manage this kind of risk in other contexts. A new employee doesn't typically receive unlimited spending authority. Financial systems impose transaction limits. Software releases can be staged rather than pushed immediately to everyone.
AI authority can be bounded in similar ways. A company can decide how many records an AI can change, how many external messages it can send, how much money it can commit, how many times it can retry a failed action, or how long it can continue working without checking in.
The right limits will depend on the use case. A system updating low-risk internal records can reasonably have different boundaries from one initiating payments or communicating with customers.
This is another place where the Bell Curve Strategy carries forward. I previously proposed that we shouldn't give everyone the most sophisticated AI simply because it is available. We can take the same approach to scale.
If an agent can process 10,000 transactions, that doesn't mean its first production assignment needs to be 10,000 transactions. It can start with 10, then 100, then more as the organization develops evidence that it performs reliably. The same approach can apply to business units, contract types, matters, or other categories of work.
There is a related question that deserves more attention as agents become operational: how do we actually stop a "rogue" agent?
A "kill switch" sounds straightforward until an AI is connected to other systems. There may already be transactions in a queue, active credentials, emails waiting to send or downstream processes that have been triggered. Turning off the model may not stop everything it has set in motion.
For higher-authority systems, organizations need a way to stop the authority to act, not merely the ability to generate another response. Depending on the workflow, that could mean cancelling queued actions, revoking credentials, freezing transactions or disabling external communications.
The objective isn't to prevent AI from operating at scale. The more useful principle is to earn scale rather than assume it.
Oversight: A Human in the Loop Is Not Enough
"Keep a human in the loop" has been standard advice since generative AI first entered the enterprise. It remains good advice, but we need to be more precise about what it means.
Having a human engaged in a workflow is not necessarily the same thing as providing meaningful human oversight.
In a randomized experiment6 involving 2,784 participants, researchers found that people were less likely to correct incorrect AI suggestions when doing so required additional effort. Pre-existing attitudes toward AI also mattered: participants who viewed AI more favorably were more likely to accept erroneous recommendations, while more skeptical participants detected errors more reliably.
That result isn't hard to imagine in practice. If an AI system produces good work repeatedly, people begin to trust it. As that trust grows, review can gradually become approval. Anyone who has reviewed repetitive work knows that if the first 50 outputs look right, the 51st probably doesn't receive quite the same scrutiny as the first.
Requiring a person to approve every AI action isn't necessarily the answer either. That can create approval fatigue and eliminate much of the value of automation. A better approach is to concentrate human judgment at the points where the AI is about to create a meaningful consequence.
In a legal workflow, that might be immediately before an external communication is sent, confidential information leaves an approved environment, a filing is submitted, a legal position is taken, money is committed or an important record is permanently changed.
This is a subtle but important shift. Instead of thinking only about keeping a human "in the loop," think about where human authorization belongs in the workflow.
The quality of that review matters too. Imagine an AI analyzes a complex matter and gives the reviewer a polished recommendation followed by two buttons: Approve or Reject. Technically, a human is in the loop. But what is that person actually reviewing?
For higher-impact decisions, the reviewer may need to see the underlying source material, relevant exceptions, what the AI has already done, and what it proposes to do next. Otherwise, we may be asking a person to validate an AI conclusion using only the information the AI itself selected and summarized.
For lawyers, this isn't merely an operational question. Formal Opinion 512 from the American Bar Association7 makes clear that using generative AI does not displace existing professional obligations involving competence, confidentiality, communication, candor, supervision, and reasonable fees. Delegating work to AI does not delegate responsibility for the work.
There is also a longer-term issue that deserves attention. If AI increasingly performs the work, will the people responsible for reviewing it remain capable of recognizing when it is wrong?
The previously mentioned 2026 International AI Safety Report1 also summarizes a different study examining whether sustained reliance on AI can affect human performance over time. In that clinical study, clinicians' unaided tumor-detection performance was approximately 6% lower after several months of using AI-assisted diagnosis.
That is a context-specific finding and should not be extrapolated directly to lawyers or other professionals. But it raises a legitimate business question: if AI performs a critical task every day, organizations may eventually need to think not only about technical backup systems, but also about how to preserve the human skills needed when the technology is unavailable or wrong.
This is another reason I continue to believe AI strategy should be designed around people rather than simply around the technology. The goal isn't to remove people from work wherever possible. It is to decide where AI adds leverage, where human judgment still matters, and how to preserve the skills necessary to exercise that judgment.
Reversibility: What Happens When We Need to Stop?
Most enterprise AI conversations focus on deployment. What are we implementing? What can we automate? How quickly can we scale?
As AI becomes more deeply embedded in business processes, we also need to ask what happens when we need to go the other direction.
There are really two issues here. The first is whether we can stop or undo what the AI does. The second is whether the business can continue if the AI itself becomes unavailable or changes.
At the individual action level, reversibility can be fairly simple. At the system level, it means being able to revoke an agent's credentials, cancel queued work, and stop its authority independently of the underlying AI model.
The second issue is business continuity.
In June 2025, OpenAI experienced an infrastructure incident in which ChatGPT error rates peaked at approximately 35% and API error rates at approximately 25%, with the highest-impact period lasting about six hours. That same month, a separate Google Cloud incident affected Agent Platform Online Inference and other services for approximately seven and a half hours.
This doesn't mean AI providers are uniquely unreliable. Technology fails and the pace of change has put these companies into perpetual startup mode. Downtime should be expected. The important question is what the business has built around the dependency.
If an AI writing assistant is unavailable for six hours, employees can probably work another way. If AI has become the mechanism through which contracts are processed, customer requests are routed, or an important control is performed, the same outage has a very different consequence.
Critical AI workflows therefore need some form of degraded mode. The fallback doesn't have to provide normal productivity. It needs to provide whatever minimum level of service the business cannot afford to lose.
Continuity also matters when the technology changes rather than fails.
AI providers are upgrading and retiring models quickly. The technical details aren't especially important for most business leaders. The implication is: the AI underneath a business process can change even when the company's own process does not.
A newer model may be better overall and still behave differently on the organization's particular work. Important workflows should therefore be tested against the model actually running them, with meaningful changes triggering another look at performance.
Organizations should also have a sense of what it would take to move a critical workflow elsewhere. Can the organization take its prompts, evaluations, knowledge, configuration and other important assets with it? What would break? How long would the transition take?
There is nothing wrong with choosing a preferred AI platform. Standardization can reduce tool sprawl, improve training and make governance easier. But there is a difference between having a preferred platform and having a platform the organization cannot leave.
Reversibility, in other words, isn't just about undoing an AI action. It is about preserving options.
Governance Has to Move When the Authority Moves
Traditional enterprise governance tends to operate on a calendar: annual reviews, quarterly committees, scheduled audits, and contract renewals.
AI can change meaningfully between those meetings.
An application that was reviewed six months ago may still have the same name and stated use case while being a very different system today.
Imagine Legal approves an AI application for contract review. Over the next six months, the vendor upgrades the underlying model. Someone connects it to the document management platform. Memory is enabled. The team increases the number of contracts it can process automatically. Then a new feature allows the AI to email proposed revisions.
The use case is still called "contract review," but the authority given to the system has changed considerably.
That is why AI governance needs an event-triggered component in addition to normal periodic reviews. The question isn't whether every software update needs to go back through a governance committee. It is whether something has changed the AI's practical authority or the consequences of an error.
A new connector or write permission should get attention. So should moving from drafting to sending, materially increasing transaction or volume limits, adding persistent memory, allowing one AI system to delegate work to another, introducing a new category of sensitive information, or making a critical business process dependent on the system. A significant incident or near miss should also prompt another look. Changes in law, regulation, court rules or professional obligations can do the same.
This suggests that inventories of AI applications will eventually tell us only part of what we need to know. "Contract AI tool" isn't a particularly useful risk description. We need to understand what the system can actually do.
What can it see? What can it change? How much can it do? Where can it communicate? What does it remember? What requires approval? What happens if it is wrong? What happens if it stops working?
Those questions are likely to remain useful even as today's models are replaced by much more capable ones.
Governance should move when authority moves.
“Governance should move when authority moves.”
Operating Under Conditions of Uncertainty
When enterprise generative AI first emerged, much of the strategic challenge was getting organizations comfortable enough to experiment. We are moving into a different stage now. AI is becoming part of everyday work, while agents are beginning to move from experiments into actual business processes.
McKinsey's 2026 global AI survey8 found that 40% of respondents at organizations with more than $1 billion in annual revenue reported scaling AI agents, up from 27% the prior year. Separately, Deloitte9 surveyed 3,235 IT and business leaders across 24 countries and found that only 21% said their organizations had a mature governance model for agentic AI. Because the studies surveyed different populations, the percentages should not be directly compared. Taken together, however, they illustrate the environment companies are navigating: agent adoption is accelerating while many organizations are still building the governance needed to manage it.
I don't think the answer is to slow AI adoption until all of the uncertainty disappears. It won't.
The better approach is to separate what AI can do from what the organization has authorized it to do.
As we move along that spectrum from "Who needs what level of AI?" to "How much authority should AI have?" access, scale, oversight and reversibility become more important. And movement doesn't have to go in only one direction. If an AI system proves reliable, the organization can expand its authority. If the model changes, the use case changes, performance deteriorates, or a new risk emerges, that authority can contract.
We are unlikely to predict the next several years of AI development particularly well. Fortunately, enterprise strategy doesn't require us to.
We can use better models without automatically giving them more access. We can adopt more capable agents without immediately allowing them to act independently. We can automate more work while preserving human authority over the decisions that matter.
So instead of asking simply whether an AI system is approved, I think the more useful question is:
Approved to do what, with what access, and with what authority?
“Approved to do what, with what access, and with what authority?”
That is a question we can continue answering no matter how capable the technology becomes.
Capability can move quickly. Authority can move at the speed the organization chooses.
Resources
- International AI Safety Report (2026). internationalaisafetyreport.org; Full report (PDF). ↩a↩b
- Microsoft. Data, Privacy, and Security for Microsoft 365 Copilot. Microsoft Learn. ↩
- OpenAI. Understanding prompt injections: a frontier security challenge (November 2025). openai.com. ↩
- National Institute of Standards and Technology. Technical Blog: Strengthening AI Agent Hijacking Evaluations. NIST. ↩
- SecurityWeek. EchoLeak: AI attack enabled theft of sensitive data via Microsoft 365 Copilot. SecurityWeek. ↩
- Beck, Jacob, Stephanie Eckman, Christoph Kern, and Frauke Kreuter. "Bias in the Loop: How Humans Evaluate AI-Generated Suggestions." Harvard Data Science Review (2026). Harvard Data Science Review. ↩
- American Bar Association, Standing Committee on Ethics and Professional Responsibility. Formal Opinion 512. Formal Opinion 512 (PDF). ↩
- McKinsey & Company. The state of AI in 2026. McKinsey. ↩
- Deloitte. The State of AI in the Enterprise 2026. Deloitte. ↩
Put the thinking into practice
Decide what AI is allowed to do.
UpLevel Ops helps legal teams define responsible AI strategy, choose practical use cases, establish governance, and match authority to the work.