An AI agent that takes action inside an enterprise must be stoppable by an authorized person, without relying on the model’s cooperation. But a stop button is only the interface. Behind it, the system needs mechanisms that block further actions, revoke permissions where necessary, and preserve a clear record of what has already happened.
Satya Nadella’s call to design agents that can be stopped mid-task sharpens a business question: does the enterprise control the process, or is it simply hoping the model behaves correctly? This is not a reason to freeze AI adoption. It is a reason to stop treating a successful demo as proof of production readiness.
What does Nadella’s call change in practice?
The three principles Nadella outlined, according to the report, are separation between the model and the mechanism that manages its actions, tamper-resistant records of significant actions, and the ability for an authorized person to stop the agent. Together, they make one point clear: control must be built into the system, not promised in a prompt.
A model can interpret a document, recommend a decision, and select a tool. It should not be the sole authority deciding whether to transfer money, change supplier details, or send information outside the enterprise. Those decisions require policies the system enforces even when the model gets something wrong.
Key insight: Asking an agent to stop is not a stopping mechanism. A real control must block actions even if the model keeps generating responses or a task has already been handed to another system.
This distinction matters because an agent is more than a chat window. It may invoke services, write to business systems, and create tasks that continue running after the conversation ends. Stopping the response on screen does not necessarily stop the business process.
An emergency brake does not put money back in the account
Consider an agent handling requests to update supplier details. It reads an email, compares it with existing documents, and prepares a change in the finance system. If impersonation is detected, stopping the agent should prevent the update and any actions that depend on it. If a payment has already been sent, stopping the agent will not automatically recover it.
That is why stopping, canceling a pending action, and remediating a completed action must be treated separately. They are different capabilities, sometimes owned by different people.
- Stopping prevents the agent from initiating further actions.
- Cancellation attempts to halt requests already submitted, where the receiving system supports it.
- Remediation addresses an existing outcome, such as restoring a record or initiating payment recovery.
- Resuming operations requires checking that the fault has been resolved and the agent’s permissions remain appropriate.
These behaviors need to be tested in failure scenarios, not merely described in a specification. If an agent keeps working through background tasks after it has been stopped, the enterprise has gained a reassuring button, not control.
The architecture: the model proposes, the system authorizes
Separating the model from the action-management layer lets an enterprise replace the model without rewriting every control rule. More importantly, it allows business rules to be enforced consistently: which tools are available, who can approve an exception, and which actions are prohibited outright.
In a sound implementation, an action passes through permission and policy checks before execution. Those checks can account for the action type, data sensitivity, process state, and need for human approval. For sensitive actions, permissions should be checked again close to execution, so a policy change or stop command actually takes effect.
- Permissions: A fragile implementation uses a broadly privileged shared account; a mature one gives the agent narrowly scoped permissions.
- Stopping: A fragile implementation sends a chat message; a mature one blocks actions at the execution layer.
- Approvals: A fragile implementation relies solely on the model’s judgment; a mature one applies defined business rules.
- Records: A fragile implementation relies on conversation history; a mature one maintains a protected action log.
- Recovery: A fragile implementation simply restarts; a mature one requires review and a controlled return to operation.
The difference is not between a weak model and a strong one. It is between a system that asks the model to behave and a system that limits the damage when it does not.
What belongs in the action log?
A useful record includes the agent’s identity, configuration version, requested action, policy-check result, approval granted, and response from the receiving system. It should be understandable to business process owners, not just developers.
Operational transparency does not require exposing the model’s internal reasoning. What matters is knowing what happened, what information supported it, which authorization allowed it, and what the outcome was. Logs need protection against retrospective changes, but they should also minimize personal information and business secrets. An audit requirement is not permission to copy every sensitive document into another repository.
Human oversight without a human bottleneck
Requiring human approval for every action may look safe, but it can erase the operational benefit. If an employee still examines every case in the same detail, the enterprise has added a system without materially changing its capacity to execute.
The goal is risk-based oversight. Routine, reversible actions can run automatically within defined limits. Unusual, sensitive, or difficult-to-remediate actions go to a person. Sampling is also necessary to identify failures that did not trigger an alert.
A human in the loop is not someone who approves everything. It is someone with the authority, context, and tools to intervene where their judgment changes the outcome.
For oversight to work, the employee needs to see the exception and its context, rather than read an entire conversation and guess what went wrong. This is where AI knowledge, process expertise, and management experience come together. A well-designed approval interface matters as much as model selection.
What should enterprise leaders require before deployment?
There is no need to wait for a complete industry policy framework before setting internal acceptance criteria. Finance, operations, and IT leaders should agree in advance on the agent’s boundaries, oversight responsibilities, and behavior when something fails.
- Map actions and risks: Identify what the agent can change, what information it can access, and which outcomes are difficult to reverse.
- Define authority and boundaries: Assign a process owner, specify permissions and approval conditions, and identify who is authorized to stop the agent.
- Test stopping under load: Verify what happens to active actions, pending tasks, and downstream systems when the agent is stopped.
- Deploy and measure incrementally: Expand operations only after reviewing business results, exceptions, and the cost of oversight.
This sequence avoids a common mistake: connecting an agent to production systems and trying to add governance afterward. Testing should also cover bad data, external service failures, and attempts to push the agent beyond its permissions through content it reads.
Financially, time saved is not the only measure. The calculation must include human review, exception handling, error correction, and infrastructure costs. A fast agent that creates a new review queue is not necessarily a good investment.
Internal capability matters more than a vendor promise
Microsoft Copilot Studio, platforms such as n8n, or custom development can provide a foundation for agents. No product choice removes the enterprise’s responsibility to test permissions, logging, and stopping within its specific process. A preference for Claude or another model is no substitute for assessing information security and business fit.
Enterprises should therefore build internal capability to create and manage agents while also training employees to communicate with models and recognize their limitations. These are complementary tracks: AI literacy improves how people work with AI; agent infrastructure enables processes to run within the existing work environment.
Selecting an implementation team requires more than an impressive demo. Relevant education, a grounding in research, and practical business and implementation experience help a team understand not only how to build an agent, but when not to let it act alone. This distinction is especially important for small and midsize businesses, which may lack multiple layers of internal control.
The right question for leadership is not just what the agent can do. It is who can stop it, what will actually stop, and what it will take to resume work. An enterprise that can answer those questions can expand automation with justified confidence, rather than replace control with hope.
