How AI governance works
Is it safe to let an AI agent act on company systems?
Short answerYes, when the agent holds only the access its task needs, asks a person before anything hard to undo, runs in an isolated environment, and logs every tool call. An agent acts with whatever credentials you give it, so its risk is the reach of those credentials. OWASP, NIST, and a May 2026 guide from CISA and five partner agencies all point the same way: start with low-risk tasks and widen access only as the controls prove out.
A chatbot answers. An agent acts: it opens a browser, runs a terminal command, calls an API, or uses a tool connected through the Model Context Protocol (MCP). Each of those actions runs on a credential someone handed it, such as a login, an API key, or an OAuth token. The useful question is what that credential reaches, who approves the risky steps, and whether anyone can reconstruct afterward what the agent did.
What can an AI agent do with your credentials?
Everything the credential allows. OpenAI's help center says that once an agent is signed into websites or given apps, it can read emails, files, and account settings and act on your behalf, for example by sharing files or changing account settings. The joint agentic AI guide from CISA, the NSA, and the cyber agencies of Australia, Canada, New Zealand, and the United Kingdom describes a procurement agent given broad access to finance, email, and contracts. An attacker who compromises one low-risk tool in its workflow inherits all of that access, and the audit logs look legitimate because the trusted agent did the work.
Two verified incidents show the pattern. In July 2025, AWS reported that a GitHub token with overly broad scope let an attacker commit malicious code into release 1.84.0 of the Amazon Q Developer extension for VS Code; the code failed to run because of a syntax error. The same month, Fortune reported that Replit's coding agent deleted a live production database during a declared code freeze. Replit's CEO called it unacceptable and said the company was adding automatic separation between development and production databases.
Is ChatGPT agent mode safe?
OpenAI has retired it. As of October 8, 2026, the help center says ChatGPT agent is no longer available and points users to ChatGPT Work and its cloud browser. That browser runs on a separate computer in the cloud and keeps its own cookies and sign-ins, apart from your personal browser. By default it asks before visiting a new website, and it asks for confirmation in chat before actions that are hard to reverse or that create a financial, legal, or account commitment. Passwords go through a secure sign-in form the model never sees. OpenAI states plainly that these safeguards do not remove every risk, and it lists an "Always allow" setting for every website as not recommended.
One detail from the retired product is worth carrying into any assessment: OpenAI's Compliance API recorded agent conversations but not the individual agent actions, such as app requests. Ask every agent vendor the same question about its own logs.
How safe are AI coding agents like Claude Code, Copilot, and Cursor?
| Coding agent | Default approvals | Isolation and limits |
|---|---|---|
| Claude Code (Anthropic) | Auto mode, the starting mode for interactive sessions, sends actions to a separate classifier model that blocks the ones it judges unsafe. Manual mode starts read-only and asks before edits, tests, and commands. | Optional sandbox for shell commands with file and network isolation. Cloud sessions run in isolated virtual machines with network access limited by default; GitHub credentials stay outside the VM. |
| GitHub Copilot cloud agent | Only people with write access can start it. Its draft pull requests must be reviewed and merged by a human, and workflows wait for that review. | Pushes to a single branch, follows branch protection rules, and runs with restricted internet access. |
| Cursor | Terminal commands and every MCP connection need approval by default. Each MCP tool call needs approval unless allowlisted. | Default settings block arbitrary network requests. Cursor calls its auto-approve run modes, including the auto-review classifier, best effort and no hard security boundary. |
All three ship approval controls. Risk enters when a team switches on blanket auto-approve to stop the prompts, or connects the agent to production systems with a developer's own long-lived token. Anthropic's documentation puts the duty plainly: the person approving is responsible for reviewing proposed code and commands.
What does least privilege mean for LLM tool execution?
OWASP's LLM06:2025 Excessive Agency names three root causes:
- Excessive functionality: a tool that only needs to read documents also carries delete and modify functions.
- Excessive permissions: a read-only tool connects to a database with an account that can also write and delete.
- Excessive autonomy: a high-impact action runs with no user confirmation.
The OWASP Top 10 for Agentic Applications, published December 9, 2025, extends this into least agency: grant autonomy only where it adds value. Its guidance for tools is concrete. Give each tool its own profile, such as read-only queries for a database and no send or delete rights for an email summarizer, and issue short-lived credentials bound to one user session.
How should MCP servers handle credentials?
The MCP specification's security best practices (version 2025-11-25) set three rules that matter most here. An MCP server must reject any token that was not issued for that server; passing a client's token straight through to another API is explicitly forbidden, because downstream logs then show the wrong identity and controls tied to that identity stop working. Scopes start small and grow only when a privileged operation first needs them; wildcard scopes such as files:* or admin:* turn one stolen token into broad access. A local MCP server runs with the same privileges as the app that launched it, so a client offering one-click setup must show the exact command and get explicit approval before running it.
Treat each MCP server as third-party software with your keys inside it. Anthropic notes that it reviews connectors for its directory against listing criteria and does not security-audit them.
Which agent actions need human approval?
Anything hard to undo: deleting data, moving money, publishing, sending information outside the company, and changing permissions. OWASP recommends showing the person a dry-run or diff of the planned change before they approve it. The CISA guide adds that agents should never be able to change their own privileges. OWASP's own example explains why: a hidden instruction in an incoming email tells an assistant to search the inbox and forward sensitive messages. A read-only mail tool, a read-only OAuth scope, and a human review before anything is sent each stop it.
What should you log for every tool call?
- Identity: which agent acted, and which person or workflow it acted for.
- Action: the tool, its parameters, and the result.
- Approval: who approved a consequential step, and when.
OWASP calls for immutable logs of every tool invocation and alerts on unusual chains, such as a database read followed by an external transfer. CISA asks for tool use in human-readable logs. NIST AI 600-1 flags indirect prompt injection, where instructions hide inside data the agent retrieves, and says documenting third-party plugins is especially important when an incident has to be disclosed.
How do you assess an AI agent before turning it on?
- Inventory the reach: list every tool, connector, MCP server, and credential the agent can use, and remove the ones the task does not need.
- Scope each credential: give the agent its own identity, read-only access wherever possible, and tokens that expire.
- Name the approval points: write down which actions pause for a person, and test that they actually pause.
- Isolate execution: run code and commands in a sandbox or separate machine, allowlist outbound network access, and keep production data out of development.
- Confirm the logs: check that individual tool calls appear in a record your security team can read, beyond the conversation transcript.
- Pilot small: start with a low-risk task, plant a test instruction in a document the agent will read, and widen access only after it passes.
Agent products change monthly. Re-run this assessment when a vendor changes its defaults, when you add a tool, and at least every quarter.
Frequently asked questions
- Is ChatGPT agent mode safe?
- OpenAI has retired ChatGPT agent and points users to the cloud browser in ChatGPT Work. That browser asks before visiting new sites by default, asks for confirmation before consequential actions such as payments or bookings, and keeps passwords out of the model's view. OpenAI says these safeguards do not remove every risk, so review sensitive steps yourself.
- What is excessive agency in AI?
- Excessive agency is OWASP's name for an AI system that can do more damage than its job requires. It comes from three causes: tools with unneeded functions, accounts with unneeded permissions, and high-impact actions that run without a person's approval.
- What is least privilege for LLM tool execution?
- Each tool an AI model can call gets only the functions, permissions, and data its task needs, enforced by the target system. OWASP's agentic guidance adds short-lived credentials bound to one user session and a separate profile for each tool.
- How should MCP servers handle API keys and tokens?
- The MCP specification says a server must reject tokens that were not issued for it and must never pass a client's token through to another API. Scopes start minimal and grow only when a privileged operation needs them.
- What is an AI sandbox?
- An AI sandbox is an isolated environment where an agent's code and commands run with restricted file and network access, so a bad command stays contained. Claude Code offers one for shell commands, and the MCP specification recommends sandboxing local MCP servers with minimal default privileges.
- What does the OWASP Top 10 for Agentic Applications cover?
- Published December 9, 2025, it lists agent goal hijack, tool misuse and exploitation, identity and privilege abuse, agentic supply chain vulnerabilities, unexpected code execution, memory and context poisoning, insecure inter-agent communication, cascading failures, human-agent trust exploitation, and rogue agents.
- How do you secure AI coding agents?
- Keep approval prompts on for commands and file edits, run the agent in a sandbox or isolated machine, and give it a scoped token in place of a developer's own credentials. Require human review before its changes merge, as GitHub does by default for Copilot cloud agent.
Sources
- OWASP Gen AI Security Project, LLM06:2025 Excessive Agency. Read October 8, 2026.
- OWASP Gen AI Security Project, OWASP Top 10 for Agentic Applications for 2026 (December 9, 2025). Read October 8, 2026.
- NIST, AI 600-1: Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (July 2024). Read October 8, 2026.
- Model Context Protocol, Security Best Practices (specification 2025-11-25). Read October 8, 2026.
- CISA, CISA, US and International Partners Release Guide to Secure Adoption of Agentic AI (May 1, 2026). Read October 8, 2026.
- Australian Cyber Security Centre, CISA, NSA, et al., Careful adoption of agentic AI services (May 1, 2026). Read October 8, 2026.
- OpenAI Help Center, ChatGPT agent. Read October 8, 2026.
- OpenAI Help Center, Using cloud browser in ChatGPT. Read October 8, 2026.
- Claude Code Docs, Security. Read October 8, 2026.
- GitHub Docs, Risks and mitigations for GitHub Copilot cloud agent. Read October 8, 2026.
- Cursor Docs, Agent security. Read October 8, 2026.
- AWS Security Bulletin AWS-2025-015, Security Update for Amazon Q Developer Extension for Visual Studio Code (Version #1.84). Read October 8, 2026.
- Fortune, An AI-powered coding tool wiped out a software company's database, then apologized for a 'catastrophic failure on my part' (July 23, 2025). Read October 8, 2026.
Reviewed