Speaking twice at Dreamforce · Sept 15-17 →

AI governance

US · UAE · India

AI agent sandboxes stop reach, not the act: Strands Box and 14 sandboxes compared

By Shivanath DevinarayananPublished 19 min read
Written forCIO, CISO, or compliance lead who approves agents on real systems“Who acted, and can we stop it?”

The short answer

What does an AI agent sandbox stop, and what does it leave to policy?

A sandbox stops reach: files, programs and hosts outside the box. It does not stop a permitted action repeated, done in the wrong order, or done after the agent read customer data. AWS's Strands Box adds rules on what the agent already did, and keeps secrets outside it. At Mindcat, we would test it on one read-only agent first.

A robot agent in a glass box: three messages leave through a gate, a red bar holds the fourth. Post 4: denied

Takeaways

  1. 01A sandbox limits where an agent reaches. It does not limit how often it acts, in what order, or after what.
  2. 02Strands Box adds Dogwood rules that read the agent's history, such as three Slack posts per 10 minutes.
  3. 03Keep the secret in a proxy outside the agent. Box, Docker, Cloudflare, E2B and Daytona document this.
  4. 04Read where each boundary stops. Claude Code's covers shell commands only, and Codex's proxy skips MCP connections.
  5. 05Box is a macOS developer preview. For multi-tenant cloud use, AWS pairs it with a microVM per session.

What did AWS ship on 7 October?

A sandbox decides where your agent can reach. It does not decide what the agent may do once it is there, how often, or in what order. On 7 October AWS released Strands Box in developer preview to close part of that gap: an open source sandbox, under Apache 2.0, that pairs operating-system isolation with rules that can depend on what the agent has already done.

Picture the security review for an on-call agent. The engineer says it runs in a sandbox. Ask two questions: what stops it posting to the incident channel forty times, and does it ever hold the Slack token? Isolation alone answers neither. Box answers the first with a rule on history and the second with a credential proxy.

Disclosure: Mindcat builds client agents on Claude and on OpenAI models. We are not an AWS, Anthropic or Google partner, and no vendor in this post pays us. We read the AWS post, Marc Brooker's companion post, the Box repository and its docs, and every other vendor's own pages on 9 October 2026.

What it doesHowFor the approver
Contains the agentOS isolation (macOS Seatbelt) on the developer's own machine, with no separate guest OSHard boundary, macOS only
Checks each actionDogwood rules at four points: network egress, a shell, a Python interpreter and a broker for MCP serversDenied unless permitted
RemembersOne event history shared by all four points, so a rule can count, order and time-box actionsLimits on repeats and order
Keeps the secret outThe egress gateway swaps a placeholder for the real credential on each permitted requestAgent never holds the token
Leaves a recordEvery policy decision written as OTLP JSON, with the rule that decided itRule ID on every decision

Four things the announcement does not say, from the Box docs and Brooker's post:

  • The preview runs on macOS with Apple silicon, macOS 15 or later, at version 0.1.x. The guide says keys, flags and policy actions can change between releases.
  • "Every process in the box can use every route." The agent, each tool and each MCP server in one box get the same placeholders, so the box, not the process, is the credential boundary.
  • For multi-tenant deployments in the cloud, AWS recommends running Box inside a dedicated microVM per session.
  • A limit counted on successful responses can be overspent by requests still in flight. A limit counted on requests is exact, and refused attempts count toward it.

So, today, Box is a control for agents on a developer's Mac, not a server control you can show an auditor. Its value now is learning what your agents actually do and writing the rules before the Linux and cloud versions arrive. We have not run Box on client work. The tests at the end are the ones we would run first.

What does a sandbox stop, and what does it not?

A sandbox stops reach: files outside the paths you grant, programs you did not allow, hosts you did not list. It does not judge an action it allows. If your agent may post to Slack, a sandbox lets it post forty times. If it may read the customer folder and reach the internet, a sandbox lets it do both, in that order.

AWS says the same in its announcement: "Containers and microVMs provide strong isolation, but isolation alone doesn't enforce contextual rules." Its example is an agent investigating a production incident, which should read logs and inspect infrastructure without being allowed to change it.

Box splits the job into two tiers. box.toml sets what the agent's own processes can open, and the operating system enforces it before the first process starts. policy.dw decides each request that goes through Box's own shell, Python, egress gateway and MCP broker, which run in Box's process outside the sandbox. Anything those four handle is refused unless a permit matches, a forbid beats a permit, and if the engine faults, the request is denied.

AWS diagram: box.toml and policy.dw feed a box with a trusted boundary of six parts, and the agent in containment
Strands Box: the trusted parts check each action, and the OS holds the agent in containment. Image: AWS, from the Strands Box announcement.

That split is where the write gets caught. If the agent runs rm -rf build/, Box's shell raises an fs:delete decision for each file it would remove, so a forbid on delete stops the command before anything is deleted, while a normal edit in the workspace still goes through. The policy sees the file and the operation, not an opaque system call. It is our line, "the write is the risk", in a form you can test.

For an approver, the record matters as much as the refusal. Box writes each decision as OTLP JSON, by default to private/telemetry/records.jsonl in the box's own directory, with the verdict, the action, the path or host, and the rule. No request from inside the box can reach that directory, so the agent cannot read or edit its own log. Two gaps to plan for: the docs say the log "records the rule and the path, and not the operation" (a cat and an ls of one file look the same), and the docs we read say nothing about retention or cost per call. OTLP is the OpenTelemetry format, so ship the file to your own log store and set retention there.

When the agent is refused, it gets the same text: "policy denied this operation", the rule's ID and its description, so a clear description tells the agent to wait instead of retry. In Box the principal is always the agent, so the log names the rule, not a person. "Who acted" still needs the identity on the other side of the call, which we come back to under Salesforce.

Two columns: the sandbox stops files, programs and hosts; policy stops a fourth post, an upload after a read, an early push
What a sandbox stops, and what it leaves to policy.
If you are deciding what has to be true before an agent touches a system of record, our AI governance work names every connector, MCP server and approver per agent, and puts the policy on the hop.

Why do rules about history matter?

Because the risky act is often a permitted act at the wrong time. Dogwood rules can read the agent's recorded history: what it did, in what order, and how long ago. In AWS's words, "permissions can depend on earlier actions, their order and limits accumulated over time."

AWS's own example is the incident agent. It may post progress to the incident channel in Slack, but no more than three times every 10 minutes, so it does not bury the updates from people. The rule counts posts that returned HTTP 200 in the last 10 minutes. The agent keeps reading logs while the cap holds.

AWS's incident-agent timeline

  1. 10:00

    Post an update: allowed

  2. 10:03

    Post an update: allowed

  3. 10:05

    Post an update: allowed

  4. 10:06

    A fourth post within 10 minutes: denied, HTTP 403

  5. 10:07

    Retrieve more logs: allowed

  6. 10:11

    Post an update: allowed

From the Strands Box announcement. 10:11 is allowed because the 10:00 post has left the 10-minute window.

The same shape covers rules an approver would want, all from AWS's pages:

  • After the agent reads a file from the customer-data directory, block further outbound HTTP requests. Because the shell and Python report reads the same way, the rule holds whichever tool did the read.
  • Allow git push only if npm test passed since the last git add, within 15 minutes. Brooker notes this is not a complete policy, since a git switch could break the promise.
  • Cap calls to one MCP server at 60 an hour, for paid APIs or a shared quota.

If you own the bill, read that last rule as a spend control. Brooker writes that caps like it help with paid APIs, shared quotas, or "simply to prevent unexpected bills", and gives a payments API capped at $100 a day as a rule Box can express. AWS also argues the interpreters cut token spend, because intermediate results stay in the sandbox instead of passing through the model. Neither page gives a latency or overhead figure. The docs warn only that a long time window slows every decision on that action, so time your own.

Timeline: posts at 10:00, 10:03 and 10:05 allowed, the fourth at 10:06 denied, logs at 10:07 and a post at 10:11 allowed
AWS's incident-agent example: three Slack posts in any 10 minutes.

Your approvals are sequences too. Our own approval flows at Mindcat (expenses, vendor payments, hiring requests) keep the request in a Slack List while it moves through the manager, the department head and, for budget items, finance and the CEO. The flow writes back to Salesforce only after the last approval. That is an ordering rule, and today the flow's design enforces it. In Box's terms it would be a forbid on the Salesforce write unless an approval appears in the history since the request. We have not built or tested that.

What we learned going live applies here. The gaps after launch were missing process cases (contractors, external approvals, audits), not model mistakes. A rule about history needs the same write-up first: you cannot count an event you never named. And the approval has to name the exact action, as the $20 approval that ran as $2,000 showed.

Should policy run inside the agent or at the boundary?

At the boundary, with the agent app's own prompts kept as a second layer. AWS's reason: the permission prompts built into agent apps "run inside the agent's own process and they judge the tool call, not its effect." A prompt sees "run this shell command". It does not see the files the command will touch or the hosts it will reach. Box enforces from outside the process, and one box.toml and policy.dw can carry across agent apps. Its docs include tutorials for Claude Code, Codex CLI and pi.

Each vendor's own page says where its boundary stops. Read that line before you approve anything:

  • Claude Code. "The sandbox covers shell commands only. Claude's file tools, MCP servers, and hooks run outside it." Anthropic's sandbox-runtime wraps the whole Claude Code process when you need more.
  • Codex. Its network proxy filters commands in the local sandbox and "does not filter web search, app or connector tool calls, MCP server connections".
  • Strands Box. Paths granted in box.toml raise no policy decision, and a program that a shell:spawn permit admits runs without its own file operations being decided.

Brooker explains why AWS put policy above the kernel. At the level of a virtual network card, "pretty much all you see on the wire is TLS ciphertexts." So Box uses the kernel sandbox to make the policy layer "the only way out of the box", and the policy layer reads HTTP methods, MCP calls and shell commands. The cost, in his words, is more code to trust: interpreters and protocol interceptors where "bugs could mean bypassed policy." AWS says it wrote that code in Rust and invested in "testing, fuzzing, and validation". The repository has a SECURITY.md for reporting a suspected problem. We found no third-party review of Box's security model.

Three cards: rules in the agent see the call; at the boundary, files and hosts; a microVM per session splits tenants
Where the rule runs decides what it can see.

So the answer is layered, and each layer has a trigger. Many tenants, or code from untrusted input: a microVM per session. An agent on a developer's machine: an OS sandbox with a policy point at each exit. Any write that moves money or changes a record: a person. Rent the isolation, since every vendor in the table sells it. Own the rules and the decision log, because they encode your approvals and have to move with you when you change vendors. On client agents we put that check on the hop where the call leaves, the pattern behind our Flex Gateway work. One caution from Box's own guide: it switches the Strands CLI's approval prompts off so the box decides each command. If you copy that, the box is your only check, so keep the person on the writes.

How do you keep the secret away from the agent?

Put the credential in a proxy outside the sandbox and give the agent a placeholder. Box's egress gateway swaps the placeholder for the real secret on each request that policy permits. It supports a bearer token, a custom header such as x-api-key, HTTP Basic, a query parameter, and AWS SigV4, which signs with credentials read outside the agent. A credential binding grants no network reach on its own: each destination still needs a permit.

A robot works inside a walled courtyard at night while the keys hang in a gatehouse outside the wall, out of its reach
The agent works inside the wall. The keys stay in the gatehouse.

Five other vendors document the same pattern. Docker's host-side proxy injects headers, and "credential values never enter the VM". Cloudflare's outbound handlers run in your Worker and add the token after the request leaves the sandbox. E2B's egress proxy fills stored secrets into HTTPS headers. Daytona swaps a placeholder in an outbound proxy. Claude Code's mask setting does the same, as an experimental option that needs the proxy to terminate TLS.

The limit to write into your approval is Box's own: "Every process in the box can use every route." The gateway cannot tell which process sent a request. To keep a credential away from a tool or an MCP server, Box's docs say to run it in a separate box. And because the agent never held the secret, revoking it does not depend on the agent.

Five steps: agent gets a placeholder, policy checks the request, the gateway swaps in the secret on the way out
Credential injection in Strands Box, in five steps.

How do the other sandboxes compare?

No offering wins every row. The first table answers "what holds the agent, and can a rule read its history?" The second answers "who holds the secret, where does it run, and can we use it today?" Where a vendor's page does not say, the cell says so. Everything was read on the vendors' own pages on 9 October 2026.

Isolation and policy, read 9 October 2026

CriterionIsolationWhere policy runsRules on history
AWS Strands BoxOS sandbox (macOS Seatbelt) on your machine, no guest OSOutside the agent: shell, Python, egress gateway, MCP broker; deny by defaultYes: counts, order and time windows in Dogwood
AWS AgentCore Runtime and PolicyMicroVM per session (Firecracker)AgentCore Policy on tool calls, in CedarNot stated in the post we read
Anthropic Claude Code sandbox, sandbox-runtimeSeatbelt on macOS, bubblewrap on Linux and WSL2Proxy outside the sandbox checks each host; built-in sandbox covers shell commands onlyNot stated
OpenAI Codex CLISeatbelt (macOS), Landlock and seccomp (Linux), restricted token (Windows)Sandbox mode plus approval policy; optional domain proxy that skips MCP connectionsNot stated
OpenAI Code InterpreterA container OpenAI calls a fully sandboxed virtual machineHosted by OpenAI; the page does not describe network policyNot stated
Google Gemini CLISeatbelt profiles (macOS), Docker or Podman, gVisor on LinuxProfiles; the default confines writes and allows networkNot stated
Google GKE Agent SandboxgVisor pods on GKE; Kata Containers also supportedAdmission policies; default-deny network, egress rules per templateNot stated
Google Agent Engine Code ExecutionManaged sandbox; technology not named on the pageHosted; limited file system and no network accessNot stated
Microsoft Azure Container Apps dynamic sessionsHyper-V boundary per sessionPool setting; egress disabled by defaultNot stated
Cloudflare Sandbox SDKContainers; technology not named on the pages we readYour Worker's outbound handlers, outside the sandbox; allowed and denied hostsNot built in; handlers are your code
E2BFirecracker microVM per sandboxEgress allow and deny lists; domains by SNI or HostNot stated
DaytonaOwn kernel per sandbox, per its READMEPer-sandbox firewall: block all, CIDR or domain allow listNot stated
Modal SandboxesgVisor, or a VM with its own kernelBlock all, CIDR allow list, domain allow list (beta)Not stated
Docker SandboxesMicroVM per sandbox with its own Docker EngineHost-side proxy, deny by default; MCP gateway policiesNot stated
Sources per row in the next table. 'Not stated' means the vendor pages we read do not say.

Secrets, platform and status, read 9 October 2026

CriterionSecret kept outsideRuns onLicence, statusSource
AWS Strands BoxYes: placeholder swapped at the gatewaymacOS on Apple silicon; Linux plannedApache 2.0; developer preview, 0.1.xAWS Open Source blog, 7 Oct; Box README and docs
AWS AgentCore Runtime and PolicyNot covered in the post we readAWS cloudManaged serviceBrooker, strandsagents.com, 7 Oct
Anthropic Claude Code sandbox, sandbox-runtimeYes, experimental: mask values swapped by the proxymacOS, Linux, WSL2; native Windows runs unsandboxedRuntime Apache 2.0, beta research previewcode.claude.com sandboxing; GitHub sandbox-runtime
OpenAI Codex CLINot stated on the pages we readmacOS, Linux, WindowsApache 2.0developers.openai.com Codex security, permissions
OpenAI Code InterpreterNot statedOpenAI cloudManaged API tooldevelopers.openai.com Code Interpreter guide
Google Gemini CLINot statedmacOS; any OS with Docker or PodmanApache 2.0; off until you turn it onGemini CLI sandbox docs
Google GKE Agent SandboxNot stated; service account token must not mountGKE on Google CloudController Apache 2.0; API v1alpha1docs.cloud.google.com GKE Agent Sandbox
Google Agent Engine Code ExecutionNot statedGoogle Cloud; Python and JavaScriptManaged; status not stated on the pagedocs.cloud.google.com Code Execution
Microsoft Azure Container Apps dynamic sessionsManaged identity, off inside sessions by defaultAzureManaged servicelearn.microsoft.com dynamic sessions
Cloudflare Sandbox SDKYes: your Worker adds the tokenCloudflareSDK Apache 2.0developers.cloudflare.com Sandbox SDK
E2BYes: egress proxy fills stored secrets into HTTPS headersE2B cloud, or self-hosted on AWS or Google CloudApache 2.0docs.e2b.dev; GitHub e2b-dev
DaytonaYes: placeholder swapped by an outbound proxyDaytona cloudCore went private in June 2026; old public repo AGPL-3.0, unmaintaineddaytona.io docs; GitHub daytonaio
Modal SandboxesNot stated on the networking pageModal cloudManaged servicemodal.com sandbox networking
Docker SandboxesYes: host-side proxy injects headersLocal microVMs or Docker's cloudOrg-wide policy on a paid plan; licence not stateddocs.docker.com Sandboxes
Licence and status as the vendor states them. A managed service has no licence to show.

Read the tables as a split by where the agent runs. On a developer's machine, Box, Claude Code, Codex, Gemini CLI and Docker Sandboxes compete, and only Box ships rules on history. In the cloud, E2B, Modal, Daytona, Cloudflare, Azure and Google's two services give you a separate machine per session and an egress list, and none of the pages we read offers a rule on what the code did earlier. Defaults differ more than you would guess: E2B and Modal sandboxes reach the internet unless you restrict them, while Azure sessions and Google's Code Execution start with no network.

For a multi-tenant agent in the cloud, our rule from these tables: pick a microVM or Hyper-V boundary per session, an egress allow list that starts empty, and a proxy that holds the secret. That is E2B, Docker's cloud sandboxes, AgentCore or Azure dynamic sessions, depending on your cloud. None of their pages we read offers a rule on history, so put that rule at your gateway, or run Box inside the session as AWS suggests.

What does this mean for Salesforce, Slack and MCP agents?

You need two logs to answer "who acted". Salesforce's hosted MCP servers run every call as the signed-in user, and the audit trail names that person, not the agent, as our Salesforce MCP server post set out. Box's decision log names the rule that allowed or refused the call. Salesforce tells you on whose authority the agent acted, and Box tells you under which rule.

Box can sit in front of both kinds of MCP server. Its MCP broker checks local servers' tool calls and their arguments, and its egress gateway raises an mcp:call decision for each MCP request over HTTP. So a rule can sit on a remote MCP call as well as a local one.

The credential needs a test before you plan on it. Box's credential sources are an environment variable or an AWS profile. The docs we read do not describe an OAuth refresh flow, which a Salesforce External Client App uses when the client should stay signed in. Check whether a short-lived token passed as a bearer covers your case.

Slack is AWS's own example. The bot token is injected on every permitted request to slack.com, and a history rule caps the posts. That is the pattern for any agent that posts approvals or alerts into a channel your people read.

Can you stop it? Box gives you three levers. A forbid refuses the next action, but the box reads policy.dw when it starts, so a new rule applies from the next run. The secret lives outside the agent, so you can revoke it at its source without the agent's help. And when the agent exits, the box stops. For a fleet of agents you still need the list of every agent, the named owner and the timed drill from our kill switch post, which also covers what DIFC Regulation 10 adds. We do not map Box's controls to a regulation clause here: no vendor page we read does, and we will not guess. For MCP controls in general, our MCP governance guide has the checklist.

Where does Strands Box stop?

AWS states most of these limits itself:

  • macOS only, for now. The preview needs Apple silicon and macOS 15 or later. Linux is planned. Your servers and CI runners are probably Linux, so today Box is a developer-machine control.
  • Early. Version 0.1.x, and keys, flags and actions can change between releases.
  • Wider trusted computing base. The interpreters and gateways that enforce policy run outside the sandbox. AWS calls this "a deliberate trade-off".
  • Direct grants are invisible to policy. Paths in box.toml are bounded by containment only and "do not appear in the policy history". A tool's own file operations are not decided either.
  • Shared routes. Every process in a box can use every credential route.
  • Rules are not checked against each other. Two permits that admit the same requests load without a word, so the docs say to write every cap as a forbid.
  • Counting has edges. A limit on successful responses can be overspent by requests in flight. A long time window slows every decision on that action.
  • Hard to read. Brooker expects your first reaction to be that the policies are "difficult to understand, and difficult to write". AWS ships a policy-authoring skill for coding agents. A policy a model wrote is a reason to review it, not to skip review.

What should your team test this month?

Start with one read-only agent on one Mac. Shivanath has put 200+ agents into production, and the method we use is Spec, shadow, gate: write the step down, run it next to today's version, then put a person on the write. Box fits the shadow stage.

  1. 01Run one read-only agent (log reading or a weekly report) in Box on one Mac. Keep the agent app's own prompts on.
  2. 02Add a forbid on fs:delete, ask the agent to clean a build folder, and read the denial in the decision log.
  3. 03Copy AWS's Slack cap to a test channel. Post four times in 10 minutes and confirm the fourth is refused and the agent waits.
  4. 04Write the customer-data rule: after a read in that folder, no outbound HTTP. Then ask the agent to upload a summary.
  5. 05List every path in agent.filesystem and every tool table in box.toml. Each one is reach that policy never sees.
  6. 06For the sandbox you run today, ask the vendor or your platform team the five questions in the tables: isolation, where policy runs, history rules, who holds the secret, which OS.
Checklist: a read-only agent in a box, a delete rule, a Slack cap, a read then upload, direct paths, revoking the secret
What we would test this month.

If your team already runs agents on Claude or on OpenAI models, nothing here asks you to switch. Ask the same four questions of any stack you run: what the boundary covers, where the rule runs, whether it can count, and whether the agent ever holds the key. Our agent stack notes and the Claude Cowork rollout post cover the approval and data side of the same stack.

Filed under

AI governanceagent sandboxingstrands boxdogwood policyclaude code sandboxcodex sandboxai agent security
Useful to someone on your team?

FAQ

Questions teams ask next

Is Strands Box ready for production?
Not as a server control yet. It is a developer preview at version 0.1.x, for macOS on Apple silicon, and its guide says keys, flags and policy actions can change between releases. For multi-tenant cloud use, AWS itself recommends running Box inside a dedicated microVM per session. Use it now to learn what your agents actually do and to write the rules you will need.
Does Strands Box work with Claude Code or Codex?
Its docs include tutorials that put Claude Code, Codex CLI and pi in a box, with the project kept out of the agent's direct reach so its shell tool is the only way to a project file. AWS's Marc Brooker says Box works with any model and is not tied to the Strands SDK.
What is Dogwood?
Dogwood is AWS's open source policy language for agents. It uses Cedar's syntax for permit and forbid rules and adds temporal expressions that read the agent's event history, such as a count within a time window or whether one event happened since another. The Dogwood Local Engine evaluates the rules on your machine, and Box embeds it.
Can a sandbox stop prompt injection?
It limits the damage, not the injection. The sandbox and a default-deny policy decide what an injected instruction can reach and do: which hosts, which files, how many posts. Anthropic notes that a proxy which checks only hostnames does not inspect encrypted traffic, so allow few domains, and keep credentials outside the agent so an injected step cannot send them anywhere.

Sources

  1. 01AWS Open Source Blog: Introducing Strands Box, AI agent sandboxes powered by Dogwood (7 October 2026)Accessed 09 Oct 2026
  2. 02Marc Brooker: Strands Box, the big picture (7 October 2026)Accessed 09 Oct 2026
  3. 03GitHub: strands-agents/box READMEAccessed 09 Oct 2026
  4. 04Strands Box docs: Security, what Box governs and what you ownAccessed 09 Oct 2026
  5. 05Strands Box docs: PolicyAccessed 09 Oct 2026
  6. 06Strands Box docs: The egress tableAccessed 09 Oct 2026
  7. 07Strands Agents: Get started with Strands BoxAccessed 09 Oct 2026
  8. 08The Dogwood GuideAccessed 09 Oct 2026
  9. 09Claude Code Docs: Configure the sandboxed Bash toolAccessed 09 Oct 2026
  10. 10Claude Code Docs: Choose a sandbox environmentAccessed 09 Oct 2026
  11. 11Claude Code Docs: Securely deploying AI agentsAccessed 09 Oct 2026
  12. 12GitHub: anthropics/sandbox-runtimeAccessed 09 Oct 2026
  13. 13OpenAI Developers: Codex agent approvals and securityAccessed 09 Oct 2026
  14. 14OpenAI Developers: Codex permissionsAccessed 09 Oct 2026
  15. 15OpenAI Developers: Code InterpreterAccessed 09 Oct 2026
  16. 16Gemini CLI docs: SandboxingAccessed 09 Oct 2026
  17. 17Google Cloud: About GKE Agent SandboxAccessed 09 Oct 2026
  18. 18Google Cloud: Enable Agent Sandbox on GKEAccessed 09 Oct 2026
  19. 19Google Cloud: Code Execution overviewAccessed 09 Oct 2026
  20. 20Microsoft Learn: Dynamic sessions in Azure Container AppsAccessed 09 Oct 2026
  21. 21Microsoft Learn: Use session pools in Azure Container AppsAccessed 09 Oct 2026
  22. 22Microsoft Learn: Use dynamic sessions in Azure Container AppsAccessed 09 Oct 2026
  23. 23Cloudflare Sandbox SDK docs: Handle outbound trafficAccessed 09 Oct 2026
  24. 24Cloudflare Sandboxes docs: Sandbox securityAccessed 09 Oct 2026
  25. 25E2B Docs: Internet accessAccessed 09 Oct 2026
  26. 26GitHub: e2b-dev/infraAccessed 09 Oct 2026
  27. 27Daytona docs: IsolationAccessed 09 Oct 2026
  28. 28Daytona docs: Network limitsAccessed 09 Oct 2026
  29. 29GitHub: daytonaio/daytona READMEAccessed 09 Oct 2026
  30. 30Modal Docs: Sandbox networking and securityAccessed 09 Oct 2026
  31. 31Docker Docs: Docker SandboxesAccessed 09 Oct 2026
  32. 32Docker Docs: Docker Sandboxes security modelAccessed 09 Oct 2026
Shivanath Devinarayanan, founder of Mindcat

Written by

Shivanath Devinarayanan

Founder, Mindcat Consulting · Salesforce MVP Hall of Fame

Runs 200+ agents in production. Reads every brief that comes in and signs the work that goes out.

About Shivanath

First Agent in Production · the written assessment

We work in real estate and right now, the AI pilot never left the sandbox.

Industry
Problem

Brief: Real estate. Leads from portals and WhatsApp. Follow-up has no owner. Stalled AI or Salesforce pilot. Need a finish, cut, or rebuild call. From: /blog/ai-agent-sandbox-policy.