Guides

Best cloud coding agents in 2026

Products that run coding agents in the cloud or side by side, grouped by how they work, with prices as of September 2026.

The short answer

Cloud coding agents come in four kinds. Anthropic's and OpenAI's own clouds run Claude Code and Codex as part of your Claude or ChatGPT plan. They suit tasks that need only your repository. Cursor, Devin, Amp, GitHub Copilot, Factory, and Vorflux run their own agent harnesses. Cloud agent platforms such as Boxes.dev, Replicas, Tembo, and Niteshift run Claude Code, Codex, and other agents on machines they prepare. Boxes.dev starts each task from a copy of a machine where your app already runs. Workbenches such as Conductor and Superset run many agents side by side, mostly on your own computers.

How we put this together

This guide covers products that run coding agents on cloud machines, whether their own agent or agents such as Claude Code and Codex. It also covers parallel agent workbenches, which run many agents at once on computers you choose. We read each product's public documentation and pricing pages between September 27 and October 2, 2026. We didn't run benchmarks or rank how well each agent writes code; the entries describe what each product does and what it costs.

Prices and plans are as of those dates. Most entries link to a longer comparison with boxes.dev, and each comparison lists its sources.

The products at a glance

ProductKindWhere the agent runsAgentsStarting price
Claude Code cloud sessionsModel labA fresh Anthropic VM for each sessionClaude CodeIncluded in Claude Pro, $20 a month
Codex cloudModel labAn OpenAI VM for each task, from an environment you publishCodexIncluded in ChatGPT Plus, $20 a month
CursorOwn agentAn isolated VM with a desktop for each cloud agentCursor's agent, with many modelsPro $20 a month
DevinOwn agentA Linux VM for each sessionDevinPro $20 a month
AmpOwn agentA remote Linux machine, called an orb, for each thread, or a machine you already haveAmp's agentFree, with orbs billed by the minute
GitHub CopilotOwn agentA GitHub Actions environment for each taskCopilot's agent, with Claude and Codex in previewFree; Pro $10 a month
FactoryOwn agentLong-lived cloud machines that sessions shareDroidPro $20 a month
VorfluxOwn agentA dedicated Ubuntu machine for each sessionVorflux's own harnessPay as you go, with signup credits
Boxes.devCloud agent platformFor each task, a full copy of a machine where your app already runsClaude Code and CodexFree trial; Starter $19 per user a month
ReplicasCloud agent platformA VM for each task, built from hook scriptsClaude Code, Codex, and six moreDeveloper $50 per seat a month
NiteshiftCloud agent platformA cloud environment for each task, defined in your repositoryClaude Code, Codex, and three moreFree with $10 of credits a month
TemboCloud agent platformA dedicated Linux VM for each sessionClaude Code, Codex, and five moreFree with a one-time $10 allowance
AQCloud agent platformGit worktrees on one shared machineClaude Code, Codex, and other agent CLIsFree for one person
EllipsisCloud agent platformA sandbox that lasts up to an hour, for each sessionClaude Code and CodexFree for individuals with a subscription
CloudCLICloud agent platformA persistent container that can hold several projectsClaude Code, Codex, and three moreHobby €7 a month
ConductorWorkbenchGit worktrees on your Mac, or cloud workspaces on paid plansClaude Code, Codex, Cursor, and OpenCodeFree for local work
T3 CodeWorkbenchYour laptop or a host you set upClaude Code, Codex, and four moreFree and open source
SupersetWorkbenchYour Mac or a host you addClaude Code, Codex, and 18 moreFree for one user
OrcaWorkbenchYour computer, an SSH host, or an Orca serverClaude Code, Codex, and more than 30 othersFree and open source
ContinuumWorkbenchYour Mac, machines you enroll, or short-lived cloud sandboxesClaude Code, Codex, and five moreFree

Cloud agents from the model labs

Anthropic and OpenAI each run their own agent in their own cloud. Both agents come with a plan you may already pay for, and both start from your repository.

Claude Code cloud sessions

Anthropic runs each Claude Code cloud session in a fresh Ubuntu VM with your GitHub repository, common toolchains, Docker, PostgreSQL, and Redis. A cloud environment adds environment variables, a network access level, and a setup script whose installed files are cached for about seven days. You can start sessions at claude.ai/code, in the Claude apps, with claude --cloud from your terminal, with @Claude in Slack, or with routines on a schedule or on GitHub events. You can't open a shell in the VM, and Anthropic reclaims it when the session sits idle.

  • Price: Included in Claude Pro ($20 a month), Max, and Team. Cloud sessions count toward your plan's usage limits.
  • Best for: Running Claude Code remotely, using Claude in team Slack channels, and starting work from GitHub events.
  • Read more: Boxes.dev vs Claude Code and How to run Claude Code in the cloud.

Codex cloud

OpenAI runs each Codex cloud task on its own VM, started from a cloud environment that Codex helps you prepare. Codex inspects your repositories, installs dependencies and tools, and tests the workflow before you publish the environment. Each task keeps its files, including uncommitted changes, so you can reopen it from another device. You can recover a task's saved state for up to seven days after its last turn. You can start tasks from Codex on the web, the ChatGPT desktop app, or the ChatGPT app on your phone. In Enterprise workspaces, you can also mention @ChatGPT in Slack or Microsoft Teams. Codex's code reviews and its Linear and GitHub integrations still use OpenAI's earlier container environments.

  • Price: Codex comes with every ChatGPT plan. OpenAI lists cloud VM sizes for Plus ($20 a month), Pro, Business, and Enterprise. Cloud tasks can use more of your allowance than local messages do.
  • Best for: Running Codex remotely on tasks that need only your repositories, and reviewing pull requests automatically.
  • Read more: Boxes.dev vs Codex and How to run Codex in the cloud.

Platforms with their own agent

These products run an agent harness they built, usually with models from several providers, and bill model usage through their own plans.

Cursor

Cursor is an AI code editor whose agent can also run in the cloud. Each cloud agent gets an isolated VM with a desktop. Cursor clones your repositories, runs an install command, and saves a disk snapshot that each VM starts from. The agent can test its work in the VM's desktop and record screenshots and videos of its checks. You can start cloud agents from the editor, the web, the Cursor CLI, @cursor in Slack, Linear, and GitHub, an API, or automations.

  • Price: Pro $20, Pro+ $60, and Ultra $200 a month, and Teams from $40 per user each month. Cloud agents use the plan's model usage first, then on-demand usage at each model's API price.
  • Best for: Working in Cursor's editor, choosing among many models, and using pull request agents like Bugbot.
  • Read more: Boxes.dev vs Cursor.

Devin

Devin is Cognition's own autonomous agent. Each session runs on a Linux VM copied from your organization's snapshot, which Devin builds from a blueprint. A session sleeps after 30 minutes of inactivity. You can take over a session's terminal, VS Code editor, and browser. A coordinating Devin can start managed Devins in their own VMs and compile their results. Devin Review explains large diffs and flags bugs.

  • Price: Pro $20 and Max $200 a month, and Teams $80 a month plus $40 per developer. Usage beyond the plan is billed at API prices.
  • Best for: Handing whole tasks to one autonomous agent, coordinating many sessions, and reviewing large pull requests.
  • Read more: Boxes.dev vs Devin.

Amp

Amp runs its own agent on an orb, a remote Linux machine for each thread. Amp's modes choose which models the agent uses. Each orb starts from a project snapshot, made by cloning your repositories and running a setup script. Idle orbs pause and keep their files and services. Agents can start other agents in other orbs, and Puck starts and tracks agents for you. Amp can also run threads on a runner, a machine you already have, such as a server or your Mac. Threads on a runner share that machine.

  • Price: Hobby is free, with orbs billed by the minute. Megawatt ($20 a month) and Gigawatt ($200 a month) include orb time and Amp credits. Model usage is billed at the providers' API prices, or you can use a ChatGPT or SuperGrok subscription or your own API keys.
  • Best for: Using Amp's agent and Puck, with a desktop in every orb.
  • Read more: Boxes.dev vs Amp.

GitHub Copilot

Copilot's cloud agent works on GitHub issues and pull requests in a fresh GitHub Actions environment for each task. A copilot-setup-steps.yml workflow prepares the environment, and each session stops after 59 minutes. You can assign issues to the agent in GitHub or Linear, mention @copilot on a pull request, or start tasks from the Copilot app, the CLI, Slack, or Microsoft Teams. Claude and Codex are also available on GitHub as third-party agents, in preview.

  • Price: Free, Pro $10, Pro+ $39, and Max $100 a month, and Business $19 and Enterprise $39 per seat each month. Tasks use Copilot AI credits and GitHub Actions minutes.
  • Best for: Assigning GitHub issues to an agent, reviewing code automatically, and starting work from GitHub events.
  • Read more: Boxes.dev vs GitHub Copilot.

Factory

Factory runs Droid, its own agent, with Factory Router choosing models from several providers. Droid works on Droid Computers, long-lived cloud machines that Factory manages. Parallel sessions share a machine, each in its own Git worktree. Missions break larger work into milestones and coordinate workers.

  • Price: Pro $20, Plus $100 with managed Droid Computers, and Max $200 a month, and Teams $60 a month plus $40 per seat. Usage is subject to rolling 5-hour, 7-day, and 30-day limits.
  • Best for: Running Missions, routing work across models, and starting automations from Slack and GitHub events.
  • Read more: Boxes.dev vs Factory.

Vorflux

Vorflux runs its own harness on a dedicated Ubuntu machine for each session. It takes a ticket through plan review, implementation by parallel subagents, testing with a test report, review, and an optional merge queue. A setup agent builds the team's machine image by cloning your repositories, installing dependencies, and smoke-testing each one.

  • Price: No subscription. Individuals get $70 of signup credits and teams $200, then pay for machine time by the hour and for model usage.
  • Best for: Taking tickets from plan to merge with little input.
  • Read more: Boxes.dev vs Vorflux.

Cloud agent platforms

These products run Claude Code, Codex, and other agents you already use on cloud machines they prepare, usually with your own subscription or API key. They differ most in how they build each machine and what it keeps between tasks.

Boxes.dev

Boxes.dev runs Claude Code and Codex on cloud machines that start as copies of your working development environment. You'll start by setting up your app on a cloud machine called the Template box, with the tools, databases, services, and data it needs. A setup agent can copy the environment from your laptop or build it from a GitHub repository. Each task then runs on its own devbox: a full, isolated copy of that machine with its own databases, services, and ports.

You'd expect an engineer to run the app locally before shipping a change; a full dev environment lets agents do the same. On a devbox, the agent can start your app, test its changes in a browser, and reproduce bugs against your real services and data. It can also take screenshots and record video for you to review.

A devbox goes to sleep after about five minutes when no agent is working and nobody is using it. Sleep keeps its files and memory, so dev servers, databases, and terminal sessions pick up where they left off when it wakes. You can work on a devbox yourself, with terminal tabs, a file editor, diff review that sends your line comments to the agent, and a browser you share with the agent. The iOS and Android apps include a terminal too. You can start work from the apps, the dvb CLI, Slack, Linear issues, or scheduled automations.

  • Price: A free trial, then Starter $19 and Pro $99 per user each month. Agents use your Claude or ChatGPT plan or an API key.
  • Best for: Letting agents test their work in your running app, using both Claude Code and Codex, and following work from your phone.
  • Read more: How boxes.dev works and Plans, seats, and box-hours.

Replicas

Replicas runs Claude Code, Codex, and six other agents in a cloud VM for each task, called a workspace. It builds workspaces from hook scripts and keeps a pool of them ready. An idle workspace sleeps after an hour and keeps its desktop and processes. You can start work from Slack, Linear, GitHub, automations, or an API.

  • Price: A 14-day free trial, then Developer $50 and Team $200 per Full seat each month, or Flex seats billed by the minute.
  • Best for: Choosing among many agents, triaging Slack channels, and starting automations from GitHub and Sentry events.
  • Read more: Boxes.dev vs Replicas.

Niteshift

Niteshift runs Claude Code, Codex, Cursor, OpenCode, and Pi in a cloud environment for each task. A setup agent writes the environment's setup, resume, and service files and commits them to your repository under .niteshift/. Niteshift's browser testing adds screenshots and demo videos to each pull request.

  • Price: Free with $10 of credits a month, then Individual at $50 and Team at $250 a month in credits. Active agent time costs $6 an hour.
  • Best for: Keeping environment config in the repository where your team can review it, and adding screenshots and demo videos to pull requests automatically.
  • Read more: Boxes.dev vs Niteshift.

Tembo

Tembo runs seven agents, including Claude Code and Codex, on a dedicated Linux VM for each session. VMs range from 2 CPUs and 2 GB of memory up to 16 CPUs and 32 GB. Each VM starts from a standard image, and Projects prepare your repositories, dependencies, and setup script in advance. Sessions can start from Slack, Linear, GitHub, schedules, and other integrations. Tembo also reviews pull requests.

  • Price: Free with a one-time $10 allowance for up to three users. Pro costs $60 a month for up to five users and Max $200 for up to ten. What you pay becomes usage allowance for models and VM time.
  • Best for: Choosing among many agents, starting work from many tools, and reviewing pull requests.
  • Read more: Boxes.dev vs Tembo.

AQ

AQ runs Claude Code, Codex, and other agent CLIs in Git worktrees on one machine. It streams each live terminal, editor, and preview to the browser, where teammates and reviewers can watch, comment, and take turns typing. On Free, the machine is a personal sandbox. On Team, AQ can run an always-on machine for you, or you can use VMs in your own cloud.

  • Price: Free for one person, and Team $50 per user each month in early access.
  • Best for: Working live with teammates and getting feedback from outside reviewers.
  • Read more: Boxes.dev vs AQ.

Ellipsis

Ellipsis runs Claude Code and Codex in cloud sandboxes that last up to an hour, with agents, environments, and mention handlers defined in YAML. An environment lists repositories, setup scripts whose results are saved for reuse, variables, secrets, and MCP servers. A follow-up resumes the session with its workspace.

  • Price: Free for individuals with a Claude or ChatGPT subscription. Organizations pay for model usage plus 20% and for CPU and memory by the hour, with no seat fee.
  • Best for: Defining agents as code, with hard spending caps and scoped GitHub tokens.
  • Read more: Boxes.dev vs Ellipsis.

CloudCLI

CloudCLI runs Claude Code, Codex, Cursor CLI, Gemini CLI, and OpenCode in persistent cloud containers called environments, with your own subscriptions or keys. Each environment is an Ubuntu container that starts from your cloned repository and can hold several projects. You or the agent can install the rest of your stack with apt. You can open an environment in CloudCLI's web app, including on a phone, in VS Code, or over SSH. An environment stops after an idle period and keeps its files.

  • Price: Hobby costs €7 a month for one active environment with 2 CPU cores and 4 GB of memory, Growth €20 for five, and Team €39 for five environments and five members. CloudCLI's open-source web UI is free to host yourself.
  • Best for: Running several agent CLIs in a low-cost, persistent environment you can reach from any device.
  • Read more: Boxes.dev vs CloudCLI and Cloud development environments for AI coding agents.

Parallel agent workbenches

Workbenches run many agents side by side, each in its own Git worktree, mostly on your own computers. Agents keep working after you close your laptop only if they run on a host that stays on.

Conductor

Conductor is a Mac app for running Claude Code, Codex, Cursor, and OpenCode in parallel, with each workspace in its own Git worktree on your Mac. Paid plans add cloud workspaces: microVMs built from your organization's Cloud Computer, which sleep after four hours without activity. Conductor also has an iPhone app and a CLI.

  • Price: Free for local work. Pro ($50 a month) and Teams ($60 per user each month) add cloud workspaces. Agents use your own subscription or API key.
  • Best for: Starting work from GitHub, managing stacked pull requests, and running Cursor and OpenCode alongside Claude Code and Codex.
  • Read more: Boxes.dev vs Conductor.

T3 Code

T3 Code is a free, MIT-licensed app for running six agents, including Claude Code and Codex, on your laptop or a host you set up, such as a cloud VM. It has desktop, web, iOS, and Android apps. Threads can use Git worktrees on the same host, and a Linux host running T3 Code's background service keeps working after you close your laptop.

  • Price: Free, and you pay for your hosts and agent plans.
  • Best for: Running many agents from a free app on a host you already have.
  • Read more: Boxes.dev vs T3 Code.

Superset

Superset is a source-available desktop app for running Claude Code, Codex, and 18 other CLI agents in Git worktrees on your Mac or a host you add, such as a Mac mini or a cloud server. Pro adds an iPhone app and scheduled automations.

  • Price: Free for one user, and Pro $20 per user each month, or $15 billed yearly. You pay for your hosts and agent plans.
  • Best for: Running many agents on your own computers, collecting team feedback with Pages, and tracking pull requests across repositories.
  • Read more: Boxes.dev vs Superset.

Orca

Orca is a free, MIT-licensed agent IDE for running Claude Code, Codex, and more than 30 other CLI agents in Git worktrees. Agents can run on your computer, an SSH host, or an Orca server you run, and iOS and Android companion apps are in beta.

  • Price: Free, and you pay for your hosts and agent plans.
  • Best for: Running many agents from a free IDE that supports stacked pull requests.
  • Read more: Boxes.dev vs Orca.

Continuum

Continuum is a free Mac, web, and iPhone app for using Claude Code, Codex, Cursor, Grok, OpenCode, Antigravity, and Z.ai with your own subscriptions or keys. Windows, Linux, and CLI versions are in beta. Agents run on your Mac or on Linux and Mac machines you enroll over SSH, and one command installs Continuum's agent and the agent CLIs on an enrolled machine. Each session can work in its own Git worktree, and a session on an enrolled machine keeps running when your laptop sleeps. Cloud Burst can also run a session in a short-lived sandbox that stops when the work is done.

  • Price: The app is free with your own subscriptions or keys. Paid plans, from Plus at $25 a month to Ultra at $500, add Continuum-hosted inference, and Team plans are priced per member.
  • Best for: Switching between several agent subscriptions, sending one prompt to several models, and tracking spending, all on machines you already have.
  • Read more: Boxes.dev vs Continuum.

How to choose

These four questions narrow the field quickly.

Which agent do you want to use?

If you already use Claude Code or Codex, then choosing a cloud agent platform or workbench lets you run them with your existing plan. The model labs' clouds do this as well, but you're locked into that one lab's models, which may be fine if you only ever use one. If you're open to trying a different harness and paying API rates, then Cursor, Devin, Amp, GitHub Copilot, Factory, and Vorflux may suit you as they run their own agents and don't always allow bringing your subscription.

Where should the work run?

Workbenches run agents on your local computer or a server you control. If that's your laptop then it needs to stay awake, and if it's a remote instance then it needs enough system resources to support parallel tasks. On a workbench agents will use worktrees to isolate their coding work, but will share CPU, memory, ports, and databases. Cloud products keep agents working while your laptop is closed. When a machine goes idle, some products delete it, some keep its files, and some, including boxes.dev, keep its running programs, so dev servers resume when it wakes. See Run coding agents in parallel.

Does your app need more than a repository checkout?

Most cloud products build each machine from your repository and setup scripts, which suits apps that a script can install from scratch. If your app needs databases with data in them, background services, or setup that lives outside the repository, check how each product builds its environment. Boxes.dev starts each devbox from a copy of a machine where your app already runs, so the agent can test its work in the running app.

Do you want to check on agents from your phone?

Most of these products have a phone app or a web app that works on phones. Check whether the app only lets you read and reply to threads, or whether you can also open a terminal on the agent's machine. The Boxes.dev iOS and Android apps let you read and reply to threads, inspect files, and open a terminal on a devbox. See Run Claude Code and Codex from your phone.

The best test is a real task. Give the same bug fix or small feature to the agent in each product you're considering, and see which agents can run your app and show you that their change works.

Frequently asked questions

What is a background coding agent?

A background coding agent works on a task without you watching, usually on a cloud machine. When it's done, it reports back with a summary, a diff, or a pull request, so you can start several tasks and review them later. Claude Code cloud sessions, Codex cloud, Cursor's cloud agents, Devin, and Copilot's cloud agent work this way, and so do Claude Code and Codex on a boxes.dev devbox.

Which products run Claude Code and Codex with my own subscription?

Claude Code cloud sessions and Codex cloud use your Claude or ChatGPT plan. Boxes.dev, Replicas, Niteshift, Tembo, AQ, Ellipsis, Conductor, T3 Code, Superset, Orca, and Continuum run both agents with your own subscription or an API key. Your Claude or ChatGPT plan is also heavily discounted. By our estimate, when fully utilized, a Claude Max plan is at least 40x cheaper than the same usage at API prices, and the $200 ChatGPT Pro plan is over 30x cheaper.

Which ones keep working after I close my laptop?

The cloud products do, within each one's limits on how long a task can run. For example, Copilot's cloud agent stops a session after 59 minutes, and Ellipsis runs a sandbox for up to an hour. Workbenches keep working only on a host that stays on, such as a Linux server running T3 Code's background service, or in Conductor's cloud workspaces on paid plans. See Keep coding agents running after you close your laptop.

Sources

Product names are trademarks of their respective owners.

Close your laptop again.

Develop in the cloud: one computer per agent, running your full app, steerable from every device you own.

Also oniOS·Android·CLI