Running agent-generated code without handing over the keys

Agents increasingly write TypeScript to coordinate tools and process results. When that code needs to do something real—authenticate, touch a database, approve an action—the typical approach of just evaling it becomes a security problem: the code inherits the full access of the application around it, and there's no clean way to pause mid-execution for input.

The Run SDK is a new package that executes untrusted JavaScript and TypeScript in a sandboxed QuickJS context. The application exposes only narrow host functions, and execution can be interrupted for authentication or human approval. After a decision, the run resumes without redoing completed work.

pnpm add run

The only path out is the one you build

Run evaluates JavaScript or type-stripped TypeScript in a fresh QuickJS context within a worker thread. There's no direct access to Node.js or the network. The application defines what the generated code can reach via hostFunctions, which become callable globals inside the sandbox:

import { run } from 'run';

const result = await run({

source: `

const orders = await store.listOrders("customer_123");

const total = orders.reduce((sum, order) => sum + order.amount, 0);

return { count: orders.length, total };

`,

hostFunctions: {

store: {

listOrders: async (customerId: string) => {

return database.orders.findMany({ customerId });

},

},

},

});

if (result.status === 'completed') {

console.log(result.value);

}

Here the program can call store.listOrders(), but the database client and its credentials stay safely in the application. Calls cross the boundary through serialization, and host functions may return promises, so existing service clients can be wrapped without being passed in. You can test this in the playground, where code can only reach the host functions on the page.

Host functions are most useful when they map to product-level actions rather than generic plumbing. Exposing orders.refund(id) gives the application a clear point to authorize the user and validate the order. A generic request function would diffuse that authority across the entire boundary.

Why code mode works

The Run SDK is the engine behind code mode tool execution in the AI SDK. Giving an agent a program changes the unit of work: one model response can describe both the calls and the logic connecting them:

const result = await run({

source: `

const accountId = "account_123";

const [account, invoices] = await Promise.all([

crm.getAccount(accountId),

billing.listInvoices(accountId),

]);

const overdue = invoices.filter(invoice => invoice.status === "overdue");

return { account: account.name, overdue };

`,

hostFunctions: {

crm: { getAccount },

billing: { listInvoices },

},

});

Two requests execute concurrently, and filtering happens inside the program. Only the useful result returns to the application. This suits agents that work across multiple internal services—a research agent can combine search results, a support agent can inspect an account without flooding its context with billing data.

The same package can drive a code interpreter or a feature accepting customer-defined transformations. In every case, the application chooses what data and operations exist.

Pausing for auth and approval

Reading invoices isn't the same as issuing a refund. When generated code hits a sensitive operation, the host function can interrupt:

import { getHostFunctionContext } from 'run';

const hostFunctions = {

documents: {

publish: async (draftId: string) => {

const context = getHostFunctionContext();

if (context.resume === undefined) {

context.interrupt({

kind: 'approval',

message: `Publish ${draftId}?`,

});

}

if (context.resume.resolution !== true) {

return { published: false };

}

return publishDraft(draftId);

},

},

};

The interrupted run returns a signed token, which the application stores with the approval request. When a decision arrives, the token resumes the run. Resumption replays the program, but completed host function calls use their recorded results—so work before the interruption doesn't repeat. The interrupted function picks up where it left off with the approval in hand.

The same mechanism handles authentication waits. The application owns the waiting period; the worker doesn't need to stay alive.

Containing runaway code

A sandbox must still cope with infinite loops and oversized results. createRunner() sets shared limits:

import { createRunner } from 'run';

const runner = createRunner({

limits: {

timeoutMs: 10_000,

memoryLimitBytes: 32 * 1024 * 1024,

},

});

Limits can also be set per run. The defaults cover the QuickJS heap and values crossing the host boundary; applications can tighten them for specific workloads. Each invocation gets a fresh context, dynamic evaluation is disabled, and built-in prototypes are hardened. That boundary applies only to generated code—host functions remain trusted application code and must do their own authorization.

Run is built for JavaScript computation inside an application. For workloads needing an OS, package installation, or process-level isolation, Vercel Sandbox is the appropriate tool.

Where the runtime came from

An earlier version lived inside just-bash as js-exec, backed by QuickJS, letting agents write TypeScript against the shell's virtual filesystem. That layer was extracted into Run and reshaped around application-defined host functions. The mechanism was tested in eve, running agent-generated TypeScript against real tools. It now powers code mode in the AI SDK, where existing AI SDK tools are mapped to host functions and Run handles the sandboxed execution.

Getting started

Run supports Node.js 22.13+ and Bun. Installation is straightforward:

pnpm add run

Documentation and the API reference are available at run-sdk.dev.