Quickstart
Install @statelyai/agent and run your first agent machine end to end.
Alpha:
@statelyai/agent2.0 is in alpha. APIs can change between releases; pin an exact version. Feedback: github.com/statelyai/agent.
The overview has the copy-paste version; this page builds the same shape up piece by piece.
Installation
pnpm add @statelyai/agent@alpha xstate@alpha zod ai@^6 @ai-sdk/openai@^3- The
@alphatag floats. Install it once, then pin what it resolved to (npm ls @statelyai/agent) so a later alpha cannot change the API under you. xstateis the one required peer, at v6 alpha.25 or newer. Node 22.18 or newer.- Provider packages must match your
aimajor.@ai-sdk/openai@^3pairs withai@^6; a bare@ai-sdk/openairesolves to@latest, whoseLanguageModelspec version may not match youraipeer. - The package is ESM-first; every entry also ships a CommonJS build, so
require()works. The examples use top-levelawait, which needs ESM: set"type": "module"inpackage.json(or use.mtsfiles).
First agent
A comment moderator. The model reads a comment and picks one of three events; the machine owns the trust threshold that decides whether publishing is even legal.
Two pieces do the work:
setupAgentdeclares the schemas and events, thencreateMachineauthors the control flow from them.createScriptedExecutorsplays back canned answers through the executor contract, so this first version runs end to end with no API key.
import { createScriptedExecutors, runAgent, setupAgent } from "@statelyai/agent";
import { z } from "zod";
const outcomeSchema = z.enum(["published", "flagged", "blocked"]);
const agentSetup = setupAgent({
context: z.object({
comment: z.string(),
trust: z.number(),
outcome: outcomeSchema,
reason: z.string().nullable(),
}),
input: z.object({ comment: z.string(), trust: z.number() }),
output: z.object({ outcome: outcomeSchema, reason: z.string().nullable() }),
events: {
PUBLISH: {}, // `{}` is shorthand for a payload-less event
FLAG: z.object({ reason: z.string() }),
BLOCK: {},
},
});
const moderationMachine = agentSetup.createMachine({
// Comments start flagged; only the machine can clear one.
context: ({ input }) => ({ ...input, outcome: "flagged", reason: null }),
output: ({ context }) => ({ outcome: context.outcome, reason: context.reason }),
initial: "reviewing",
states: {
reviewing: {
invoke: {
src: "agent.decide",
input: ({ context }) => ({
model: "fast",
system: "PUBLISH harmless comments, FLAG borderline ones with a reason, BLOCK abuse.",
prompt: `Comment: ${context.comment}\nAuthor trust score: ${context.trust}`,
allowedEvents: ["PUBLISH", "FLAG", "BLOCK"],
}),
},
on: {
// The guard owns the threshold, not the model: under 50 this returns
// undefined, so PUBLISH is illegal and the model has to choose again.
PUBLISH: ({ context }) =>
context.trust >= 50
? { target: "published", context: { outcome: "published" } }
: undefined,
FLAG: ({ event }) => ({ target: "flagged", context: { reason: event.reason } }),
BLOCK: () => ({ target: "blocked", context: { outcome: "blocked" } }),
},
},
published: { type: "final" },
flagged: { type: "final" },
blocked: { type: "final" },
},
});
const result = await runAgent(moderationMachine, {
input: { comment: "honestly this update is terrible", trust: 20 },
executors: createScriptedExecutors({
decisions: [{ type: "FLAG", reason: "Borderline tone." }],
}),
});
if (result.status === "done") {
console.log(result.output); // { outcome: 'flagged', reason: 'Borderline tone.' }
}Save it as agent.ts in a project with "type": "module" in its package.json, then run it:
npx tsx agent.tsIt prints { outcome: 'flagged', reason: 'Borderline tone.' }. The scripted decision took the place of a model call. Everything else (the guard, the retry, the final state, the schema-checked output) is the real machine.
Real model executors
Swap the executors. The machine does not change: add a models registry (which also types the model: keys), then hand runAgent the AI SDK adapter instead of the script.
import { openai } from "@ai-sdk/openai";
import { createAiSdkExecutors, defineModels } from "@statelyai/agent/ai-sdk";
const models = defineModels({ fast: openai("gpt-5.4-mini") });
const agentSetup = setupAgent({
models, // adds typed `model` keys; everything else is unchanged
// …the same schemas and events
});
const result = await runAgent(moderationMachine, {
input: { comment: "honestly this update is terrible", trust: 20 },
executors: createAiSdkExecutors({ models }),
});export OPENAI_API_KEY=sk-...
npx tsx agent.tsNow the model picks the event, and reason is whatever it wrote. Keep the scripted set around: it is what your tests run against (see Testing without an API key).
What the machine does that a prompt call cannot:
- The model only ever picks a legal event.
allowedEventsis intersected with the events the current state actually accepts, so the choice set moves with the machine. - A guard can overrule the model. A
PUBLISHon a low-trust author is rejected before it reaches state (failure: 'rejected-by-guard') and the decision retries with that feedback. The threshold lives in one place and cannot be prompted away. - The outcome is a state, not a parsed string. Every path ends in a known final state with schema-checked output, and the whole graph renders as a diagram.
The setupAgent surface
setupAgent returns a setup (not a running agent) you author machines from, like XState's setup(). Context, input, output, and event payloads are Standard Schemas, so Zod works directly and the machine's types come from them.
The model value is a key into the models registry, or any string your host resolves at run time.
Every machine can invoke these built-in actor sources. They are reserved src strings; the invoke's input shapes each call.
src | Purpose |
|---|---|
agent.generateText | Inline one-shot text (or structured-output) model call. |
agent.streamText | Same, streamed chunk by chunk through onChunk. |
agent.decide | Model picks exactly one currently-legal event. |
agent.userInput | Gather human input mid-run without settling. |
Named requests
A request is a typed, reusable model call: named schemas, a model, and a prompt built from its input. It is the testable counterpart to an inline agent.generateText.
const agentSetup = setupAgent({
models,
// …schemas as above
requests: {
moderatorNote: {
schemas: {
input: z.object({ comment: z.string() }),
output: z.object({ note: z.string() }),
},
model: "fast",
prompt: ({ input }) => `Write a one-line moderator note for: ${input.comment}`,
},
},
});Each request key becomes an invocable src:
blocked: {
invoke: {
src: "moderatorNote",
input: ({ context }) => ({ comment: context.comment }),
onDone: ({ output }) => ({ target: "done", context: { reason: output.note } }),
},
},See Text requests for tools, streaming, and messages.
Running
runAgent drives the machine and calls your executors whenever it needs a model. It settles with a status:
done: reached a final state;result.outputmatches your output schema.idle: waiting on a human. See Human in the loop.error: something threw.
Every variant carries result.events, a versioned, JSON-safe AgentLogEntry[] of the replayable external inputs (identity, timestamp, machine version, state/effect hashes). Pass it to replay(machine, result.events) to reconstruct the final snapshot without re-running model or tool calls; verifyReplay requires every hash. See The event log.
For a run that must go straight through to a final state, generateResult(machine, options) resolves with the done result: result.output plus metadata (result.snapshot, replayable result.events, aggregated result.usage), like generateText's text + call metadata. It throws AgentIdleError if the machine pauses.
Other ways to run the same machine
runAgent is the default, not the only option. provideExecutors binds the executors onto the machine and hands it back for a plain createActor; the step path lets a durable host own the loop and the persistence. The machine is identical in all three. See Choosing a run mode.
Testing without an API key
createScriptedExecutors is the same helper the first run used: FIFO queues of scripted answers, one entry consumed per model call. decisions feeds agent.decide, text feeds every text request.
const executors = createScriptedExecutors({
decisions: [{ type: "FLAG", reason: "Borderline tone." }],
text: [{ note: "Repeat offender." }],
});
const result = await generateResult(moderationMachine, {
input: { comment: "…", trust: 20 },
executors,
});
expect(result.output).toEqual({ outcome: "flagged", reason: "Borderline tone." });- An entry may be a function of the request, for branching or looping machines:
(request) => request.name === "moderatorNote" ? … : …. - A
PUBLISHrejected by the guard consumes an entry and retries with the next one, so scripts assert retry behavior directly. - Running dry throws an error naming the pending request.
Executors are plain functions, so a hand-written mock is just an object; see Hosts and executors. To assert a whole playthrough without a run loop at all, simulateAgent walks the step path from a by-src script. See Testing and verification.
Live visualization
The Stately Inspector is a web UI that draws a running machine as a diagram and highlights each state as it is entered. Install @statelyai/sdk, create an inspector, and pass its handler to runAgent's inspect option. The run opens the diagram in your browser and updates it live:
import { createInspector } from "@statelyai/sdk";
// Uses Stately's hosted relay by default. Pass `url` for a self-hosted relay.
const inspector = createInspector();
await runAgent(moderationMachine, {
input: { comment: "honestly this update is terrible", trust: 20 },
executors: createAiSdkExecutors({ models }),
inspect: inspector.inspect, // opens the diagram and lights it up live
});It is the same machine you authored, so it also renders in Stately Studio and the VS Code extension. See Observability for production tracing.
Related
- Agent machines: authoring states, transitions, and typed context.
- Decisions: let the model choose one of several legal machine events.
- Hosts: model aliases, the AI SDK adapter, and writing your own executors.
- Choosing a run mode:
runAgent,provideExecutors, or the step path. - Where state lives: the event log, snapshots, and the shipped stores.