The agent is the workflow.
Flowe is a visual builder for AI agents. You describe an agent on a canvas: what triggers it, which model answers, what tools it may call, which skills and delegates it has, what guards admit or refuse a caller, and what shape its answer takes. Flowe compiles that canvas into one validated JSON document, a GraphDoc, and a single Flue agent interprets that document at runtime.
In this build
- Licence
- MIT
- Default model
- cloudflare/@cf/zai-org/glm-4.7-flash
- Models that need no key
- 5 of 15
- Channels that can start a conversation
- 5 of 7
- Tools per agent
- up to 32
- MCP servers per agent
- up to 8
Version 0.0.1, and early. STATUS is the present tense: what works, what is half-built, and what is not deployed. Read it before you depend on any of this. docs/STATUS.md

There is no orchestration engine
Most agent builders compile a canvas into a graph and then ship a second program to walk that graph: a scheduler, a step executor, a state machine, a retry loop. The agent becomes one node inside it.
Flowe inverts that. There is exactly one Flue agent in the whole system, called Runner, and it has no behaviour of its own. It reads a GraphDoc out of initialData and configures itself from it: the model, the tools, the skills, the MCP connections, the delegates, the persistent state fields, the structured output. What the canvas produces is data, not code, and the thing that runs it is a model in a loop rather than an interpreter of your boxes and arrows.
- One thing to version
- A published agent is one immutable JSON row, agent_versions.graph. Publishing copies it and rollback repoints at it. There is no second artifact that can drift from the first.
- The runtime stays small
- The whole interpreter is apps/runtime/src/graph/: compose the instructions, resolve the model route, decide which tools are mounted, plan the delegates. Nothing schedules anything.
- The render is synchronous
- Runner cannot await, so it cannot read the database. The resolved document arrives in initialData, and everything asynchronous happens inside a tool run callback. That is a Flue constraint, not a preference: returning a promise from an agent function throws.
What it costs
You do not get arbitrary control flow. If you need a general-purpose node graph with fan-out, joins and a scheduler you can reason about independently of a model, you want a workflow engine, and Flowe is not one.
What is in the box
Everything below is read from the same contract the editor and the runtime read, so a channel that goes dark in one place cannot stay advertised in another.
Channels that can start a conversation
API
liveAlways on, and the one trigger every document must have. Callers create a conversation over /v1 with a per-agent key.
Chat widget
liveA script tag on your own site, with a short-lived token your server mints per signed-in user. For an agent with no site and no server behind it, the same widget is served full-page from a link Flowe hosts, and the editor names what opening that up costs before you open it.
Schedule
liveA cron expression in a time zone you pick. One Worker cron fires each minute and the scheduler dispatches whichever agents are due, read from D1.
Discord
liveSlash commands from a Discord app you connect, verified against its Ed25519 public key. Ordinary channel messages arrive over a gateway WebSocket, which a Worker cannot hold open. That is a limit of the transport rather than a stub, and the editor says so at the switch.
Through an open-wa host you run. The session lives on that host and never here, which is what makes it shippable without a Business API account.
In the contract, not built
- Webhook
- A verified ingress for a provider with no blueprint. The generic channel router is not mounted.
- Event
- The landing pad for a hand-off from another agent. Nothing dispatches one.
What an agent is made of
- Tools, up to 32 of themdocs/GRAPH.md
- Five implementations in the document: http, integration, mock, code and task. http interpolates its template, checks SSRF before the first byte leaves and again on every redirect hop, and enforces a capped timeout and a 256 KB read cap. integration calls a connected app, and the two implemented are Slack and Stripe. mock returns its configured JSON and is badged as a mock on the trace row. code runs in a sandbox you connect with your own account. task parses, saves and publishes, and is not executed yet: the call returns a failed tool result.
- Skills that load on activationdocs/RUNTIME.md
- Only a name and a description sit in the prompt. The instructions arrive as the result of the model calling activate_skill, so a skill costs one catalog line per turn until the model decides the task matches. Restating the instructions in the prompt would pay their whole context cost on every turn, which is the only reason a skill is not simply more instructions.
- Delegates, dispatched by the modeldocs/RUNTIME.md
- Delegation goes through a framework-owned task tool that every agent has. Declared delegates are catalogued by name and description in the system prompt; the child runs in a fresh context and only its final message comes back. An agent with no delegates therefore needs no special case.
- Guards that decide before admissiondocs/THREAT-MODEL.md
- rate and budget windows, scoped to the agent or to a declared variable. On exceeding a window a guard either ends the request with 429 and a Retry-After header, or forces the cheapest model, which outranks even a pinned one. Budget guards are request-count windows labelled est. in the UI, not cost metering: a guard has to decide before admission and token cost is only known after the turn.
- Branches that fail closeddocs/GRAPH.md
- A tool behind a branch whose condition is false is not mounted at all. The model cannot call a tool that does not exist, which is a stronger guarantee than instructing it not to.
- Structured outputdocs/GRAPH.md
- Give the document an output section and the runtime mounts one more tool, submit_result, whose arguments are the typed response. Calling it writes a client-facing data part and terminates the turn, so the model cannot keep talking after the answer has been submitted.
- MCP servers, mounted per agentdocs/RUNTIME.md
- Up to 8 remote servers per agent. A GraphDoc names them by row id only: no URL, no transport and no credential ever enters a document, because a document is deep-copied into every published version and stored forever. The credential is decrypted from D1 per request, inside the auth closure. Flowe consumes MCP and does not serve it.
- Approvals bound to the argumentsdocs/THREAT-MODEL.md
- Every tool is auto or ask. An ask approval is bound to sha256 of the canonicalised arguments, so approving a refund of one amount can never authorise a refund of another: different arguments hash differently and need their own approval. Approvals are single use, consumed by a compare-and-set, and time-bounded.
- Model routing and a fallbackdocs/GRAPH.md
- First-match rules choose a model per turn, a pin bypasses the router entirely, and a fallback model layers on top of the framework’s own recovery advisories. The catalog is one list both the editor and the runtime read.
- A trace with nothing to redactdocs/OBSERVABILITY.md
- Seventeen row types come off the framework observe stream, and a compile-time assertion fails pnpm check:types if a row is added to the union and not to the list. Whether a row reaches a caller is structural rather than filtered: on an end-user surface the row is never written, so there is nothing to redact.

Quickstart
Requires pnpm 10 (this repo pins pnpm@10.33.0) and Node 22 or newer. CI runs Node 24. The Node and Docker target needs 23.4 or newer, because it reaches SQLite through node:sqlite.
git clone https://github.com/khalilelghoul01/flowe.git flowe
cd flowe
pnpm install
pnpm run setup
pnpm devThen open http://localhost:3000, sign up, and create an agent from a template.
pnpm run setup, spelled out. setup is one of pnpm’s own commands, so a bare pnpm setup runs that one instead of scripts/setup.mjs.
What setup does
- 01Generates local secrets with node:crypto and writes apps/web/.dev.vars and apps/runtime/.dev.vars from their templates. It never overwrites a value you already set.
- 02Keeps SECRETS_KEY and INTERNAL_KEY byte-for-byte identical across the two files. Those two disagreeing is the most common way local dev breaks, so the script refuses to guess when it finds a mismatch and tells you to fix it by hand.
- 03Applies the local D1 migrations with wrangler.
It is plain Node with no dependencies and safe to re-run: scripts/setup.mjs
The rest of the commands
- pnpm run setup
- Generate local secrets, write both .dev.vars, migrate the local database.
- pnpm dev
- Both apps: web on :3000 through Next, runtime on :5173 through Vite.
- pnpm test
- Vitest across every workspace package.
- pnpm check:types
- tsc --noEmit across the workspace.
- pnpm build
- Build both apps for the Cloudflare target.
- pnpm run deploy
- Deploy the runtime Worker, then the web Worker.
A fresh agent answers with no provider key
Cloudflare Workers AI runs through the runtime Worker’s own AI binding and needs no API key at all, so the product works end to end with zero provider credentials configured: install, sign up, pick a template, get a reply. DEFAULT_MODEL and every shipped template point at a keyless entry for exactly that reason.
The default model
cloudflare/@cf/zai-org/glm-4.7-flash5 keyless Workers AI models are in the catalog, spanning fast, balanced, frontier, so switching models is a real decision and not a paywall.
OpenRouter, Anthropic and OpenAI are opt-in upgrades. Store a key from the dashboard under Integrations, or set it as an environment variable on the runtime Worker. The Node and Docker target has no Workers AI, so a provider key is required there.
A published agent is an HTTP endpoint
Publish, mint a key, and the agent answers at /v1/agents/<slug>. This is the same round trip the dashboard prints on the API page. The send is accepted immediately and the reply arrives on the conversation, so the second call reads it back; the read is a passthrough onto the framework’s own stream, so view=updates and live=sse work on the public endpoint too.
AGENT="https://<your-runtime-worker>/v1/agents/support"
KEY="fl_live_..."
# 1. create a conversation with the first message on it
CID=$(curl -sS -X POST "$AGENT/conversations" \
-H "Authorization: Bearer $KEY" \
-H "Content-Type: application/json" \
-d '{"body":"Where is order RG-10482?","metadata":{"user":{"id":"u_123"}}}' \
| jq -r .conversationId)
# 2. the send is accepted immediately; poll until the turn settles
curl -sS "$AGENT/conversations/$CID" -H "Authorization: Bearer $KEY"metadata is projected onto the document’s declared variables only, so a caller cannot inject a variable that a router rule or a branch condition reads.
What you can see when it runs


Two Workers, one database
Two Workers rather than one because they cannot share a build: the runtime is compiled by Vite with the Flue plugin, which generates a Durable Object class per registered agent, and the dashboard is built by OpenNext into its own bundle.
apps/web
Next 15 on the App Router, built by OpenNext. The dashboard, BetterAuth, and the tRPC admin API. Owns every human-initiated write to D1, and every admin query is owner-scoped in its WHERE clause rather than fetched and then checked.
apps/runtime
Hono. Public /v1, /internal behind a shared key, /channels for provider ingress. One Flue agent, one Durable Object class, and the Workers AI binding.
One D1, both Workers
Both wrangler configs declare the same database. They talk over a service binding on Cloudflare and over plain HTTP in local dev, because a service binding cannot reach a separately running dev server.
Everything that guards a request is pre-admission middleware
It can live nowhere else. After admission the framework issues internal deterministic requests that preserve no headers, cookies, query parameters or body, so a check inside the agent would not see the caller. In order:
- 01resolve the slug to an agent and a version, 404 if either is missing
- 02verify the API key, 401 or 403
- 03run the guards, 429 with Retry-After, or downgrade the model
- 04meter usage, on waitUntil, never on the hot path
- 05compose the instance id and attach the resolved document as initialData
What is not true yet
This is the short form. STATUS carries the long one, dated, with paths for every claim. docs/STATUS.md
- No server-side run recording
- run_events is written by exactly one caller: the test console, and the server stamps every row as a draft test run so a browser cannot file something that looks like production traffic. Production traffic emits trace rows to the Worker log and nowhere durable, which is why there is no runs page you can open days later.
- Provider keys are deployment-wide
- The framework resolves credentials through a callback that receives no request, agent or owner context, so it cannot know whose key to use. If two owners have stored a key for the same provider the runtime refuses to guess and returns unconfigured, rather than signing one user’s traffic with another’s credential.
- No teams and no roles
- The schema is multi-user-ready through ownerId. The interface is single-user.
- API keys are per agent, not per account
- api_keys.agent_id is a non-null foreign key to agents, and the admin surface has no machine credential at all. This is what blocks an agent that builds agents, and the roadmap says so.
- Knowledge is inlined, not retrieved
- Entries are pasted into the prompt and the corpus is capped as a whole. There is no chunking, no vectors and no retrieval. The cap is mechanical: the render is synchronous, so the whole document travels to the instance on every conversation start.
- Parts of the document parse without running
- A task tool, the phases section, continuations, and the skill, delegate and instruction arms of a branch all validate, save and publish, and the runtime does not interpret them. GRAPH has the table, field by field.
- Flowe consumes MCP and does not serve it
- An agent can mount remote MCP servers. Nothing can drive Flowe over MCP, and there is no MCP server endpoint on either Worker.
- Channel redelivery is suppressed in memory
- The suppression cache is an in-isolate LRU with a ten-minute TTL, and it is honest about being one. A provider retry that landed on a different isolate would run a second turn and send a second reply.
Documentation
10 documents in the repository, changed in the same pull request as the code they describe.
- docs/STATUS.md
What works today, what is half-built, what is not deployed. Dated.
- docs/ROADMAP.md
What comes next, in rough order, with the blockers named.
- docs/ARCHITECTURE.md
The two Workers, the database, the request pipeline.
- docs/RUNTIME.md
The Runner agent, the synchronous render, tool execution.
- docs/GRAPH.md
The GraphDoc contract, section by section, and how it evolves.
- docs/DECISIONS.md
Why the architecture is shaped this way.
- docs/THREAT-MODEL.md
Secrets, keys, approvals, SSRF guards, the trust boundaries.
- docs/OBSERVABILITY.md
Trace rows, the console, and what is deliberately not logged.
- docs/DEPLOY.md
Cloudflare end to end, plus the Node and Docker target.
- docs/DEVELOPMENT.md
The engineering handbook: how to run it, test it, and the house rules.