Blog Cloud Platform

OpenClaw Architecture: From LLM to Stateful AI Agent

18 min read

OpenClaw turns a large language model (LLM) into a stateful AI agent by wrapping model inference in a persistent runtime. The model handles reasoning and tool selection. OpenClaw handles routing, sessions, memory, tools, browser control, computer access, execution policy, and persistent storage.

The distinction matters. A basic prompt-response application returns an answer to a prompt. OpenClaw runs the model inside a continuing agent loop where previous actions, tool results, session history, memory, and machine state influence what happens next.

As of October 8, 2026, the latest stable release is OpenClaw v2026.9.9. In a standard Gateway deployment, OpenClaw uses SQLite for active session and transcript state, supports persistent memory, and can recover eligible interrupted work after restarts. Depending on configuration, the runtime connects model reasoning with browser, file, shell, messaging, plugin, Model Context Protocol (MCP), and computer-control capabilities.

Key Takeaways

OpenClaw works as a stateful computer-using agent because several runtime layers sit around the LLM. Each layer solves a different part of the agent lifecycle.

  • The Gateway receives events and selects the appropriate agent and session.
  • The agent runtime assembles workspace instructions, history, relevant memory, a skills catalog, and permitted tools.
  • The model decides whether to respond or request an action.
  • Tool results return to the agent loop for another reasoning step.
  • Sessions and memory preserve useful state beyond one model response.

This architecture keeps model reasoning separate from persistence and execution. The runtime gives each model call access to the context and capabilities needed for the current task.

What Turns An LLM Into An OpenClaw Agent?

An LLM predicts tokens from supplied context. An OpenClaw agent adds routing, persistent state, tools, memory, policies, and execution surfaces around those model calls.

A basic prompt-response application follows a short cycle:

Prompt → model → response

A tool-using OpenClaw run follows a longer cycle:

Request → routing → context assembly → model → action → observation → model → state update → response

The Gateway receives an event and maps it to an agent and session. The agent runtime builds working context from workspace instructions, session history, relevant memory, skills information, and tool definitions. The selected model processes this context and either returns a response or requests a tool.

A tool result then becomes new input for the same run. The model reviews the result and, when needed, selects another step. This loop turns inference into ongoing work.

Model intelligence and agent state also live in separate layers. Changing the inference model does not inherently require rebuilding session storage, workspace files, or browser profiles. Memory embeddings have their own configuration: changing the provider or model may require rebuilding the vector index.

The OpenClaw Gateway Provides The Persistent Control Plane

The OpenClaw Gateway connects users, sessions, agents, tools, channels, and remote execution surfaces. A single Gateway process owns session state and coordinates connected clients.

In a Gateway-backed workflow, a Slack message, Telegram message, web request, terminal interaction, or scheduled event passes through the Gateway before reaching the model.

The Gateway first identifies where the event belongs. Routing information links the request to:

  • An agent
  • A conversation
  • An account
  • A channel
  • A session

This design matters when you move between devices. Multiple clients reference Gateway-owned session state. A conversation started through one supported client keeps the same server-side session identity when another authorized client attaches to that session.

The Gateway also provides the control boundary for connected nodes, approvals, authentication, and session access. Local execution paths, such as openclaw agent execcan run an embedded agent turn without connecting to a Gateway.

A Request Becomes A Stateful Agent Loop

The OpenClaw agent loop converts one request into a sequence of observations, decisions, and actions. Each completed tool call changes the information available to the next reasoning step. Runs are serialized per session key to reduce conflicts between work using the same conversation state.

Consider this request:

Open our staging dashboard, check the failed deployment, read the latest error, and save a short incident note.

Step 1: OpenClaw Selects The Session

The Gateway associates the incoming request with the appropriate agent and session. This step determines which conversation history and runtime state belong to the request.

Step 2: The Runtime Builds Context

OpenClaw draws on relevant workspace files when preparing agent context. Common files include:

  • AGENTS.md
  • SOUL.md
  • IDENTITY.md
  • USER.md
  • MEMORY.md

Their inclusion depends on the runtime and session. Some content is injected into the prompt; other content is retrieved when needed. The skills catalog identifies available instructions, while the agent loads relevant SKILL.md files for the task.

Session history, recalled memory, model settings, and visible tool definitions complete the working context. The model therefore receives more than the latest user message.

Step 3: The Model Selects An Action

The LLM reads the request and available context. For the staging example, the model might request a browser action first.

Step 4: OpenClaw Executes The Tool

The runtime sends the browser request through the permitted tool surface. If the action succeeds, the browser opens the target page and returns page state, text, screenshots, or another observation. A refused or failed action returns information the model can use to decide how to proceed.

Step 5: The Model Reviews The Observation

The browser result enters the active agent loop. If the page shows a failed deployment, the model selects the next relevant tool, such as shell access, a log, or a file reader.

Step 6: OpenClaw Persists Useful State

The final response is only one result from the run. Session rows, transcript events, workspace changes, selected memory updates, downloaded files, and external system changes can provide continuity for later work.

This repeated cycle explains the “stateful” part of a stateful AI agent.

Sessions Hold Conversation And Runtime State

Sessions preserve the active conversation path and related runtime metadata. Current OpenClaw releases store session rows and transcript events in a per-agent SQLite database on the Gateway host.

The default per-agent path is:

~/.openclaw/agents/<agentId>/agent/openclaw-agent.sqlite

Replace <agentId> with the configured agent’s identifier. The default agent is main.

OpenClaw uses two persistence layers inside this storage model:

  • Session rows track mutable runtime metadata such as session identity, activity timestamps, settings, and token counters.
  • Transcript events store conversation entries, tool calls, and compaction summaries. OpenClaw later rebuilds model context from these events.

Older OpenClaw releases used sessions.json and JSON Lines (JSONL) artifacts more directly. The current runtime uses SQLite, while older files serve migration, archive, recovery, and offline maintenance roles.

Why SQLite Matters

SQLite gives OpenClaw a structured local persistence layer for active session rows and transcripts. OpenClaw uses this storage to implement:

  • Bounded history reads
  • Transcript search
  • Restart recovery
  • Retention controls
  • State management for long-running agents

For users, the practical benefit is continuity. Conversation state does not depend on keeping every earlier message inside one model context window.

Compaction helps manage that window by summarizing older conversation content while retaining recent messages. The runtime can therefore build a bounded working context from a longer stored conversation.

Memory Gives OpenClaw Durable Recall

Sessions answer “what happened in this conversation?” Memory answers “what information should remain useful later?”

OpenClaw keeps these jobs distinct. The built-in memory engine indexes durable workspace memory into a per-agent SQLite database.

Memory sources include:

  • MEMORY.md
  • A root USER.md, when present
  • memory/*.md files

The built-in engine supports FTS5 full-text keyword search and, with a working embedding provider, vector search. Hybrid retrieval combines keyword and semantic matching. Ranking also considers relevance, recency, recorded importance, and diversity. Without an embedding provider, keyword search remains available.

This design avoids using the LLM context window as permanent storage. A long-running agent accumulates far more information than a single prompt can carry. Memory search retrieves relevant information when needed, keeping the working context focused on the current task.

Session History Vs. Memory

Session history records the conversation and tool activity. Memory stores selected information intended for future reuse.

For example:

  • A deployment error from today belongs in session history.
  • A stable staging URL fits long-term memory.
  • A preferred deployment process fits long-term memory.
  • A recurring project constraint also fits long-term memory.

The distinction helps organize context and control token usage. The search paths can overlap: memory search can also include session material when configured to do so.

Computer Use Extends The Loop Beyond Text

A computer-using agent needs two capabilities:

  • The ability to observe a user interface
  • The ability to perform actions against that interface

OpenClaw exposes both browser automation and computer control as runtime surfaces.

Browser Control

The OpenClaw-managed browser uses a separate Chromium-family profile for agent activity. The browser tool supports:

  • Tabs
  • Page reading
  • Clicking and typing
  • Dragging
  • Screenshots and snapshots
  • PDF output
  • Downloads

OpenClaw also supports selected existing browser sessions through separate attachment paths. Feature availability varies by backend: PDF export and download interception require a Playwright-backed profile, such as the managed openclaw profile.

The important architectural detail is observation continuity. After an action, the agent checks a fresh observation to determine whether the page changed as intended. A successful download produces a file that a later tool can read. Each result informs the next step.

Computer Control

The built-in computer tool works with supported Gateway desktops, paired machines, and desktop-enabled cloud sessions. Availability depends on the selected provider, desktop configuration, tool policy, and operating-system permissions.

Current support includes macOS, Windows, and Linux paths. Windows and Linux use the optional cua-computer provider, with local computer use still described as experimental. Controlling a server’s own desktop requires a configured graphical environment. A headless Gateway can also connect to a supported desktop on another machine.

Depending on the active provider, computer control supports actions such as:

  • Screenshots and visual observations
  • Pointer movement and mouse clicks
  • Keyboard input
  • Scrolling and dragging
  • Window interaction
  • Element actions, where supported

A vision-capable model interprets the observed interface. Policy and operating-system permissions govern access.

Tools, Skills, Plugins, and MCP Serve Different Roles

OpenClaw separates actions, workflow instructions, runtime extensions, and external protocol integrations. This separation makes the execution surface easier to understand and restrict.

Tools Perform Actions

A tool is a callable function the model can use. Examples include:

  • Shell execution through exec
  • Browser control
  • Web search
  • Messaging
  • File operations
  • Image generation
  • Computer control

Tool policy decides which functions reach the model for a specific run. Individual tools can also depend on configured providers, credentials, or other prerequisites.

Skills Teach Repeatable Workflows

A skill is an instruction package built around a SKILL.md file. Skills describe how the agent should approach a recurring task and use its available tools.

Common skill use cases include:

  • Defining task-specific procedures
  • Providing reusable instructions
  • Describing tool usage patterns
  • Adding domain-specific workflow guidance

A skill can include supporting scripts and assets. Those scripts still run through an available execution tool; the skill itself does not create a new low-level execution primitive.

Plugins Extend The Runtime

Plugins add runtime capability. A plugin might add:

  • Tools
  • Model providers
  • Channels
  • Hooks
  • Search services
  • Speech features
  • Packaged skills

MCP Connects External Capability Servers

Model Context Protocol provides another path. OpenClaw supports it in two directions: openclaw mcp serve exposes Gateway-backed conversations to external MCP clients, while saved outbound server definitions let supported OpenClaw runtimes connect to external MCP servers.

MCP integrations might expose:

  • External application tools
  • Databases
  • Development services
  • Internal systems
  • Specialized APIs

Actual tool availability depends on the runtime, server configuration, connectivity, credentials, and tool policy. Sandboxed sessions apply an additional tool-policy gate, so saving a server definition does not guarantee that its tools will be available to the agent.

Keeping these layers separate gives administrators more control over the instructions the model receives, the tools it sees, and the operations it can execute.

Where The OpenClaw Gateway Runs Matters

OpenClaw supports local and server-hosted deployment models. The Gateway host becomes part of the agent’s security, availability, storage, and execution boundary.

A laptop works for personal testing and local workflows. An always-on server fits workloads involving:

  • Remote messaging
  • Scheduled work
  • Persistent sessions
  • Continuous availability
  • Remote agent access

Our one-click OpenClaw image prepares a Cloud Server and the base application on the Atlantic.Net On-Demand Cloud platform. You connect a supported model provider, authorize users, and configure your preferred channels. OpenClaw remains self-managed, including its application configuration, updates, and tool policies.

The hosting choice preserves the underlying OpenClaw agent model. The Gateway still owns routing and session state, while you must explicitly configure models, tools, credentials, storage, and policies.

When choosing a host, decide:

  • Which machine receives credentials
  • Where session data stays
  • Where tools execute
  • Which clients can reach the Gateway
  • Which network boundaries protect the runtime

Security Sits Between Model Intent And Real Action

An LLM with shell, browser, file, messaging, and desktop access has meaningful authority. OpenClaw therefore places execution policy outside model reasoning.

Three layers deserve close attention:

  • Sandbox configuration controls where eligible operations run. Sandboxing is optional and disabled by default; the Gateway process remains on its host.
  • Tool policy controls which tools the agent can call. A permitted general-purpose tool, such as a shell, can perform several kinds of actions.
  • Execution approvals and allowlists control selected host operations according to the effective policy.

OpenClaw’s security trust model treats one Gateway as one trust boundary. The supported model assumes one operator or a mutually trusting group per Gateway. Mutually untrusted users require separate Gateways, preferably under separate operating-system users or hosts.

This distinction matters for team deployments. Session keys select conversation context; the Gateway’s access controls handle authorization.

Current defaults allow broad session access within a Gateway: tools.sessions.visibility defaults to alland ordinary agent-to-agent access is enabled. Tool availability and sandbox restrictions still apply. Sandboxing limits the sandboxed caller’s reach, but does not automatically hide its transcripts from other permitted agents.

Administrators who need narrower separation should explicitly review:

  • Session visibility
  • Agent-to-agent access
  • Sandbox policy
  • Tool exposure
  • Execution approvals
  • Credential access

Why Human Approval Still Matters

A model proposes actions from prompts and observations. Runtime policy decides which actions reach real systems.

Approval requirements deserve particular attention around high-impact operations such as:

  • Command execution
  • File modification
  • Credential access
  • Remote machine control
  • External messaging

Human review depends on configuration. An unconfigured Gateway or node uses a permissive host-execution baseline, so administrators should check the effective approval policy before granting access to sensitive systems.

Host-command approvals, plugin approvals, and native runtime permissions are separate controls. File changes and outbound messages are not automatically covered by one universal approval gate. Operating-system permissions, credentials, tool restrictions, and destination controls must support the intended access boundaries.

Basic LLM Chat Vs. OpenClaw Agent

A basic chat application centers the conversation. OpenClaw centers a persistent runtime where conversation becomes one input alongside tools, files, memory, scheduled work, browser state, and machine state.

The table uses a basic prompt-response application for comparison. Features depend on how each application is implemented.

Area Basic Prompt-Response Application OpenClaw Agent
Execution model Prompt then response Multi-step agent loop
Session state Application dependent Per-agent SQLite state
Long-term memory Application dependent Workspace memory plus indexed recall
Tool execution Requires tool Runtime tool layer with policy
Browser control Optional Managed and attached browser paths, with different capabilities
Computer control Optional Supported Gateway, node, and cloud desktop paths
External integrations Application dependent Plugins and MCP
Recovery Depends on stored state and recovery logic Persistent conversations and recovery of eligible interrupted turns
Security boundary Application, service, and host controls Gateway access, configured sandboxing, tool policy, and permissions
Model layer Selected by the application Multiple model-provider paths

Browser automation forms one part of OpenClaw’s system. Sessions, memory, files, shell access, plugins, MCP, routing, and remote computers are part of the broader runtime model.

Who Should Use OpenClaw?

OpenClaw fits users who need persistent agent workflows. Typical use cases include:

  • Developers who need an agent to work with repositories, terminals, browsers, and project files.
  • Infrastructure and operations teams that need agents to inspect systems, follow runbooks, check dashboards, or summarize incidents.
  • Power users who want continuity across messaging channels and devices.
  • Teams building multi-step workflows around tools, memory, and persistent sessions.
  • Users who need scheduled or remote agent workflows on an always-on Gateway.

OpenClaw fits less naturally when a task only needs occasional question answering. A basic chat application can meet that need with fewer components to configure.

The decision therefore depends on state and action requirements. If your workload needs memory, persistent sessions, tool use, remote access, or multi-step execution, an agent runtime provides the supporting infrastructure.

End-To-End Example: From Message To Machine Action

A full workflow shows how these layers work together. Suppose a developer sends this message through Slack:

Check the staging site after the latest deploy. If the homepage returns an error, inspect the deployment log and write a short incident summary.

With the necessary tools, credentials, and access configured, the workflow could proceed like this:

  1. The Gateway receives the event and resolves the channel, sender, agent binding, and session.
  2. The agent runtime loads workspace instructions and relevant session history.
  3. Memory retrieval supplies available project context, such as the staging URL or previous deployment notes.
  4. The runtime exposes the tools permitted by the active policy.
  5. The model selects the browser tool.
  6. The browser opens the staging site and returns a snapshot or page observation.
  7. If the model identifies a failure, it selects a log-access path.
  8. A shell tool, deployment, plugin, or MCP tool returns the relevant log output.
  9. The model reads the error and asks a file tool to save an incident note.
  10. The response goes back through the original conversation route.

Meanwhile, several forms of state remain available:

  • The transcript records model and tool activity.
  • Workspace changes remain on disk.
  • Generated incident files persist.
  • Selected durable information can be written to memory for future recall.
  • External systems retain changes made through permitted tools.

A later question such as “Did we see the same failure last week?” can draw on this saved state, subject to access and retention policies.

Session search provides conversation evidence. Memory search provides indexed project knowledge and, where configured, relevant session material. Stored files provide artifacts from earlier work. The runtime carries state from one action into the next.

Why Stateful AI Agent Architecture Matters

State changes the model’s role. The LLM becomes one reasoning component inside a larger software system, with the runtime supplying the history and operational context needed for each turn.

OpenClaw separates several forms of state:

  • Session state tracks conversation continuity.
  • Transcript state records messages and tool events.
  • Workspace files hold operating context and agent instructions.
  • Memory stores durable knowledge.
  • Browser and desktop sessions hold interface state.
  • External systems keep their own state behind tools, plugins, and MCP servers.

This separation also supports model portability. A different inference provider can use the same core concepts, sessions, workspace files, memory, tools, and Gateway routing, subject to its supported runtime and configuration.

State makes restart recovery meaningful. OpenClaw can resume or reconcile eligible interrupted conversation turns from saved state after a Gateway restart. The parent agent decides how to continue interrupted subagents. Process-local terminals are not restored, and uncertain tool outcomes must be checked before actions are repeated.

For long-running agents, persistence is part of the execution model.

Common Misconceptions About OpenClaw

OpenClaw Is More Than A Chatbot

Chat is one input surface. The broader runtime includes Gateway routing, sessions, memory, tools, browser control, computer access, plugins, and automation.

Memory Is Not The Same As Session History

Session history records conversation and tool events. Memory stores durable knowledge selected for reuse. Keeping those purposes clear helps organize what the agent should retrieve for a task.

Tool Use Does Not Mean Unlimited Autonomy

Available tools, configured sandbox rules, operating-system permissions, credentials, session scope, and approvals determine what an agent can reach. The model chooses from the capabilities the runtime exposes, and the effective configuration determines the limits.

Computer Use Is More Than Browser Automation

Browser control targets web pages and browser sessions. Computer control can interact with desktop environments, applications, windows, pointer input, keyboard input, and screen observations, according to the active provider’s capabilities.

OpenClaw From LLM To Persistent Agent

OpenClaw turns an LLM into a stateful, computer-using agent by wrapping model inference with routing, persistent sessions, memory, tools, machine access, and an execution policy.

The Gateway receives and routes work. The agent runtime assembles context. The model chooses the next step, and tool results return to the reasoning loop. Sessions, memory, and stored files preserve useful state after the immediate response ends.

Together, these layers let a prompt become one event in a continuing software process. The practical value comes from being able to return to earlier work, retrieve relevant context, and continue a task through the tools and permissions you have configured.

For an always-on deployment, start by deciding where session data will live, which systems the agent can access, and how tool execution and recovery will be controlled. If you are evaluating a cloud-hosted Gateway, contact our team to discuss the infrastructure for your OpenClaw workload.

Share

Written by

View all articles →

Ready to Deploy?

Launch secure, compliant, enterprise-grade infrastructure with confidence.