Argento Agents
One brain, many bodies
A behaviour and dialogue engine for characters that live in games, chat and calls. Agents do not just talk: they choose what to do, from a catalogue the body itself declares — so a simulation only ever receives commands it already knows how to carry out.

Why a character needs more than dialogue
Most conversational NPC systems produce speech and the occasional one-shot action. A training scenario needs a character that holds a position, follows a trainee, breaks contact, remembers what happened last time, and reacts in the same instant a door slams — without a language model trying to steer any of it frame by frame.
The body declares what it can do
The engine has no built-in idea of “move” or “attack”. A body sends a catalogue of actions, each with a JSON Schema, and every call the model makes is validated against that schema before it leaves the engine. A guard in Unreal, a shopkeeper on Discord and a tutor on a call run the same loop with different catalogues.
Actions and behaviours, kept apart
Discrete actions have a lifecycle and an outcome: if one blocks, the turn waits, so the agent learns the door was locked and adapts within the same turn. Persistent behaviours — follow, guard, patrol, flee — run in the game's own AI at frame rate. The model only chooses which one is active.
Reflexes without a round trip
Flinching, turning to face an attacker, raising a weapon: profile reflexes fire immediately with no model call at all. A reflex can stand as the whole response, so the things that must be instant never wait on inference.
Memory that survives the session
Every perception becomes an episodic memory weighted by salience, alongside emotions and per-character relationships that persist across turns and respawns. A character a trainee met last week knows it.
It extends your AI, it does not replace it
The Unreal Engine 5.8 plugin adds server-only components to an existing game. Perception streams out; speech and chosen actions route back through delegates into the systems you already have — AIController, StateTree, Blackboard, montages, interactables.
Attention, so cost is bounded
A turn runs only when something is addressed to the agent or rises above its attention threshold, and ambient turns are rate-limited per character. A crowded scene does not become a crowded bill.
Cognition is an interface
The engine does not depend on any one model vendor. Cognition is a provider interface with implementations for Anthropic's API, any OpenAI-compatible endpoint — OpenAI, Ollama, LM Studio, vLLM, Groq — and an offline mock used for deterministic tests.
For a customer who cannot send character dialogue to a commercial API, that is the whole point: point the provider at a model running on your own hardware and the engine makes no external calls at all. The same build, the same profiles, the same protocol.
Providers can be set per character as well as globally, so a scenario can mix a local model for anything sensitive with a hosted one for the rest.

Where agents can live
The same character, reachable four ways
In the simulation
Unreal pawns over a realtime WebSocket protocol.
In chat
A Discord channel bridged straight to the same agent.
Over REST
A headless body for tooling, tests and after-action review.
As an MCP tool
Talk to, perceive for, and command an agent from an MCP client.
Status
Built and running in Argento's own environment: API-key gated, with character profiles, a knowledge base and the realtime protocol in service. The Unreal plugin targets Engine 5.8. Cognition runs against Anthropic, any OpenAI-compatible endpoint, or a model on your own hardware — we will demonstrate against whichever of those your programme actually has to use.
Put a character in your scenario
Tell us what your trainees need to talk to, and what your simulation already knows how to do. The action catalogue is yours; the brain is ours.
CAGE: 9M1W1 | UEI: FG59XKRYKCW7 | SDVOSB / VOSB