The Autonomous User Architecture project
Pursuing Goals in the Metaverse is a container project. The first is the Autonomous User Architecture (PGM-AUA) project specifying the architecture, functions and interfaces of an Autonomous User (A-User). An A-User is a Process that operates in a metaverse instance (M-Instance) under the responsibility of a human and interacts, through a Speaking Avatar, with another User. That User may be another A-User or a Human-User under the direct control of a human, in the same or in another M-Instance.
The standard enables an A-User to perceive the Space and the User it interacts with, understand the User’s condition, decide how to respond, and then act in the M-Instance, with a degree of autonomy set by the responsible human.
The Reference Architecture
The A-User is an instruction-driven system of interacting AI Modules (AIMs), organised in a Module and orchestrated by a central controller.
The system includes:
- A-User Control (AUC), which governs the A-User’s operation and is its sole interface with the M-Instance.
- Context Description (CXC), the front-end that describes the audio and visual Space and derives a first-pass User State.
- Domain Access (DAC), which supplies domain semantics: object classes, relations, constraints and affordances.
- Prompt Creation (PRC), which assembles the A-User’s understanding of the situation into a structured PC-Prompt.
- Basic Knowledge (BKN), the reasoning core that determines the A-User’s communicative behaviour and produces the Final Response.
- User State Refinement (USR), which produces a stable and coherent User State.
- Personality Alignment (PAL), the sole producer of the A-User State, aligning the A-User’s Personality with the User State.
- A-User Formation (AUF), which renders the speaking avatar.
- A-User Storage (AUS), which keeps the authorised Interaction History.
Each AIM has a defined role, so AIMs can be developed and replaced independently while the A-User remains interoperable.
Main Functions of the Autonomous User
The MPAI-PGM-AUA standard enables the following essential functions:
- Perception – Captures text, audio, visual and 3D information from the User and the surrounding Space.
- Scene Description – Produces aligned Audio and Visual Scene Descriptors, including object identity, depth and affordances.
- User Understanding – Derives the User State, a structured description of the User’s observable cognitive, emotional and interactional condition.
- Deliberation – Determines the stance to take towards the Space and the User, and the response to give.
- Personality Alignment – Derives an A-User State consistent with the A-User’s Personality and congruent with the User State.
- Rendering – Expresses the response through the speech, facial expression, gaze and gesture of a Speaking Avatar.
- Action – Performs Actions and Process Actions in the M-Instance, as specified by MPAI-MMM.
- Escalation – Refers intermediate results or unresolvable conflicts to the responsible human.
These functions allow an A-User to take part in a conversation as a coherent, goal-directed participant rather than a scripted responder.
Operation Model
PGM-AUA separates two interaction paradigms and two modes of execution.
- Operational Interfaces carry stateless, deterministic exchanges, such as Directives and Status messages or single-shot domain queries.
- Semantic Interfaces, implemented with the Model Context Protocol (MCP), carry session-level exchanges in which meaning is established and refined.
- Reactive execution gives low-latency responses from already available descriptors and states, without MCP interactions.
- Deliberative execution uses MCP-based reasoning to produce the Final Response and the updated A-User State.
A-User Control steers the AIMs:
- It sets goals and constraints through Directives, organised in eight Instruction Types from perception to avatar rendering.
- Each AIM reports what it understood, what it achieved and why, in a Status.
- A-User Control may then issue a revised Directive.
This model keeps the A-User responsive in real time while its deliberation proceeds.
Powered by the MPAI AI Framework
PGM-AUA relies on the MPAI AI Framework (MPAI-AIF V3.0), which provides:
- A modular and interoperable execution environment for AIMs.
- Configuration and orchestration of AIMs in a Module.
- Platform-independent implementation.
The A-User operates in M-Instances specified by MPAI Metaverse Model – Technologies (MMM-TEC).
Open and Flexible Design
PGM-AUA follows key design principles:
- Goal authority is separated from reasoning. A-User Control holds the goals but does not interpret or deliberate; the specialised AIMs do.
- A single, accountable interface to the metaverse. Only A-User Control holds the A-User’s Rights and Wallet and acts in the M-Instance.
- Perception is not rewritten by deliberation. Scene Descriptors are not modified by reasoning.
- Governed memory. A-User Control authorises what is stored, by whom, for how long, and who may read it.
- Human oversight. The responsible human can start, stop, authorise, deny and adjust the A-User’s autonomy.
Benefits
The PGM-AUA standard will support:
- Metaverse operators to populate M-Instances with autonomous Users made of interoperable components.
- Developers to build and improve individual AIMs without redesigning the whole A-User.
- Integrators to compose A-Users from AIMs of different providers.
- Responsible humans to delegate tasks to an A-User while keeping control and accountability.
- Users to interact naturally with autonomous agents that understand and respond to them.
A New Paradigm for Autonomous Agents
PGM-AUA promotes a shift from monolithic conversational agents to:
- Autonomous agents built from standard, replaceable components.
- Agents whose reasoning, memory and actions are governed and traceable.
- Agents that combine real-time responsiveness with deliberative reasoning.
- Agents whose autonomy is set and supervised by a responsible human.
Conclusion
PGM-AUA V1.0 provides an interoperable architecture for Autonomous Users that:
- Perceive and understand the Space and the User they interact with.
- Deliberate and respond coherently with a defined Personality.
- Act in the metaverse within their Rights and the M-Instance Rules.
- Remain under the governance of a responsible human.
The standard enables a new generation of autonomous agents built from reusable AI components.