The Human and Machine Communication standard

The Human and Machine Communication (MPAI-HMC) standard addresses a communication problem that substantially extends the ordinary video paradigm by including autonomous agents in a conversation without sharing a space, a language, or even a physical form. The standard specifies a single AI Module — Communicating Entities in Context (HMC-CEC) — built from standardised components that work together in a clear and interoperable way.

Entities and Contexts

What communicates is an Entity. An Entity may be a human in a real audio-visual scene, a Digital Human representing either a human or a machine in a virtual one, or a machine with no visible form at all. Each Entity carries a Context — its language, its culture, and the other attributes that shape how what it says should be understood.

The point of the distinction is that the same conversation may span different Entities in different Contexts. A human speaking in a room can be understood by a machine that answers as a speaking avatar, and what one Entity produces can be converted so that its meaning fits the Context of another.

Main functions of the workflow

The Reference Model composes seven AI Modules, three of which are themselves composed of further AIMs:

  • Capturing the scene – builds a digital description of the real audio-visual scene around the Entity, or integrates a received avatar into a virtual scene.
  • Understanding the Entity and its Context – identifies who or what is communicating, what was said, what it means, and the Personal Status behind it.
  • Translating – converts text and speech between languages while preserving what was meant.
  • Responding – produces the machine’s reply as text and a Personal Status congruent with what was received.
  • Rendering – displays the response as a speaking Virtual Human in an audio-visual scene.

One item carries the whole message

An Entity communicates by emitting a Communication Item — an implementation of the Portable Avatar Data Type, which carries an avatar together with its Context. Because the item is self-contained and standardised, a receiver can render the sender’s avatar as the sender intended, without prior agreement between the sender and receiver implementations.

Open and flexible design

An AI Module is specified only by its function and its interfaces. Implementers choose their own technologies freely, and may split one module into several or combine several into one, provided that the combined function and interfaces still conform. Every input and output validates against a published JSON Schema and its Qualifier, which is what makes a Module from one supplier substitutable for another’s.

Benefits

The MPAI-HMC standard helps:

  • Researchers and Developers to select individual components — speech recognition, translation, avatar rendering — with a fixed interface.
  • Component manufacturers to bring standard-conforming modules to a market rather than just to a single customer.
  • System builders to assemble conversational systems from components of known operation.
  • Regulators to establish procedures and assess conformance for systems that represent humans and machines to each other.
  • Users to communicate across languages and forms of presence with systems whose operation can be understood and explained.

Ultimately, MPAI-HMC promotes a competitive market of interoperable components for human-machine communication leveraging continuously improving technologies.