Go to CAV-TEC V2.0 AI Modules

Functions
Ref. Model
I/O Data
SubAIMs
JSON MData
Profiles
Ref. Software
Conformance
Performance

1 Functions

The Human-CAV Interaction (CAV-HCI) AIM:

Receives Basic Audio Object From the cabin, through the User Agent.
Point of View From the cabin.
Basic Offline Map Object The Offline Map.
AMS-HCI Message Response From the Autonomous Motion Subsystem.
Ego-Remote HCI Message From Remote CAVs in range.
Produces AMS-HCI Message Request To the Autonomous Motion Subsystem.
Basic Speech Object To the passengers.
Basic Text Object To the passengers.
Recognised Text To the passengers.
Face Descriptors To the passengers.
User ID Outside the CAV.
Ego-Remote HCI Message To Remote CAVs in range.

The AIM is the CAV’s interface with the humans in its cabin. It provides these services:

The dialogue about where to go. It hears the passengers, understands the Destination they ask for – checked against the named places of the Offline Map before anything is sent – and requests it of the Autonomous Motion Subsystem; it proposes the Routes the AMS gives, sends the Route the passenger agrees to, relays Suspend, Resume and Stop, and tells the Route’s state and the arrival.

User identification. It recognises who speaks by the voice and who is in the cabin by the face, and reconciles the two into one User ID.

Personal Status. It extracts the passengers’ Personal Status – from their text, speech, face and gesture – and takes it into account in what it answers.

The Speaking Avatar. It renders what the CAV says as speech, text and a face, with the CAV’s own Personal Status.

Remote CAVs. It exchanges Ego-Remote HCI Messages with the HCIs of the CAVs in range.

Other services – translation of what the passengers say, for instance – may be added with the corresponding AIMs of MPAI-MMC.

2 Reference Model

Figure 1 depicts the Reference Model of the Human-CAV Interaction (CAV-HCI) AIM.

Figure 1 – The Human-CAV Interaction (CAV-HCI) AIM

Figure 1 – The Human-CAV Interaction (CAV-HCI) AIM

3 I/O Data

Table 1 specifies the Input and Output Data of the Human-CAV Interaction (CAV-HCI) AIM.

Table 1 – I/O Data of the Human-CAV Interaction (CAV-HCI) AIM

Input Data Type Description
Basic Audio Object Basic Audio Object What the passengers say.
Point of View (optional) Point of View Where the passenger looks from.
Basic Offline Map Object Basic Offline Map Object The named places the passenger may ask for.
AMS-HCI Message Response AMS-HCI Message Response The Routes proposed, the Route’s state (Route Status), the arrival.
Ego-Remote HCI Message (optional) Ego-Remote HCI Message What another CAV’s HCI sends.
Output Data Type Description
AMS-HCI Message Request AMS-HCI Message Request The Destination, the Route selected, the Route Command.
Basic Speech Object Basic Speech Object What the CAV says.
Basic Text Object Basic Text Object What the CAV says, as text.
Recognised Text (optional) Basic Text Object What the CAV heard.
Face Descriptors (optional) Face Descriptors Object The face of the CAV’s avatar.
User ID (optional) Instance Identifier Who the passenger is.
Ego-Remote HCI Message (optional) Ego-Remote HCI Message What the CAV’s HCI sends.

4 SubAIMs

4.1 Functions of SubAIMs

Table 2 specifies the Functions of the SubAIMs of the Human-CAV Interaction (CAV-HCI) AIM.

Table 2 – Functions of the SubAIMs

SubAIM Function
Basic Audio Scene Description Describes the cabin’s Basic Audio Object as Basic Audio Scene Descriptors.
Basic Visual Scene Description Describes the cabin’s Basic Visual Object as Basic Visual Scene Descriptors.
Basic LiDAR Scene Description Describes the cabin’s LiDAR Object, where the cabin has one.
Audio Qualifier Conversion Converts a Basic Audio Object from one Audio Qualifier to another.
Audio-Visual Alignment Aligns the Identifiers of Speech, Audio and Visual Objects with the same Spatial Attitude.
Audio Scene Object Identification Classifies each audio object of the scene.
Visual Scene Object Identification Tags the faces among the visual objects of the scene.
Automatic Speech Recognition Recognises what the passenger says.
Natural Language Understanding Extracts the meaning of what was said.
Audio Instance Identification Identifies a non-speech audio object.
Visual Instance Identification Identifies a visual object.
Speaker Identity Recognition Identifies the speaker from the speech.
Face Identity Recognition Identifies a face.
Identity Reconciliation Reconciles the Face ID and the Speaker ID into a single User ID.
Personal Status Extraction Extracts the passenger’s Personal Status from text, speech, face and gesture.
Entity Dialogue Processing Holds the dialogue with the passenger and with the AMS: understands the Destination with a language model, checks it against the named places of the Offline Map, proposes the Routes, tells the Route’s state and the arrival.
Response and Scene Rendering Renders the CAV’s answer as a Speaking Avatar: speech, text and face.

4.2 Operation

The cabin’s audio is described and aligned with what is seen; the passenger’s speech is recognised and its meaning extracted; the passenger’s identity and Personal Status are recognised where the cabin sees the passenger.

Entity Dialogue Processing holds the dialogue: it understands the Destination, asks about a place named ambiguously, tells a place the map does not have, and sends the AMS-HCI Message Request; with the AMS-HCI Message Response it proposes the Routes, and tells the Route’s state and the arrival. What it says is rendered by Response and Scene Rendering.

4.3 I/O Data of SubAIMs

Table 3 specifies the Input and Output Data of the SubAIMs.

Table 3 – I/O Data of the SubAIMs

SubAIM Input Output
Basic Audio Scene Description Basic Audio Object – Basic Audio Object
Spatial Attitude – Spatial Attitude
Basic Environment Descriptors – Basic Environment Descriptors
Basic Audio Scene Descriptors – Basic Audio Scene Descriptors
Alert – Alert
Basic Visual Scene Description Basic Visual Object – Basic Visual Object
Spatial Attitude – Spatial Attitude
Basic Environment Descriptors – Basic Environment Descriptors
Basic Visual Scene Descriptors – Basic Visual Scene Descriptors
Alert – Alert
Basic LiDAR Scene Description Basic LiDAR Object – Basic LiDAR Object
Spatial Attitude – Spatial Attitude
Basic Environment Descriptors – Basic Environment Descriptors
Basic LiDAR Scene Descriptors – Basic LiDAR Scene Descriptors
Alert – Alert
Audio Qualifier Conversion Input Audio Object – Basic Audio Object Output Audio Object – Basic Audio Object
Audio-Visual Alignment Speech Scene Descriptors – Basic Speech Scene Descriptors
Audio Scene Descriptors – Basic Audio Scene Descriptors
Visual Scene Descriptors – Basic Visual Scene Descriptors
LiDAR Scene Descriptors – Basic LiDAR Scene Descriptors
Aligned Audio Scene Descriptors – Basic Audio Scene Descriptors
Aligned Visual Scene Descriptors – Basic Visual Scene Descriptors
Basic Audio Scene Geometry – Basic Audio Scene Geometry
Basic Visual Scene Geometry – Basic Visual Scene Geometry
Audio Scene Object Identification Audio Scene Descriptors – Basic Audio Scene Descriptors Basic Speech Object – Basic Speech Object
Audio Object – Basic Audio Object
Visual Scene Object Identification Visual Scene Descriptors – Basic Visual Scene Descriptors Visual Object – Basic Visual Object
Automatic Speech Recognition Language Selector – Selector
Auxiliary Text – Text Object
Speaker ID – Instance Identifier
Speech Overlap – Speech Overlap
Image Speech – Basic Speech Object
Response Speech – Basic Speech Object
Recognised Text – Basic Text Object
Speech Descriptors – Speech Descriptors Object
Text Question – Basic Text Object
Natural Language Understanding Input Text – Basic Text Object
Recognised Text – Basic Text Object
Text Descriptors – Text Descriptors Object
Refined Text – Basic Text Object
Text Personal Status – Text Personal Status
Audio Instance Identification Input Audio Object – Basic Audio Object Audio Object ID – Instance Identifier
Visual Instance Identification Target Visual Object – Basic Visual Object Visual Instance ID – Instance Identifier
Speaker Identity Recognition Auxiliary Text – Text Object
Basic Speech Object – Basic Speech Object
Speech Overlap – Speech Overlap
Speech Scene Geometry – Speech Scene Geometry
Audio Scene Descriptors – Basic Audio Scene Descriptors
Speaker ID – Instance Identifier
Face Identity Recognition Auxiliary Text – Text Object
Visual Object – Basic Visual Object
Visual Scene Geometry – Visual Scene Geometry
Visual Scene Descriptors – Basic Visual Scene Descriptors
Face ID – Instance Identifier
Bounding Box – Bounding Box
Identity Reconciliation Face ID – Instance Identifier
Speaker ID – Instance Identifier
User ID – Instance Identifier
Personal Status – Entity Personal Status
Response – Basic Text Object
Identification – boolean
Personal Status Extraction Text Object – Text Object
Text Descriptors – Text Descriptors
Basic Speech Object – Basic Speech Object
Speech Descriptors – Speech Descriptors Object
Face Object – Basic Visual Object
Face Descriptors – Face Descriptors Object
Body Object – Visual Object
Gesture Descriptors – Gesture Descriptors Object
Text Personal Status – Text Personal Status
Personal Status – Entity Personal Status
Entity Dialogue Processing Summary – Summary
Text Object – Basic Text Object
Text Descriptors – Text Descriptors Object
Personal Status – Entity Personal Status
User ID – Instance Identifier
Visual Object IDs – Instance Identifier
Audio Object IDs – Instance Identifier
Visual Scene Geometry – Basic Visual Scene Geometry
Audio Scene Geometry – Basic Audio Scene Geometry
Audio-Visual Scene Descriptors – Basic Audio-Visual Scene Descriptors
Memory Text – Basic Text Object
Memory Personal Status – Entity Personal Status
AMS-HCI Message Response – AMS-HCI Message
Basic Offline Map Object – Basic Offline Map Object
Ego-Remote HCI Message – Ego-Remote HCI Message
Machine Text Object – Basic Text Object
Machine Personal Status – Entity Personal Status
Edited Summary – Summary
HCI Memory Data – HCI Memory Data
AMS-HCI Message Request – AMS-HCI Message
Ego-Remote HCI Message – Ego-Remote HCI Message
Response and Scene Rendering Text Object – Basic Text Object
Text Object – Basic Text Object
Text Object – Basic Text Object
Personal Status – Entity Personal Status
Full Environment Descriptors – Basic Audio-Visual Scene Descriptors
Point of View – Object Point of View
Machine Speech – Basic Speech Object
Machine Face Descriptors – Face Descriptors Object
Machine Body Descriptors – Body Descriptors Object

4.4 AIMs and JSON Metadata

Table 4 gives the AIMs and their JSON Metadata.

Table 4 – AIMs and JSON Metadata

AIM Name JSON
CAV-HCI Human-CAV Interaction HumanCAVInteraction.json
OSD-BAS Basic Audio Scene Description BasicAudioSceneDescription.json
OSD-BVS Basic Visual Scene Description BasicVisualSceneDescription.json
OSD-BLS Basic LiDAR Scene Description BasicLiDARSceneDescription.json
CAE-QCV Audio Qualifier Conversion AudioQualifierConversion.json
OSD-AVA Audio-Visual Alignment AudioVisualAlignment.json
CAE-ASI Audio Scene Object Identification AudioSceneObjectIdentification.json
CVE-VSI Visual Scene Object Identification VisualSceneObjectIdentification.json
MMC-ASR Automatic Speech Recognition AutomaticSpeechRecognition.json
MMC-NLU Natural Language Understanding NaturalLanguageUnderstanding.json
CAE-AII Audio Instance Identification AudioInstanceIdentification.json
CVE-VII Visual Instance Identification VisualInstanceIdentification.json
MMC-SIR Speaker Identity Recognition SpeakerIdentityRecognition.json
PAF-FIR Face Identity Recognition FaceIdentityRecognition.json
OSD-IDR Identity Reconciliation IDReconciliation.json
MMC-PSE Personal Status Extraction PersonalStatusExtraction.json
MMC-EDP Entity Dialogue Processing EntityDialogueProcessing.json
PAF-RSR Response and Scene Rendering ResponseAndSceneRendering.json

5 JSON Metadata

https://schemas.mpai.community/CAV2/V2.0/AIMs/HumanCAVInteraction.json

6 Profiles

No Profiles.

7 Reference Software

The reference software of Human-CAV Interaction is part of the reference software of CAV-TEC V2.0 (see N4195, An introduction to the reference software): C# on .NET 10, released under the BSD 3-Clause licence, obtained from MPAI.

Its provider builds every Sub-AIM of the L3 of Human-CAV Interaction, its composites (Personal Status Extraction, Response and Scene Rendering) included; the Controller builds each one when the Module starts. The engines: Automatic Speech Recognition with whisper.cpp (model ggml-small); Entity Dialogue Processing with a language model (Llama 3.2, 3B parameters, served by Ollama) that only understands – a Destination, yes, a choice of Route, no, suspend, resume, stop – everything it understands checked against the named places of the Offline Map before anything is sent, and what the CAV says composed from what the AMS gave, not by the model; Response and Scene Rendering speaking with Piper. The state of the dialogue is kept in the Private Storage of Entity Dialogue Processing.

Implemented and tested: the speech path – Audio Qualifier Conversion, the audio scene, Automatic Speech Recognition, Natural Language Understanding, Entity Dialogue Processing, Response and Scene Rendering. The vision and identity AIMs are built and receive nothing where the cabin has no camera. A dialogue turn takes about 3.4 s (median) on a PC without GPU.

8 Conformance Testing

Table 5 provides the Conformance Testing Method for the Human-CAV Interaction (CAV-HCI) AIM: the data it receives and produces. If a schema references other schemas, conformance of data for the primary schema implies that data referencing a secondary schema shall also validate against that schema, if present, and conform with the Qualifier, if present.

Table 5 – Conformance Testing Method for the Human-CAV Interaction (CAV-HCI) AIM

Receives Basic Audio Object Shall validate against the Basic Audio Object schema.
Point of View Shall validate against the Point of View schema.
Basic Offline Map Object Shall validate against the Basic Offline Map Object schema.
AMS-HCI Message Response Shall validate against the AMS-HCI Message schema.
Ego-Remote HCI Message Shall validate against the Ego-Remote HCI Message schema.
Produces AMS-HCI Message Request Shall validate against the AMS-HCI Message schema.
Basic Speech Object Shall validate against the Basic Speech Object schema.
Basic Text Object Shall validate against the Basic Text Object schema.
Basic Text Object Shall validate against the Basic Text Object schema.
Face Descriptors Object Shall validate against the Face Descriptors Object schema.
Instance Identifier Shall validate against the Instance Identifier schema.
Ego-Remote HCI Message Shall validate against the Ego-Remote HCI Message schema.

9 Performance Assessment