Go to PGM-AUA V1.0 AI Modules

Function
Ref. Model
I/O Data
SubAIMs
JSON MData
Profiles
Ref. Software
Conformance
Performance

1 Functions

The Cross‑Modal Interpretation (PGM‑CMI) AI Module integrates the evidence derived separately from the words, the delivery, and the conduct of an Entity. It composes a single reading supported by that evidence, taken together. It does so by:

  • Resolving the references that no single modality resolves alone.
  • Weighing concurring and conflicting evidence according to the Cross‑Modal Interpretation AIM’s confidence and to the quality of the observation that produced it.
  • Integrating the evidence over the capture interval, separating momentary effects from persistent ones.
  • Characterising the disagreements between modalities, distinguishing those that are meaningful from those that arise from imperfect observation.
  • Producing the Cross‑Modal Interpretative Evidence.

The module is integrative. It adjudicates between the sources of evidence it receives and states what that evidence supports. The AIM should not derive new modality‑specific evidence, since this is the function of Linguistic‑Paralinguistic Analysis and Behavioural and Expressive Analysis. Moreover, it should not construct a state of the Entity, since this is the function of Entity State Creation.

Where the evidence supports more than one reading, Cross‑Modal Interpretation produces the readings it supports, ordered by the confidence associated with each. It does not discard the discounted readings and does not resolve all disagreements: an unresolved disagreement is reported, so that Entity State Creation and the downstream AI Modules become aware of the disagreement.

2 Reference Model

Figure 1 depicts the Reference Model of the Cross‑Modal Interpretation (PGM‑CMI) AIM.

Figure 1 – The Cross‑Modal Interpretation (PGM‑CMI) AIM

The Cross‑Modal Interpretation (PGM‑CMI) AI Module operates as follows:

  1. Receives the Harmonised Multimodal Context, the Linguistic‑Paralinguistic Evidence, and the Behavioural and Expressive Indicators resulting from the upstream Sub-AIMs.
  2. Resolves cross‑modal reference by:
    • Binding the referring expressions of an utterance to the Objects and Entities indicated by deictic gestures and gaze.
    • Binding the utterance to the Entity it addresses, based upon the direction of gaze and posture.
    • Reporting the referring expressions that remain unresolved.
  3. Weighs the evidence according to:
    • The confidence carried by each item.
    • The quality of the observation that produced it, including the visibility of the Entity and the acoustic conditions of the capture.
    • The corroboration received by the item from other modalities.
  4. Integrates the evidence over the capture interval by:
    • Distinguishing the momentary from the sustained.
    • Identifying the point at which the evidence changes.
    • Relating the evidence of the current capture to that of the preceding captures.
  5. Characterises each disagreement between modalities as:
    • Meaningful, in cases where the divergence is itself evidence, for example, when the delivery or the conduct qualifies, softens, or contradicts the words.
    • Apparent, in cases where the divergence arises from imperfect observation and the weaker evidence is discounted accordingly.
    • Unresolved, in cases where the evidence does not resolve the disagreement.
  6. Constructs the Cross‑Modal Interpretative Evidence, comprising the readings the evidence supports, their order of confidence, the established reference bindings, and the identified disagreements.
  7. Outputs the Cross‑Modal Interpretative Evidence for use by downstream AI Modules.

The AIM operates in real time, without recourse to the reasoning capability of Basic Knowledge and without Model Context Protocol interactions. Domain knowledge is obtained by query, and the Response to a query is normally received during a subsequent processing cycle.

The Domain Response supports the reading of evidence in the terms of the domain: the expected conduct of the contained Entities, and the conventional combinations of words and conduct which would elsewhere be read otherwise.

3 I/O Data

Table 1 specifies the Input and Output Data of the Cross‑Modal Interpretation (PGM‑CMI) AIM.

Table 1 – I/O Data of the Cross‑Modal Interpretation (PGM‑CMI) AIM

Input Description
Harmonised Multimodal Context Audio, visual, textual, and paralinguistic evidence of the scene, placed on a common time base and attributed to harmonised Entities.
Linguistic‑Paralinguistic Evidence Evidence carried by the content of the utterances of an Entity and the manner of their delivery.
Behavioural and Expressive Indicators Indicators carried by the conduct of an Entity and the expression of its body.
User‑USD Directive Control directive specifying scope, depth, or policy constraints for cross‑modal interpretation.
User Domain Response Domain‑specific knowledge supporting the reading of evidence in the terms of the domain.
User Interaction History Response Prior interaction‑history content supporting the integration of the evidence of the current capture with that of the preceding captures.
Output Description
Cross‑Modal Interpretative Evidence Readings supported by the evidence of all modalities, in order of confidence, with the established  reference bindings and the identified disagreements.
User‑USD Status Status information describing the execution and outcome of Cross‑Modal Interpretation processing.
User Domain Request Query to Domain Access for domain‑specific knowledge.
User Interaction History Request Query to A‑User Storage for prior interaction‑history content.

4 SubAIMs

No SubAIMs.

5 JSON Metadata

https://schemas.mpai.community/PGM1/V1.0/AIMs/CrossModalInterpretation.json

6 Profiles

No Profiles.

7 Reference Software

Not part of this specification.

8 Conformance Testing

Table 2 provides the Conformance Testing Method for the Cross‑Modal Interpretation (PGM‑CMI) AIM.

If a schema contains references to other schemas, conformance of data for the primary schema implies that any data referencing a secondary schema shall also validate against the relevant schema, if present, and conform with the Qualifier, if present.

Table 2 – Conformance Testing Method for the Cross‑Modal Interpretation (PGM‑CMI) AIM

Receives Harmonised Multimodal Context Shall validate against Harmonised Multimodal Context schema.
Linguistic‑Paralinguistic Evidence Shall validate against Linguistic‑Paralinguistic Evidence schema.
Behavioural and Expressive Indicators Shall validate against Behavioural and Expressive Indicators schema.
User‑USD Directive Shall validate against User‑USD Directive schema.
User Domain Response Shall validate against User Domain schema.
User Interaction History Response Shall validate against Interaction History schema.
Produces Cross‑Modal Interpretative Evidence Shall validate against Cross‑Modal Interpretative Evidence schema.
User‑USD Status Shall validate against User‑USD Status schema.
User Domain Request Shall validate against User Domain schema.
User Interaction History Request Shall validate against Interaction History schema.

9 Performance Assessment

Not part of this specification.

Go to PGM-AUA V1.0 AI Modules