Go to PGM-AUA V1.0 AI Modules

Function
Ref. Model
I/O Data
SubAIMs
JSON MData
Profiles
Ref. Software
Conformance
Performance

1 Functions

The Cross‑Modal Interpretation (PGM‑CMI) AI Module integrates the evidence derived separately from the words, the delivery, and the conduct of an Entity into a single reading of what that evidence, taken together, supports, by

  • Resolving the references that no single modality resolves alone
  • Weighing concurring and conflicting evidence according to its confidence and to the quality of the observation that produced it
  • Integrating the evidence over the capture interval, separating what is momentary from what persists
  • Characterising the disagreements between modalities, distinguishing those that are meaningful from those that arise from imperfect observation
  • Producing the Cross‑Modal Interpretative Evidence

The module is integrative. It adjudicates between the evidence it receives and states what that evidence supports; it SHALL NOT derive new modality‑specific evidence, which is the function of Linguistic‑Paralinguistic Analysis and Behavioural and Expressive Analysis, and it SHALL NOT construct a state of the Entity, which is the function of Entity State Construction.

Where the evidence supports more than one reading, Cross‑Modal Interpretation produces the readings it supports, ordered by the confidence attaching to each. It does not discard the readings it does not favour, and it does not resolve a disagreement it cannot resolve: an unresolved disagreement is itself reported, so that Entity State Construction and the AI Modules downstream of it know that the evidence did not agree.

2 Reference Model

Figure 1 depicts the Reference Model of the Cross‑Modal Interpretation (PGM‑CMI) AIM.

Figure 1 – The Cross‑Modal Interpretation (PGM‑CMI) AIM

The Cross‑Modal Interpretation (PGM‑CMI) AI Module operates as follows:

  1. Receives the Harmonised Multimodal Context, the Linguistic‑Paralinguistic Evidence, and the Behavioural and Expressive Indicators
  2. Resolves cross‑modal reference by:
    • Binding the referring expressions of an utterance to the Objects and Entities indicated by deictic gesture and by gaze
    • Binding the utterance to the Entity it addresses, from the direction of gaze and of posture
    • Reporting the referring expressions that remain unresolved
  3. Weighs the evidence by:
    • The confidence carried by each item
    • The quality of the observation that produced it, including the visibility of the Entity and the acoustic conditions of the capture
    • The corroboration the item receives from the other modalities
  4. Integrates the evidence over the capture interval by:
    • Distinguishing the momentary from the sustained
    • Identifying the point at which the evidence changes
    • Relating the evidence of the current capture to that of the captures preceding it
  5. Characterises each disagreement between modalities as:
    • Meaningful, i.e., the divergence is itself evidence, as when the delivery or the conduct qualifies, softens, or contradicts what the words assert
    • Apparent, i.e., the divergence arises from imperfect observation, and the weaker evidence is discounted accordingly
    • Unresolved, i.e., the evidence does not settle which of the two holds
  6. Constructs the Cross‑Modal Interpretative Evidence, comprising the readings the evidence supports, their order of confidence, the reference bindings established, and the disagreements as characterised
  7. Outputs the Cross‑Modal Interpretative Evidence for use by downstream AI Modules

The module operates in real time, without recourse to the reasoning capability of Basic Knowledge and without Model Context Protocol interactions. Domain knowledge is obtained by query, and the Response to a query is in general received during a subsequent processing cycle.

The Domain Response supports the reading of evidence in the terms of the domain: the conduct the setting expects of the Entities in it, and the combinations of word and conduct that are conventional in it and that would be read otherwise elsewhere.

3 I/O Data

Table 1 specifies the Input and Output Data of the Cross‑Modal Interpretation (PGM‑CMI) AIM.

Table 1 – I/O Data of the Cross‑Modal Interpretation (PGM‑CMI) AIM

Input Description
Harmonised Multimodal Context Audio, visual, textual, and paralinguistic evidence of the scene placed on a common time base and attributed to harmonised Entities.
Linguistic‑Paralinguistic Evidence Evidence borne by the content of the utterances of an Entity and by the manner of their delivery.
Behavioural and Expressive Indicators Indicators borne by the conduct of an Entity and by the expression of its body.
User‑USD Directive Control directive specifying scope, depth, or policy constraints for cross‑modal interpretation.
User Domain Response Domain‑specific knowledge supporting the reading of evidence in the terms of the domain.
User Interaction History Response Prior interaction‑history content supporting the integration of the evidence of the current capture with that of the captures preceding it.
Output Description
Cross‑Modal Interpretative Evidence Readings the evidence of all modalities supports, in order of confidence, with the reference bindings established and the disagreements as characterised.
User‑USD Status Status information describing the execution and outcome of Cross‑Modal Interpretation processing.
User Domain Request Query to Domain Access for domain‑specific knowledge.
User Interaction History Request Query to A‑User Storage for prior interaction‑history content.

4 SubAIMs

No SubAIMs.

5 JSON Metadata

https://schemas.mpai.community/PGM1/V1.0/AIMs/CrossModalInterpretation.json

6 Profiles

No Profiles.

7 Reference Software

Not part of this specification.

8 Conformance Testing

Table 2 provides the Conformance Testing Method for the Cross‑Modal Interpretation (PGM‑CMI) AIM.

If a schema contains references to other schemas, conformance of data for the primary schema implies that any data referencing a secondary schema shall also validate against the relevant schema, if present, and conform with the Qualifier, if present.

Table 2 – Conformance Testing Method for the Cross‑Modal Interpretation (PGM‑CMI) AIM

Receives Harmonised Multimodal Context Shall validate against Harmonised Multimodal Context schema.
Linguistic‑Paralinguistic Evidence Shall validate against Linguistic‑Paralinguistic Evidence schema.
Behavioural and Expressive Indicators Shall validate against Behavioural and Expressive Indicators schema.
User‑USD Directive Shall validate against User‑USD Directive schema.
User Domain Response Shall validate against User Domain schema.
User Interaction History Response Shall validate against Interaction History schema.
Produces Cross‑Modal Interpretative Evidence Shall validate against Cross‑Modal Interpretative Evidence schema.
User‑USD Status Shall validate against User‑USD Status schema.
User Domain Request Shall validate against User Domain schema.
User Interaction History Request Shall validate against Interaction History schema.

9 Performance Assessment

Not part of this specification.

Go to PGM-AUA V1.0 AI Modules