Definition Functional Requirements Syntax Semantics

Definition

The Linguistic–Paralinguistic Evidence (LPE) Data Type:

  • Is produced by the Linguistic and Paralinguistic Analysis (LPA) SubAIM of the User State Description (USD) AIM.
  • Represents structured, temporally anchored evidence derived from recognised speech and associated audio descriptors.
  • Encodes what is said (linguistic content) and how it is said (observable paralinguistic cues), without performing semantic, affective, or intentional inference.
  • Serves as an evidential input for downstream cross‑modal interpretation and Entity State construction.

Functional Requirements

LPE conveys the following main information elements:

Function Description
Linguistic Structuring Encodes recognised speech into structured linguistic units, preserving explicit grammatical and communicative form.
Paralinguistic Observation Captures observable speech‑related cues (e.g. loudness, speech rate, pauses) derived from audio descriptors.
Temporal Grounding Associates all linguistic and paralinguistic elements with precise temporal anchors.
Evidence Separation Maintains a strict separation between observable evidence and inferred mental or affective states.
Modal Correlation Enables correlation between speech content and audio objects through explicit references.
Interpretation Neutrality Avoids encoding emotion, intention, belief, or motivation.
Reasoning Substrate Provides structured evidence for subsequent cross‑modal interpretation stages within ESD.
Auditability Includes Data Exchange Metadata to support provenance, traceability, and confidence assessment.

Syntax

Semantics

The table below specifies the semantics of all top‑level and nested keys defined in the Linguistic–Paralinguistic Evidence schema.

Label Description
Header LPE header identifying the data type and version, formatted as MMC-LPE-Vx.y.
MInstanceID Identifier of the M‑Instance producing the LPE data.
LPEID Unique identifier of the Linguistic–Paralinguistic Evidence instance.
EvidenceTime Time reference indicating the temporal scope of the evidence set.
LinguisticUnits Collection of structured linguistic units derived from recognised speech.
– TextOrTextID Either the recognised text string or a TextSegmentID referencing a text segment produced by an ASR AIM.
– UtteranceType Grammatical or communicative form of the utterance (e.g. Declarative, Interrogative, Imperative).
– TemporalAnchor Temporal anchor indicating when the linguistic unit was uttered.
ParalinguisticCues Collection of observable speech‑related cues derived from audio descriptors.
– AudioRef Audio Object or AudioObjectID identifying the audio source associated with the cue.
– CueType Type of observed paralinguistic cue (e.g. Loudness, SpeechRate, Pause, Emphasis).
– CueValue Structured representation of the observed cue value, including a numerical value and the scale used for its interpretation.
  – Value
Scale of the cue value (e.g. Normalized, Relative, Absolute).
  – Scale Scale of the cue value (e.g. Normalized, Relative, Absolute).
  – TemporalAnchor Temporal anchor indicating when the paralinguistic cue was observed.
DataXMData Data Exchange Metadata providing provenance, source AIM identification, confidence, legality, and rights information.
DescrMetadata Human‑readable descriptive metadata associated with the LPE instance.