<-Introduction Go to ToC Definitions ->
Technical Specification: Context-based Audio Enhancement (MPAI-CAE) – Audio Scene Management (CAE-ASM) V1.0 – in the following also called CAE-ASM V1.0, or simply CAE-ASM.
CAE-ASM specifies:
- the capture of Basic Audio Objects;
- the editing of Basic Audio Objects;
- the composition of Basic Audio Objects and Audio Objects into Audio Objects;
- the playing of Basic Audio Objects and Audio Objects;
- the composition of Basic Audio Scene Descriptors and Audio Scene Descriptors into Audio Scene Descriptors, and of Basic Audio Objects into Basic Audio Scene Descriptors;
- the playing of Audio Objects and Audio Scene Descriptors.
The same functions apply to Speech Objects and Speech Scenes. Speech and audio heard together are composed in Multimodal Objects and Multimodal Scene Descriptors (MSD).
CAE-ASM specifies four AI Modules (Audio Object Acquisition, Audio Object Editing, Audio Scene Editing, Audio Object Delivery), the Composite AIM Audio Scene Management, and the User Command.
CAE-ASM V1.0 relies on the following MPAI standards:
- Technical Specification: Artificial Intelligence Framework (MPAI-AIF) V3.0 specifying the AI Framework (AIF) where Audio and Speech Data are processed by AI Modules (AIM).
- Technical Specification: Process Instance Trust Framework (MPAI-PTF) V1.0 specifying the data structures, processes, and protocols enabling AIMs to establish, evaluate, and maintain trust in the AIF.
- Technical Specification: Object and Scene Description (MPAI-OSD) V1.5 specifying the Audio, Speech, and Multimodal Objects and Scene Descriptors.
- Technical Specification: Data Types, Formats, and Qualifiers (MPAI-TFA) V1.5 specifying the Audio and Speech Qualifiers enabling CAE-ASM to specify Data Types that are independent of specific formats.
CAE-ASM was developed by the CAE-DC Development Committee. MPAI may develop CAE-ASM extensions or new Technical Specifications in the CAE-ASM work area.