<-Foreword Go to ToC Scope ->

Audio is increasingly produced and consumed as Objects placed in Scenes rather than as fixed channels: each Object carries its sound and what describes it, and a Scene says where each Object is, how the space sounds, and from where it is heard. This makes a presentation adaptable to the user, the device, and the listening position.

Technical Specification: Context-based Audio Enhancement (MPAI-CAE) – Audio Scene Management (CAE-ASM) V1.0 specifies how Audio and Speech Objects are captured, edited, and composed into Objects and Scenes, and how Objects and Scenes are played, so that a User can build, change, and hear an audio presentation, and independent implementations can exchange what they build. Each User action is a User Command.

CAE-ASM operates in the AI Framework (MPAI-AIF) as AI Modules exchanging Data Types of Object and Scene Description (MPAI-OSD).

In all Chapters and Sections, Terms beginning with a capital letter are defined in Table 1. All MPAI-defined Terms are accessible online. All Chapters are Normative unless they are labelled as Informative.

<-Foreword Go to ToC Scope ->