Function
Ref. Model
I/O Data
SubAIMs
JSON MData
Profiles
Ref. Software
Conformance
Performance
1 Functions
The Audio Scene Enhancement (PGM‑ASE) AIM interacts with three external entities through two distinct patterns. A‑User Control issues the Audio CXT Directive and receives the Audio CXT Status reporting the goal understood and the outcome reached. Domain Access and A‑User Storage are queried: the Audio Domain Request and the Audio Interaction History Request are issued by the AIM, and the Audio Domain Response and the Audio Interaction History Response are received in reply, in general during a subsequent processing cycle
Audio Scene Enhancement
- Operates on the Audio Object, under the Audio CXT Directive received from A‑User Control, using the Audio Domain Response received from Domain Access and the Audio Interaction History Response received from A‑User Storage.
- Produces the Enhanced Audio Scene Descriptors and the Audio CXT Status, sent to Context Description Multiplexing, together with the Audio Domain Request sent to Domain Access and the Audio Interaction History Request sent to A‑User Storage.
The Enhanced Audio Scene Descriptors carry the perceptual semantics of the scene, augmented with derived and semantic information under the control of CXT Directive.
2 Reference Model
Figure 1 gives the Reference Model of the Audio Scene Enhancement (PGM‑ASE) AIM.

Figure 1 – Reference Model of the Audio Scene Enhancement (PGM‑ASE) AIM
3 I/O Data
Table 1 gives the Input and Output Data of the Audio Scene Enhancement (PGM‑ASE) AIM.
| Input | Description |
|---|---|
| Audio IH Response | Response obtained from A-User Storage. |
| Audio Object | Audio of the scene captured according to the Audio CXT Directive. |
| Audio CXT Directive | Control directive specifying scope, depth, or policy constraints for audio description. |
| Audio Domain Response | Domain‑specific knowledge supporting audio interpretation and semantic classification. |
| Output | Description |
| Audio IH Request | Request made to A-User Storage. |
| Enhanced Audio Scene Descriptors | Description of the audio scene, carrying perceptual semantics augmented with derived and semantic audio properties. |
| Audio CXT Status | Status information describing the execution and outcome of Audio Scene Enhancement processing. |
| Audio Domain Request | Query to Domain Access requesting domain‑specific knowledge. |
4 Sub-AIMs (informative)
This section is informative. The decomposition of Audio Scene Enhancement into Sub-AIMs described below specifies one conformant architecture that produces the normative outputs of Audio Scene Enhancement (PGM‑ASE). Implementers may adopt alternative internal structures, provided they satisfy the conformance requirements of Section 8. An implementer may develop a Composite PGM-ASE AIM as specified below and claim conformance to it, provided the individual Sub-AIMs conform with the respective specifications.
4.1 Reference Model
An implementation of the Audio Scene Enhancement (PGM‑ASE) AIM may be based on the architecture of Figure 2.

Figure 2 – Reference Model of the Audio Scene Enhancement (PGM‑ASE) Composite AIM
4.2 Operation
The Audio Scene Enhancement operation extracts the Audio Scene Descriptors from the captured Audio Object, providing audio objects and their spatial attitudes; analyses their motion, proximity, and acoustic environment; identifies object types with optional domain knowledge; maps salience with respect to the CXT Directives, Domain Response and Interaction History Response; and constructs the Enhanced Audio Scene Descriptors, together with the execution status (Audio CXT Status), Audio IH Request, and Domain Request.
4.3 Functions of Sub-AIMs
Table 2 specifies the functions performed by the Audio Scene Enhancement (PGM‑ASE) AIM Sub-AIMs in the current example.
| SubAIM | Function |
|---|---|
| Audio Scene Description | Produces an initial Audio Scene Description. |
| Acoustic Environment Analysis | Characterises the acoustic conditions affecting the audio scene, using signal‑derived measures producing the Acoustic Profile. |
| Audio Object Identification | Assigns semantic object‑type labels to audio objects, using classification models and optional domain knowledge. |
| Audio Salience Mapping | Determines the relevance of audio objects with respect to user interaction and context. |
| Enhanced Audio Multiplexing | Aggregates perceptual and enriched evidence into the Enhanced Audio Scene Descriptors and issues the execution status. |
4.4 I/O Data of SubAIMs
Table 3 specifies the Input and Output Data of the Audio Scene Enhancement (PGM‑ASE) AIM SubAIMs.
4.5 AIMs and JSON Metadata
Table 4 provides the links to the AIM specifications and JSON schemas. AIM1 indicates the Composite AIM and AIM2 its SubAIMs.
| AIM1 | AIM2 | Name | JSON |
|---|---|---|---|
| PGM‑ASE | Audio Scene Enhancement | X | |
| PGM-ASD | Audio Scene Description | X | |
| PGM-AEA | Acoustic Environment Analysis | X | |
| PGM-AOI | Audio Object Identification | X | |
| PGM-ASM | Audio Salience Mapping | X | |
| PGM-EAM | Enhanced Audio Multiplexing | X |
5 JSON Metadata
https://schemas.mpai.community/PGM1/V1.0/AIMs/AudioSceneEnhancement.json
6 Profiles
No Profiles.
7 Reference Software
Not part of this specification.
8 Conformance Testing
Not part of this specification.
9 Performance Assessment
Not part of this specification.