Go to PGM-AUA V1.0 AI Modules

Function Ref. Model I/O Data SubAIMs JSON MData Profiles Ref. Software Conformance Performance

1. Function

The Visual Scene Enhancement (PGM‑VSE) AIM derives the Visual Scene Descriptors from the Visual Object captured according to the Visual CXT Directive and refines them using contextual directives, domain knowledge, and interaction history, producing the Enhanced Visual Scene Descriptors. The derivation of the initial Visual Scene Descriptors is performed by the Visual Scene Description (PGM‑VSD) Sub‑AIM, which Visual Scene Enhancement contains.

Visual Scene Enhancement:

  1. Operates on the Visual Object, the Visual CXT Directive received from A‑User Control, the Visual Domain Response resulting from queries made to Domain Access, and the Visual Interaction History Response resulting from queries made to A-User Storage.
  2. Identifies and classifies Visual Objects in the scene, using contextual and domain directives.
  3. Infers Affordance Tags and Interaction Potential for identified objects.
  4. Produces a Salient Object Matrix ranking objects by perceptual importance and interaction relevance.
  5. Produces the Enhanced Visual Scene Descriptors sent to Audio‑Visual Alignment and Visual Enhancement Multiplexing, the Visual CXT Status, the Visual Domain Request when querying Domain Access, and the Visual Interaction History Request when querying A-User Storage.

Affordances and environmental characterisation are inferred from the visual modality alone. The possible actions an object offers derive from its observable geometry, material properties, and spatial relations, which the audio modality does not convey; Audio Scene Enhancement therefore has no affordance counterpart, and its Acoustic Environment Analysis Sub‑AIM characterises the acoustic environment only.

The Enhanced Visual Scene Descriptors carry the perceptual semantics of the scene, augmented with identified objects, affordances, and salience rankings, under Visual CXT Directive control.

2. Reference Model

Figure 1 depicts the Reference Model of the Visual Scene Enhancement (PGM‑VSE) AIM.

Figure 1 – Reference Model of the Visual Scene Enhancement (PGM‑VSE) AIM

3. Input/Output Data

Table 1 lists the Input and Output Data of the Visual Scene Enhancement (PGM‑VSE) AIM.

Table 1 – Input/Output Data of the Visual Scene Enhancement (PGM‑VSE) AIM

Input Description
Visual Object Input Visual Object to the PGM-VSE AIM.
Visual CXT Directive Contextual directive specifying scope, depth, and policy constraints for visual enhancement.
Visual Domain Response Domain‑specific knowledge supporting visual object classification and semantic interpretation.
Visual IH Response Interaction History response, providing temporal context for scene understanding.
Output Description
Enhanced Visual Scene Descriptors Refined Visual Scene Descriptors integrating identified objects, affordances, and salience rankings.
Visual CXT Status Multiplexed status signals from Visual Object Identification, Affordance Inference, and Visual Salience Mapping, reporting their processing outcomes.
Visual Domain Request Multiplexed domain requests from Visual Object Identification, Affordance Inference, and Visual Salience Mapping.
Visual IH Request Multiplexed Interaction History requests from Visual Object Identification, Affordance Inference, and Visual Salience Mapping.

4. Sub-AIMs (Informative)

This section is informative. The decomposition into Sub-AIMs described below specifies one conformant architecture that produces the normative outputs of the Visual Scene Enhancement (PGM‑VSE) AIM. Implementations may adopt alternative internal structures, provided they satisfy the conformance requirements of Section 8. An implementer may develop a Composite Visual Scene Enhancement AIM as specified below and claim conformance to that Visual Scene Enhancement AIM, provided the individual Sub-AIMs conform with the respective specifications.

4.1 Reference Model

Figure 2 depicts the Reference Model of the Visual Scene Enhancement (PGM‑VSE) Composite AIM.

Figure 2 – Reference Model of the Visual Scene Enhancement (PGM‑VSE) Composite AIM

In general, Salience is based on the Context as represented by Visual CXT Directive (priorities set by A-User); Interaction History as derived from A-User Storage; and Domain as derived from querying Domain Access.

4.2 Operation

The Visual Scene Enhancement operation progressively enriches the initial Visual Scene Descriptors instance through a sequence of specialised Sub-AIMs:

  1. The Visual Scene Description Sub-AIM constructs the initial scene representation.
  2. Visual Object Identification assigns semantic labels to detected objects.
  3. Visual Affordance Inference determines the possible interactions with objects and their feasibility.
  4. Visual Salience Mapping ranks objects by perceptual importance and interaction relevance and produces the Salient Object Matrix.
  5. The Visual Enhancement Multiplexing Sub-AIM collects the outputs of all Sub-AIMs and assembles the Enhanced Visual Scene Descriptors together with the multiplexed status and request signals.

Each of these Sub-AIMs –  Visual Scene Description, Visual Object Identification, Affordance Inference, and Visual Salience Mapping – independently generates a Visual CXT Status, a Visual Domain Request, and a Visual Interaction History Request. These are multiplexed by Visual Enhancement Multiplexing into the corresponding external output signals.

4.3 Functions of Sub-AIMs

Table 2 specifies the functions of the Sub-AIMs of the Visual Scene Enhancement (PGM‑VSE) Composite AIM.

Table 2 – Functions of the Sub-AIMs of the Visual Scene Enhancement (PGM‑VSE) Composite AIM

Name Function
Visual Scene Description Receives the external inputs, constructs the initial Visual Scene Descriptors from the best‑effort VSD input.
Visual Object Identification Assigns semantic object‑type labels to Visual Objects using classification models, domain knowledge, and Interaction History.
Affordance Inference Infers Affordance Tags and Interaction Potential describing the possible interactions with Visual Objects and their feasibility, combining object properties, spatial constraints, domain information, and Interaction History.
Visual Salience Mapping Determines the relative perceptual importance of Visual Objects with respect to user interaction and context, and produces the Salient Object Matrix.
Visual Enhancement Multiplexing Collects outputs from all Sub-AIMs and assembles the Enhanced Visual Scene Descriptors and the multiplexed Visual CXT Status, Visual Domain Request, and Visual Interaction History Request.

4.4 Input/Output Data of SubAIMs

Table 3 lists the Input and Output Data of the SubAIMs of the Visual Scene Enhancement (PGM‑VSE) Composite AIM.

Table 3 – Input/Output Data of the SubAIMs of the Visual Scene Enhancement (PGM‑VSE) Composite AIM

Name Input Data Output Data
Visual Scene Description Visual Object
Visual CXT Directive
Visual Domain Response
Visual IH Response
Visual Scene Descriptors
Visual CXT Status
Visual Domain Request
Visual IH Request
Visual Object Identification Visual CXT Directive
Visual Scene Descriptors
Visual Domain Response
Visual IH Response
Visual CXT Status
Visual Object Types
Visual Domain Request
Visual IH Request
Affordance Inference Visual CXT Directive
Visual Scene Descriptors
Visual Domain Response
Visual IH Response
Affordance Tags
Interaction Potential
Visual CXT Status
Visual Domain Request
Visual IH Request
Visual Salience Mapping Visual Object Type IDs
Visual CXT Directive
Visual Scene Descriptors
Visual Domain Response
Visual IH Response
Affordance Tags
Interaction Potential
Salient Object Matrix
Affordance Tags
Visual CXT Status
Visual Domain Request
Visual IH Request
Visual Enhancement Multiplexing Visual Object Type IDs
Visual CXT Directive
Visual Scene Descriptors
Visual Domain Response
Visual IH Response
Affordance Tags
Interaction Potential
Visual Object Type IDs
Enhanced Visual Scene Descriptors
Visual CXT Status
Visual Domain Request
Visual IH Request

4.5 AIMs and JSON Metadata

Table 4 gives the Visual Scene Enhancement (AIM1) and its SubAIMs (AIM2).

Table 4 – Visual Scene Enhancement (AIM1) and its SubAIMs (AIM2)

AIM1 AIM2 Name JSON
PGM‑VSE Visual Scene Enhancement X
PGM‑VSD Visual Scene Description X
PGM‑VOI Visual Object Identification X
PGM‑AFI Affordance Inference X
PGM‑VSM Visual Salience Mapping X
PGM‑VEM Visual Enhancement Multiplexing X

5. JSON Metadata

https://schemas.mpai.community/PGM1/V1.0/AIMs/VisualSceneEnhancement.json

6. Profiles

No Profiles.

7. Reference Software

Not part of this specification.

8. Conformance Testing

A VSE implementation conforms with Visual Scene Enhancement (PGM‑VSE) if:

  1. The implementation includes all SubAIMs listed in Table 2.
  2. All I/O Data listed in Table 1 are present and conform with their respective Data Types.
  3. All SubAIM I/O Data listed in Table 3 conform with their respective Data Types.
  4. The implementation produces Enhanced Visual Scene Descriptors that refine the input Visual Scene Descriptors using the outputs of Visual Object Identification, Affordance Inference, and Visual Salience Mapping.

9. Performance Assessment

Not part of this specification.

Go to PGM-AUA V1.0 AI Modules