Next Article in Journal
Improving the Energy Efficiency of Radio Access Networks by Using an Adaptive URLLC Slot Structure Within the 5G Advanced Architecture
Previous Article in Journal
Enhancing Network Traffic Monitoring Through eXplainable Artificial Intelligence Methodologies
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

AI-Driven Reliability in 6G Networks: Enhancing QoE of Real-World Video Streaming

by
Christos Betzelos
1,
Dimitrios Uzunidis
2,
Anastasios Vetsos
1 and
Panagiotis A. Karkazis
1,*
1
Department of Informatics and Computer Engineering, University of West Attica, 12243 Athens, Greece
2
Department of Electrical and Electronics Engineering, University of West Attica, 12243 Athens, Greece
*
Author to whom correspondence should be addressed.
Telecom 2026, 7(2), 35; https://doi.org/10.3390/telecom7020035
Submission received: 2 January 2026 / Revised: 9 February 2026 / Accepted: 19 March 2026 / Published: 30 March 2026

Abstract

This paper advances user-centric Artificial Intelligence (AI) frameworks for reliability in fifth-generation and beyond (B5G) networks by examining their use in high-demand services such as video streaming. The proposed framework can leverage multi-layer monitoring across the edge–cloud continuum, application-layer metrics, and 5G core performance data to evaluate reliability through Quality of Experience (QoE) optimization. Results demonstrate that improved frame delivery can be achieved via dynamic resource prediction and proactive resource allocation. The study validates the framework’s scalability in dynamic workload conditions, emphasizing its role in mission-critical video services.

1. Introduction

The sixth generation (6G) of wireless communication systems is expected to revolutionize mobile network services with its promise of Ultra-Reliable and Low-Latency Communications (URLLC) [1]. New Key Performance Indicators (KPIs) have been defined, including sub-millisecond-level latency and increased data transmission capacity. According to the International Telecommunication Union’s (ITU) vision for international mobile telecommunications 2030 (IMT-2030), reliability over the air interface could range between 10 5 and 10 7 or even 10 9 for industrial applications [2]. Furthermore, the integration of Artificial Intelligence (AI) into network functions will enhance the user experience by enabling intelligent decision-making and management of the complexity of 6G. Reliability is further enhanced by AI-powered predictive maintenance, which enables real-time communication among machines, anticipates potential failures and optimizes operational performance within industrial environments [3]. This advancement is particularly important for mission-critical applications, such as remote surgery, autonomous vehicles and extended reality (XR), where users anticipate a consistently high Quality of Experience (QoE) [4]. As video streaming represents a significant share of global data traffic, enhancing this QoE for end users has become a key performance metric in emerging network paradigms. However, achieving reliable and dynamically adaptive video delivery across a massive number of heterogeneous devices presents significant challenges, due to the distributed nature of the whole edge–cloud continuum infrastructure.
Addressing these challenges requires a fundamental rethinking of network architectures, where computation, intelligence, and control are no longer centralized but distributed across the edge–cloud continuum. In this context, the architecture of 6G networks is decentralized, integrating advanced technologies such as Multi-Access Edge Computing (MEC) and Edge AI to satisfy stringent requirements for high availability, reliability, and low latency in mission-critical applications. By allowing computation near end users, MEC significantly shortens response times, thus supporting latency-sensitive use cases, including gaming and intelligent vehicular communication systems [5]. Complementing this, Edge AI empowers the network edge with localized processing capabilities, facilitating real-time data analysis and predictive maintenance to optimize performance. This approach helps prevent service disruptions by identifying potential issues and thereby increasing service reliability [3]. However, the distributed deployment of Edge AI brings technical challenges, such as synchronization issues between servers, necessitating the need for a resilient coordination framework. The continuous exchange of real-time information on processing and memory resources is crucial to maintaining seamless functionality [6]. As a result, reliability in 6G systems increasingly depends on precise orchestration between distributed computing entities and on communication mechanisms capable of efficiently handling large volumes of real-time monitoring data. While such architectures enable unprecedented responsiveness and scalability, they also expose new challenges related to coordination, resilience, and service continuity across heterogeneous environments.
Beyond architectural and infrastructural advances, the evolution from 5G to 6G also introduces a fundamental shift in the way network performance is evaluated and optimized. To address the recognized limitations of 5G networks, the vision for 6G moves towards a more user-centric approach [7]. Rather than focusing on the operational approach used today to solely enhance network capabilities, 6G aims to place users in the center, allowing them to define, customize and control their interaction with the services and applications. This shift requires networks to support highly personalized and adaptive services that respond intelligently to the needs and situational contexts of users [8]. The user-centric approach can ensure a high QoE for each individual user, whereas 5G networks focus mainly on guaranteeing the Quality of Service (QoS), independent of the user that utilizes the service [9]. This distinction highlights a critical gap in existing reliability approaches, which focus primarily on network-centric metrics and fail to capture the subjective and contextual dimensions of the user experience. Realizing and supporting this level of personalization requires a network infrastructure that is continuously stable, highly scalable and inherently reliable, allowing efficient processing of contextual and behavioral user data. A key enabler of this architecture, inspired by the 5G core’s openness and standardized Application Programming Interfaces (APIs), is the ability to deploy network functions decoupled from fixed physical locations. This flexibility allows for the creation, scaling, and migration of services throughout the edge–cloud continuum. Consequently, reliability mechanisms in 6G must evolve from static, infrastructure-oriented guarantees to adaptive, AI-driven functions that explicitly take into account user intent and perceived QoS.
In this paper, expanding on recent user-centric AI frameworks for 6G reliability, we propose and implement a user-centric reliability function by analyzing its impact on a video streaming application. The main contributions of this work are summarized as follows:
  • User-centric reliability modeling:We propose a novel reliability function that explicitly integrates QoE metrics into the reliability control loop, shifting the focus from traditional network-centric indicators to user-perceived service quality.
  • Multi-layer monitoring framework: We design a comprehensive monitoring approach that correlates telemetry from the edge–cloud continuum, network, and application planes, enabling holistic and context-aware reliability assessment in 6G environments.
  • Real-world video streaming validation: We implement a cloud-native video streaming platform using open-source tools and conduct Proof of Concept (PoC) experiments under heterogeneous workloads and memory constraints, demonstrating how reliability degradation can be predicted and mitigated in practice.
  • AI-driven reliability prediction and adaptive resource management: We perform a machine learning evaluation, comparing classical and deep learning models that predict reliability degradation in video streaming services and support dynamic application-layer resource scaling across the edge–cloud continuum to preserve QoE under high traffic loads and constrained resource conditions.
This approach to reliability differs from traditional ones, which rely on monolithic or single-layer solutions that assess reliability using isolated QoS metrics, such as packet loss. Although prior studies have explored AI-assisted reliability, QoS-to-QoE mapping, and resource management in 6G networks, these mechanisms typically treat QoE as an indirect or post-hoc indicator and do not incorporate it as an explicit control signal within an orchestration framework. In this work, reliability is instead grounded in the users’ perspective and their unique needs and intentions by examining their subjective QoE-related metrics that are directly used to drive orchestration decisions. By focusing on these user-defined requirements and leveraging correlated multi-layer telemetry, the proposed function dynamically selects and deploys appropriate predictive models to improve latency reduction, resource utilization, and overall network dependability. This user-centric predictive approach fosters trust in automation while strengthening the reliability of the system [10]. At the application layer, reliability is achieved by ensuring adequate provisioning of resources for Virtual Network Functions (VNFs), which are required to operate continuously throughout the service lifecycle [11]. Unlike static scaling policies, the proposed intelligent scaling mechanisms leverage QoE-driven predictions to anticipate traffic variations and potential failures, balancing performance and energy efficiency without excessive resource activation. As a result, the system achieves proactive, intent-aware reliability assurance across the edge–cloud continuum, creating a robust and efficient communication environment that users can rely on confidently.
The remainder of the paper is organized as follows. In Section 2, we present in detail the proposed architecture of network function for AI-driven user-centric reliability on 6G networks, also including the multi-layered monitoring and data collection techniques. Section 3 provides experimental scenarios designed as a PoC to evaluate the performance and adaptability of the intelligent video streaming system under memory-constrained conditions and varying workloads. Next, in Section 4, we conduct an ML-based analysis and discuss the results obtained from the experimental setup while trying to enhance the end-user’s QoE. Finally, Section 5 concludes the paper by summarizing the key findings and proposing future directions.

2. User-Centric Reliability Function

Considering that each user or tenant has different requirements in terms of trust level from a 6G system, it is necessary for the 6G system to be capable of being adapted to the specific needs and requirements of each user. These unique needs and intents of the user correspond to the desired Level of Reliability (LoR). Therefore, the 6G system should not only be reliable in a static way, but it should also become user-centric and capable of dynamically adapting the reliability level to the trust level requirements of each tenant. It should be noted that, in certain advanced implementations beyond the scope of this work, the LoR could also be derived through advanced human-machine interaction channels. Such channels can include conversational agents or Large Language Models (LLMs) capable of eliciting, interpreting, and formalizing user intent and requirements.
The user-centric reliability function aims to preserve the overall trustworthiness of the network by dynamically managing and continuously optimizing service reliability based on the provided LoR. It operates dynamically according to the intent of the user, leveraging AI-based mechanisms to respond to variations in the network environment and infrastructural capabilities. It oversees reliability across the full lifecycle of network services, from deployment to decommissioning. Through this lifecycle-oriented management, the desired LoR is consistently preserved as the network evolves and transforms.

2.1. High-Level Architecture

The reliability function is structured around three core components: the AI Agent, the vApp, and the NetworkApp. All components are implemented in Python 3.12 and correspond to version v1.0 as defined in [12]. They are deployed as cloud-native, containerized microservices that support high availability and horizontal scaling by design across the heterogeneous and distributed edge–cloud continuum. The AI Agent acts as the orchestration entity, selecting and deploying the appropriate vApp according to the LoR value received. Three LoR flavors are supported, corresponding to low (1–40%), medium (41–80%) and high (81–100%) trust requirements. The vApp component is responsible for executing the inference phase of the different AI/ML models and is deployed separately for each network service. Designed for scalability and modularity, vApps store their models within container images, enabling flexible deployment and lifecycle management. The vApp receives monitoring input from the NetworkApp via a message broker and subsequently forwards its prediction results to the AI Agent through the same messaging mechanism. The NetworkApp, also implemented and deployed as a containerized component, is designed to interface with various data sources from the multi-layered monitoring and data collection framework, with one instance per component. Further details of this monitoring framework are presented in Section 2.5. A high-level overview of the reliability function is illustrated in Figure 1.

2.2. Requirements

Given the amount of data and the complexity of tasks, the availability of high-performance computing (HPC) is a fundamental requirement for the ML-based reliability functions. This infrastructure is essential to accelerate both training and inference processes, reducing execution times for real-time reliability estimation and enabling the processing of large-scale, feature-rich datasets originating from multiple planes of the 6G network. In this context, the capability to dynamically select ML models according to different reliability levels provides the necessary intelligence to align reliability mechanisms with user intent.
In addition, integration with an MLOps framework for post-deployment evaluation enables the continuous assessment of deployed AI models with respect to performance, accuracy, and latency, thereby embedding mechanisms for ongoing improvement and accountability within the operation of the reliability function. Security requirements are addressed by enforcing that reliability operations are executed within isolated and protected environments inaccessible from public networks, protecting both the AI workflow and sensitive monitoring data from exposure or compromise. Interoperability further requires that the reliability function can seamlessly ingest streaming telemetry related to infrastructure and service status, supporting real-time, data-driven adaptation.
One of the key requirements includes access to performance data across the entire cloud continuum, the 5G/6G core, and the user application to ensure that reliability models are trained and deployed across multiple layers, thus covering the complete operational context of the 6G network. Fault tolerance and scalability are equally critical, allowing the function to remain robust under dynamic scaling conditions and resilient to failures, which is particularly important for mission-critical and high-availability services, such as video streaming. Moreover, the presence of standardized and robust interfaces to the MLOps framework enhances both system intelligence and long-term maintainability. Near-real-time monitoring enabled through a publish-subscribe paradigm is necessary to support rapid analysis and mitigation actions. User applications are also required to support management operations, such as scaling and migration, to maintain service continuity in direct response to reliability-related events. Finally, the last requirement highlights the need for reliable data persistence mechanisms, ensuring that adequate storage is available for telemetry data, models, and algorithms in distributed AI-driven deployments.

2.3. Operation

The user-centric reliability function is designed to maintain the overall trustworthiness of the system and the network through dynamic management and continuous optimization of service reliability according to the specified LoR. The operational workflow follows a well-defined sequence to ensure dependable and adaptive performance, as depicted in Figure 2. The main steps are described below [12].
  • The AI Agent receives the LoR score from the user and maps it to the appropriate reliability level (low, medium, high).
  • Based on this level, the AI Agent selects the appropriate vApp and identifies the required NetworkApp(s).
  • The AI Agent deploys the selected vApp and NetworkApp(s) through a Meta-OS orchestrator. More details on this are presented in Section 2.4.1.
  • The NetworkApp(s) retrieves input data from the multi-layered monitoring and data collection framework.
  • The vApp processes the input data from the NetworkApp(s) and generates predictions.
  • The vApp sends the predictions to the AI Agent.
  • Based on the predictions, the AI Agent determines the appropriate actions or outputs across the whole 6G edge–cloud continuum to improve service reliability.

2.4. Functionalities per Component

2.4.1. AI Agent

The AI Agent serves as the core decision-making component of the reliability function, responsible for the selection and coordinated orchestration of the AI/ML models [12]. Its primary responsibility is to translate the user-defined reliability score into the corresponding reliability level (low, medium, high). Depending on the level derived, the system determines the appropriate vApp flavor along with the associated NetworkApps from the available deployment options. The final deployment of the vApp and NetworkApp(s) is carried out via a Meta-OS orchestrator [13,14]. The AI Agent has been designed to maintain compatibility with emerging MetaOS orchestrators, including aerOS [15], enabling seamless integration across the edge–cloud continuum. This compatibility aligns with the architectural principles outlined in recent frameworks that emphasize modular orchestration and interoperability across heterogeneous environments [16,17]. The AI Agent exchanges information with the vApp and receives its predictive outputs via a message broker. Based on these predictions, it decides on the appropriate actions across the 6G edge–cloud continuum to maintain service reliability. Such actions may include resource reallocation, service migration, security policy reinforcement, or mitigation of emerging operational risks. In coordination with the vApp and the NetworkApp, the AI Agent continuously monitors system and network conditions, dynamically adjusting reliability through real-time decisions and actions to ensure the application’s trustworthiness.
The above behavior of the AI Agent can be expressed formally as follows:
Let L [ 1 , 100 ] denote the user-defined LoR. The AI Agent maps L to a reliability tier according to:
Tier ( L ) = Low , 1 L 40 , Medium , 41 L 80 , High , 81 L 100 .
The selected tier determines the deployed vApp flavor, the associated inference budget, and the aggressiveness of the control policy used by the AI Agent to trigger mitigation actions. The AI Agent enforces a tier-dependent control policy:
P ( L ) = { M ( L ) , S ( L ) , C ( L ) } ,
where M is the number of consecutive unreliable predictions required to trigger mitigation, S is the size of the action step (for example, the number of replicas added or the migration priority), and C is a cooldown period that prevents oscillations. Typical actions include scaling, reallocating, or migrating a service instance across the edge–cloud continuum, depending on the capabilities of the underlying MetaOS orchestrator. Table 1 summarizes the formal mapping between user-defined LoR values and the orchestration policies enforced by the AI Agent. The abstract service-level objectives (SLOs) are translated into measurable KPIs depending on the service type and monitoring capabilities.
Table 2 summarizes the decision logic executed by the AI Agent, including LoR-driven vApp selection and prediction-based action triggering.

2.4.2. vApp

The user-centric reliability function in 6G networks is structured to guarantee dependable and adaptive operation, aligning the vApp flavor selection with the requested LoR [12]. It implements a tier-based deployment strategy in which vApp instances are deployed according to the specified LoR, enabling user-centric reliability while balancing performance, efficiency, and security in accordance with the declared trust requirements. Distinct vApp flavors are employed, progressively scaling in computational complexity, predictive accuracy, and embedded security mechanisms as the LoR varies within the 1–100% range. These flavors are tailored according to system constraints and user requirements.
Under low-trust conditions (LoR below 40%), the system prioritizes rapid execution and efficient use of resources. Lightweight ML techniques, including k-Nearest Neighbors (k-NN) algorithms and simplified neural networks, are utilized to achieve adequate accuracy while limiting computational burden. This setup supports essential operational tasks, including basic device monitoring, data analysis and maintenance of an acceptable level of QoE. The main characteristics include:
  • Low-complexity algorithms and lightweight models.
  • Basic device monitoring for operational awareness.
  • Assurance of an acceptable QoE to ensure reliable service.
  • Sufficient predictive accuracy to meet baseline reliability needs.
This tier is tailored for environments characterized by lower trust requirements, emphasizing core reliability functionalities without incorporating advanced features.
At intermediate trust levels (LoR between 41% and 80%), the system adopts more advanced implementations based on deep neural networks. These models are designed to strike a balance between computational cost and predictive performance, achieving elevated accuracy while monitoring multiple network planes and devices. The additional computational demand ensures seamless operation and improved QoE. In particular, the main characteristics include:
  • Continuous operation across network components.
  • Utilization of deep neural networks to support advanced processing capabilities.
  • Generation of multiple alert signals to enable proactive monitoring.
  • High accuracy without excessive computational burden.
This tier is suited for applications that demand consistent performance and an enhanced user experience.
In environments with high trust (LoR exceeding 80%), the system incorporates advanced federated learning (FL) techniques combined with strengthened security mechanisms. These advanced and sophisticated implementations achieve very high predictive accuracy and support the identification of malicious activities, including Distributed Denial-of-Service (DDoS) attacks and intrusion attempts. The high LoR tier incorporates enhanced functionalities, such as:
  • Detection of malicious actions for proactive threat mitigation.
  • Security profiling mechanisms aimed at strengthening network resilience.
  • Deployment of highly complex algorithms for superior performance.
This tier targets mission-critical applications, guaranteeing secure and reliable performance under strict trust requirements.
Overall, the categorization of vApp flavors allows 6G systems to deploy progressively sophisticated ML models in alignment with the defined trust levels, ranging from fundamental service assurance to comprehensive security threat mitigation mechanisms.

2.4.3. NetworkApp

The NetworkApp continuously executes a monitoring cycle in which metrics are collected, processed, and disseminated in near-real time across the 6G edge–cloud continuum [12]. Each NetworkApp component is designed to gather data from a specific plane of the multi-layered monitoring and data collection framework described in Section 2.5. The collected data are subjected to intelligent filtering and preprocessing procedures that convert raw telemetry data into standardized representations suitable for consumption by vApps, thereby supporting interoperability and efficient resource usage within the reliability evaluation pipeline. Data delivery from the NetworkApp to the vApp is achieved through advanced message queuing protocols, where message brokers provide reliable and asynchronous publishing mechanisms capable of maintaining system responsiveness while handling high-volume metric streams. This communication layer ensures that the reliability function can maintain persistent visibility of system state variations across all operational layers and planes, supporting early detection of reliability degradation trends and enabling timely mitigation actions to preserve service quality objectives.

2.5. Multi-Layered Monitoring and Data Collection

The multi-layered monitoring and data collection framework is designed to feed the user-centric reliability function, and specifically the NetworkApps, with data from all planes of the 6G ecosystem [12]. Monitoring probes are placed in all possible layers, including the infrastructure (edge–cloud continuum), the network (5G/6G core system), and the application/service. Regarding the infrastructure plane, monitoring data may originate from physical nodes, hypervisors, virtual machines, containerized environments, and orchestration systems, covering hardware-level indicators such as Central Processing Unit (CPU) utilization, memory consumption, network interface activity, and energy metrics. Devices like User Equipments (UEs) may interface with portable network servers that integrate both Radio Access Network (RAN) and core functions, enabling telemetry from both physical and virtual resources. At the network plane, analytics can be derived from entities including the Network Data Analytics Function (NWDAF) and the User Plane Function (UPF), providing visibility into UE mobility patterns, traffic load conditions, session management behavior, traffic class differentiation, and network slice utilization. Finally, at the application/service plane, workload-related indicators are captured, such as concurrent request volume, user-to-application traffic exchanges, service latency measurements, and QoS/QoE metrics, which collectively reflect the perceived user experience.
The proposed framework integrates state-of-the-art open-source technologies to enable observability across the entire 6G ecosystem. OpenCAPIF, which is one of the most widely used Common API Framework (CAPIF) [18] implementations, provides standardized and secure API exposure within the 5G/6G core, ensuring unified access to network functions. Prometheus [19] serves as the core telemetry engine, collecting real-time metrics across distributed cloud and edge environments. To enhance observability, durability and global querying, Thanos [20] extends Prometheus with long-term storage, high availability and federated data views across geographically dispersed components. Kepler [21] is integrated for energy telemetry collection at the infrastructure level, enabling sustainability-aware decision-making in the reliability function. Together, these open-source tools provide a unified and extensible telemetry pipeline that feeds the reliability function with fine-grained data from almost all operational planes.
A key feature of the proposed reliability function is its ability to correlate this telemetry originating from different operational planes. For instance, variations in user location or device characteristics at the application plane can influence radio access conditions at the network plane and consequently impact resource utilization and performance within the infrastructure plane. By gathering indicators such as Channel Quality Indicator (CQI), Signal-to-Interference-plus-Noise Ratio (SINR), end-to-end delay, and traffic volume, and associating them with user-centric QoE metrics, such as frame rate in video streaming services or real-time latency, the reliability function can trace how lower-layer dynamics affect perceived service quality, while simultaneously detecting early signals of degradation or anomalous behavior. This holistic cross-plane correlation enhances the training process of AI/ML models, enabling the prediction of reliability-related anomalies or performance-breaking points and facilitating timely alerts that trigger proactive mitigation measures, such as scaling, migration, or service reconfiguration.
The monitoring framework also allows for user-centric interpretation of reliability by incorporating service and user-related characteristics. For example, different video quality demands in terms of resolution and frame rate, variations in radio conditions such as Reference Signal Received Power (RSRP) and Reference Signal Received Quality (RSRQ), or changes in 5G QoS Identifiers (5QI) values, reflecting different service requirements, can all affect the perceived QoE. Correlating these indicators with infrastructure and application metrics allows the reliability function to adapt its decisions to individual user intents.

3. Experimental Setup

In this section, we present the cloud-native experimental video streaming infrastructure used to validate the proposed user-centric reliability function. The PoC scenarios have been designed to evaluate the performance and adaptability of the system under memory-constrained conditions and varying workloads.

3.1. Architecture

The main infrastructure of the proposed video streaming platform is built around MediaMTX [22], a modern open-source media server designed for real-time video and audio streaming. MediaMTX supports a wide range of protocols, making it highly adaptable to diverse streaming scenarios. In our implementation, the Real-Time Streaming Protocol (RTSP) [23] is used as the primary protocol for media delivery. The architecture, as depicted in Figure 3, is composed of the following modular components, each serving a distinct role:
  • Publisher: Acts as the origin of the media stream. It generates or receives video content from an external source and publishes it to the MediaMTX infrastructure.
  • Main Server: Serves as the central configuration and control node. It manages routing logic, stream registration, and coordination between publishers and streamers.
  • Streamers: Distributed nodes that receive the media stream from the publisher and broadcast it to end clients. Streamers operate independently and can be horizontally scaled to accommodate increased loads. They support multiple streaming options and can be configured to operate under computational resource constraints.
  • End Clients: Connect directly to streamers through RTSP to consume media content. Clients may include desktop applications, embedded devices, or browser-based video players.
This modular and distributed architectural design ensures scalability and fault tolerance, enabling both small- and large-scale deployments while providing a controlled environment for video stream quality evaluation.

3.2. Data Collection and Metrics

To optimize data collection and enhance system observability, the proposed experimental video streaming validation infrastructure incorporates elements of the previously described multi-layered monitoring framework. The architecture is deployed as a set of containerized microservices that provide modularity, scalability and fault isolation. Each component exposes performance metrics in compliance with Prometheus standards, which support near-real-time data monitoring and analysis. At the generic infrastructure and container level, the monitoring system captures metrics such as CPU utilization, memory usage, network throughput and energy consumption, providing a foundation for assessing the efficiency and reliability of the underlying computing environment.
At the application level, the MediaMTX components extend observability by exposing specific protocol-aware metrics. These include the number of active streamers, the number of active RTSP sessions, the frames per second and the packets received and sent per session. Error-related metrics remain relevant for diagnosing anomalies and failures, although packet loss and jitter are largely mitigated by the Transmission Control Protocol (TCP) retransmission and flow control. By combining infrastructure and container-level with protocol-aware metrics, we enable a fine-grained assessment of user-centric QoE, particularly the performance and reliability of the video streaming perceived by end clients.
In this experimental setup, all the metrics mentioned ensure that the monitoring is focused on the most critical factors that affect users’ performance and reliability. They are structured as feature vectors and exported as JavaScript Object Notation (JSON) objects, enabling integration with appropriate AI/ML algorithms designed for streaming performance optimization, anomaly detection, and proactive resource migration and scaling. Consequently, they form the backbone of our dataset, offering clear insights into system and network behavior.

3.3. Validation Scenarios

For validation, we employ the open-source dataset and framework presented in [24], which enables multi-layer monitoring and proactive scaling of 6G video streaming applications, based on the experimental architecture described previously. The dataset models dynamic scaling behavior, where the number of concurrent streamers increases from one to three, enabling the observation of system performance during both scale-up and scale-down operations. The goal of this setup is to stress the video streaming application in order to identify critical memory thresholds that may result in performance degradation or application failures. To ensure a realistic and distributed workload, the stress evaluation is conducted across three independent instances, each hosted in different environments, with client streams increasing progressively from one to thirty in each instance. Building on this foundation, our validation is divided into two options that differ in memory allocation per streamer: 100 MiB and 200 MiB, respectively. Therefore, since the number of streamers is increased from one to three and the memory constraint per streamer is either 100 MiB or 200 MiB, the final dataset of the experimental setup consists of six folds, including all possible scenarios. For each experimental scenario, the monitoring stack generates a CSV file that contains the memory utilization ratio of each streamer and a binary reliability label that indicates whether, within that sampling interval, at least one stream has experienced a failure or severe degradation. These labels originate directly from QoE-related metrics at the application level, based on the continuity of frame delivery and the successful maintenance of active RTSP sessions.
Across all six scenarios, the resulting datasets exhibit a moderate to strong class imbalance, with unreliable states representing between 8.1% and 27.0% of the samples, as shown in Table 3. This imbalance reflects realistic operational conditions, where failures and severe QoE degradation events occur less frequently than normal service operation. Importantly, the degree of imbalance varies between the scenarios, allowing the evaluation of model robustness under heterogeneous class distributions.
In addition to varying the number of streamers and memory constraints, the validation scenarios are designed to reflect heterogeneous user-centric demands at the application level. Specifically, the end clients connected to the different streamers consume different video qualities, corresponding to distinct resolutions and bitrates. This setup emulates realistic video streaming conditions, where users may request different content qualities based on device capabilities, network conditions, or individual preferences. By serving multiple video qualities simultaneously, the experimental setup allows the proposed reliability function to be evaluated under diverse QoE requirements, demonstrating its ability to maintain service reliability across heterogeneous user profiles rather than optimizing for a single uniform demand.

4. ML-Based Analysis

To deepen the analysis of the validation methodology described previously, we conducted an extensive ML-based evaluation using all the different scenarios. This analysis aimed to quantify how memory constraints, the number of streamers and the number of clients influence system reliability and to determine whether early indicators of degradation can be learned directly from this telemetry exported by the experimental streaming infrastructure. In particular, we focused on low and medium vApp flavors, which correspond to lightweight and moderately complex inference pipelines within the proposed reliability function. These vApp flavors are representative of practical deployments where reliability predictions are generated at runtime and directly consumed by the AI Agent to trigger orchestration actions according to the LoR-dependent control policy. Although the control policy parameters associated with each LoR tier (e.g., mitigation thresholds, action step sizes and cooldown periods) are formally defined in the proposed architecture, the present evaluation aimed at the reliability inference capability of the vApp flavors and not at the quantitative analysis of closed-loop control behavior and policy parameter tuning. The entire workflow was implemented in Python, leveraging a classical machine learning pipeline designed to process raw monitoring traces, extract time-dependent features, and train multiple models capable of distinguishing reliable from unreliable operational states.

4.1. Data Pre-Processing and Feature Extraction

To capture temporal patterns inherent in the streaming workload, the data was processed using a sliding-window scheme of 40 consecutive samples. For each window, a set of statistical descriptors was computed for memory usage, resulting in a set of 14 aggregated statistical features, such as mean, standard deviation, minimum, maximum, skewness, kurtosis, and percentiles. The resulting feature vectors were scaled using Min-Max normalization to facilitate the training of all models. This windowing methodology allowed the models to capture the gradual buildup of resource pressure, which is particularly important in constrained-memory environments. The binary reliability label was then assigned per window according to the failed streams, ensuring an alignment between the machine-learning analysis and the QoE goals defined in the previous section.
Sliding windows were constructed using a time-ordered procedure providing short-range temporal context for each inference step. Each window aggregated telemetry samples within a bounded interval and the associated reliability label reflected the service state observed in that same window. No information from future evaluation windows or unseen scenarios is used during training or inference. In addition, training, validation, and test splits preserve scenario-level isolation, ensuring the absence of label leakage across datasets.
Although the monitoring framework provides a wide range of metrics across multiple planes, memory utilization was selected as the primary feature for this analysis, as memory exhaustion was empirically observed to be the dominant precursor to service degradation and streaming failures in the evaluated scenarios. This choice reflects the objective of the PoC evaluation, which focuses on isolating a dominant and reproducible degradation factor, while the proposed framework remains agnostic to the specific metric and can incorporate network and RAN-level telemetry in future validations.

4.2. Cross-Scenario Cross-Validation Strategy

To ensure a scenario-agnostic evaluation, we adopted a cross-validation scheme in which each of the six dataset files served as a hold-out test set once. In each fold, four files were used for training and one for validation, while the remaining file was reserved for testing. This rolling configuration guarantees that all combinations of memory allocation (100 MiB and 200 MiB) and number of streamers (one to three) are observed both during training and evaluation, but never simultaneously in the same role. As a result, the models are systematically exposed to heterogeneous operating conditions and are required to generalize to an unseen scenario in each fold, closely reflecting the practical requirements of AI-driven reliability mechanisms in 6G environments.
For each fold, we trained four classical ML algorithms along with one deep-learning model:
  • Support Vector Classifier (SVC);
  • Decision Tree Classifier (DT);
  • Random Forest Classifier (RF);
  • k-Nearest Neighbors (k-NN);
  • Deep Neural Network (DNN).
These models were initialized with a basic configuration and then tuned using a grid search procedure in the validation set. The grid search explored model-specific hyperparameters, such as the kernel type and regularization strength for SVC, the maximum tree depth and minimum samples per split or leaf for DT and RF, the number of neighbors and the distance metric for k-NN, and the hidden layers and the learning rate for DNN. The search was guided by a cross-validated F1-score (weighted) in the validation set, and the best-performing configuration was then selected as the final model. The optimal hyperparameters for each model are presented in Table 4. The implementation automatically tracked accuracy, precision, recall, F1-score and per-class accuracy (for reliable and unreliable states) in the held-out test scenario, allowing an extensive assessment of each model’s ability to detect degradation patterns.
The DNN, which warrants a more detailed architectural examination, followed a fully connected structure with an explicit input layer matching the dimensions of the feature space, followed by the hidden layers with rectified linear unit (ReLU) activations and dropout regularization, and a single sigmoid output neuron for binary classification. The network was optimized using the Adam optimizer and trained on the training set while monitoring the validation loss. To avoid overfitting and mimic realistic resource-aware deployments, early stopping was used, restoring the best-performing weights observed during training. This design choice reflects the role of DNN as a more computationally expensive but still practical model that can be mapped to a medium LoR of our reliability function, where additional computational complexity is acceptable in exchange for higher predictive quality.

4.3. Results

The cross-scenario evaluation provides a perspective on how the different learning models behave under varying memory constraints and workload intensities. Figure 4 presents the aggregated results across all six scenarios, summarizing global accuracy, precision, recall, and F1-score (including standard deviations) for each model across all folds. In parallel, Figure 5 shows the corresponding confusion matrices computed from the same aggregated results. Together, these figures reveal consistent and interpretable patterns that highlight both the strengths of the underlying monitoring design and the ability of AI-driven methods to detect reliability degradation in resource-constrained video streaming environments.
Across all scenarios, the DNN provides the most stable and high-performing behavior, achieving accuracy levels that remain consistently above 0.97. As shown in Figure 4, the DNN trace displays minimal variability across scenarios, indicating its representational capacity enables it to generalize effectively across heterogeneous operational contexts. Furthermore, Figure 5 illustrates that the DNN produces the fewest false negatives (i.e., instances where unreliable states are misclassified as reliable), aligning with the expectation that deep learning models, given sufficient non-linear structure, can better capture the temporal evolution embedded in the derived statistical features.
From a reliability perspective, the most important advantage of the DNN is its improved recall, particularly for the unreliable class. While classical ML models tend to exhibit slightly lower recall, indicating a higher tendency to misclassify early degradation signals, DNN demonstrates a better ability to correctly detect windows preceding reliability failures. As we already mentioned, this is especially important for proactive reliability management, where false negatives (i.e., missing an imminent degradation event) are considerably more harmful than false positives.
Although the mentioned metrics provide an initial indication of performance, the datasets considered in this work exhibit strong class imbalance, with unreliable states representing a minority of samples in all scenarios. To provide a threshold-independent evaluation and assess robustness to imbalance, we additionally report Receiver Operating Characteristic (ROC) and Precision-Recall (PR) curves for the DNN model, as presented in Figure 6. The achieved area under the ROC curve (AUROC) of 0.9965 confirms that the DNN maintains excellent separability between reliable and unreliable operational states across a wide range of decision thresholds. However, as ROC analysis can remain optimistic in imbalanced settings, we further examine the area under the PR curve (AUPRC). The DNN achieves an AUPRC of 0.9841, significantly higher than the baseline precision of 0.1717 corresponding to the class prevalence of unreliable states. This result demonstrates that the model is highly effective in identifying rare reliability degradation events while maintaining strong precision, even under severe class imbalance.
Beyond classification accuracy, the reliability function requires well-calibrated probability estimates, as these outputs are directly consumed by the AI Agent to trigger orchestration actions according to the LoR-dependent control policy. To assess calibration quality, we computed the Expected Calibration Error (ECE) for the DNN across all six scenarios, as reported in Table 5. The obtained ECE values are consistently low, with an average of 0.029, indicating that the predicted probabilities closely match the observed empirical frequencies. In particular, calibration improves in the less constrained 200 MiB scenarios, where the system exhibits more stable behavior. These results demonstrate that the DNN provides reliable confidence estimates, which is a critical requirement for trust-aware and user-centric reliability orchestration in 6G networks.
To complement the aggregated comparison, Figure 7 provides a detailed scenario-by-scenario analysis of the DNN’s performance in the held-out test sets. Specifically, the plots visualize the predicted reliability labels compared to the actual ones, along with the corresponding average memory usage ratio, allowing us to interpret how the DNN responds to the temporal evolution of resource constraints. Across all scenarios, the model demonstrates a remarkable ability to correctly predict transitions from reliable to unreliable states, confirming its strength under unseen operational conditions.
Across all six scenarios, the visual representation of the DNN predictions reveals another important feature: the model consistently identifies the first transition into the unreliable state with high accuracy, effectively providing an estimate of the Time Till First Major Failure (TTFMF) or, equivalently, the Mean Time to Failure (MTTF) at the window granularity. This capability is clearly visible in Figure 7, where the predicted label switches to the unreliable state at the very first occurrence of a degradation event, regardless of whether the system operates under strict (100 MiB) or relaxed (200 MiB) memory constraints. The DNN neither delays the detection nor exhibits premature triggering. Instead, it aligns almost perfectly with the actual onset of reliability degradation. Although the reliability label corresponds to the service state at the window endpoint, the temporal analysis shows that the model identifies the transition into the unreliable state at its earliest manifestation in the monitoring data. As a result, the effective forecast horizon can be interpreted as the number of windows by which degradation precursors are detected before a fully manifested failure. In the evaluated scenarios, this corresponds to a minimum lead time of one window before the first failure event. As such, the DNN not only classifies reliability states accurately, but also offers early situational awareness, enabling the reliability function to predict and mitigate upcoming QoE degradation. From a user-centric perspective, the ability to detect degradation early translates directly into maintaining perceived service continuity, particularly for users who consume higher quality video with a lower tolerance for interruptions.
Overall, the results indicate that the proposed ML pipeline is capable of discriminating between reliable and unreliable operational states, with DNN consistently outperforming the classical ML baselines across most metrics and scenarios. Reliability degradations are consistently detected with near-perfect recall, which is crucial for proactive reliability management in video streaming applications. No significant over-prediction is observed, confirming that the DNN captures degradation patterns with the stability required for deployment as part of a real-time vApp within the proposed user-centric reliability function. These results demonstrate that the DNN is well-positioned to support medium LoR vApp flavors, providing accurate and reliable predictions across the full range of scenarios considered in this work. From an orchestration perspective, the ability of DNN to accurately detect the onset of service degradation and track transitions between reliable and unreliable states provides a timely and robust control signal to the AI Agent, enabling effective dynamic service-level adaptation under the LoR-driven orchestration framework.

5. Conclusions

This work presented a user-centric reliability function designed to enhance the QoE of emerging 6G environments, validated through a real-world video streaming application. By integrating multi-layer monitoring and AI/ML mechanisms, the proposed framework demonstrates how reliability can be aligned with user intent through the LoR. The architecture links telemetry from the edge–cloud continuum, the 5G/6G core and the application plane, enabling a holistic assessment of service conditions and providing the foundation for dynamic reliability management. The modular design of the AI Agent, the scalable deployment of vApp flavors and the continuous data flow provided by the NetworkApp highlight the feasibility of a flexible and user intent-driven reliability framework suitable for 6G networks. The experimental evaluation, conducted as a PoC on an open-source dataset from a 6G video streaming application setup, confirmed that resource-constrained scenarios can be effectively predicted and mitigated through the integration of ML-based analysis within the reliability function. Specifically, the deep learning model consistently identified degradation patterns across heterogeneous workloads, validating the suitability of the proposed monitoring and inference pipeline for user-centric reliability control. While the experimental evaluation focuses on memory-driven degradation scenarios, the proposed framework is independent of the specific resource type and can support additional metrics from all layers, such as network traffic or radio conditions, in future validations. Therefore, the methodology can be extended to additional mission-critical applications that exhibit similar reliability-sensitive characteristics.
Looking forward to future work, we plan to extend the function by incorporating continuous learning techniques to further enhance AI/ML adaptability. As 6G systems evolve and applications become increasingly heterogeneous, the ability of vApps to retrain or refine their models based on recent operational data will be essential to maintain prediction accuracy and minimize drift. In parallel, future work can explore the refinement of the LoR-dependent control policy parameters used by the AI Agent, aiming to better align mitigation aggressiveness and responsiveness with the confidence and temporal characteristics of the predicted reliability signals. Furthermore, integration with real-time telemetry data from the RAN and 5G/6G network functions could also enable a more comprehensive reliability estimation, particularly in mobility-intensive environments. Another promising direction is the exploration of distributed or federated ML methods that can preserve data locality, strengthening privacy, an important consideration for high-LoR scenarios. Extending the current validation approach to additional application areas, such as XR or industrial control, would allow a broader assessment of how user-centric reliability mechanisms perform under various QoE requirements. Finally, a detailed quantitative characterization of end-to-end inference latency and compute/energy overhead per vApp flavor on representative edge hardware could also constitute an important direction for future work.

Author Contributions

Conceptualization, C.B. and P.A.K.; Methodology, C.B. and D.U.; Software, C.B. and A.V.; Validation, C.B., D.U. and P.A.K.; Formal analysis, C.B. and D.U.; Investigation, C.B., D.U. and P.A.K.; Resources, D.U. and A.V.; Data curation, D.U.; Writing—original draft, C.B., D.U. and A.V.; Writing—review & editing, P.A.K.; Visualization, C.B. and A.V.; Supervision, P.A.K.; Project administration, P.A.K.; Funding acquisition, P.A.K. All authors have read and agreed to the published version of the manuscript.

Funding

The work presented in this paper is (partially) supported by the SAFE-6G project that has received funding from the Smart Networks and Services Joint Undertaking (SNS JU) under the European Union’s Horizon Europe research and innovation programme under Grant Agreement No 101139031.

Data Availability Statement

The original data presented in the study are openly available in Zenodo at https://doi.org/10.5281/zenodo.17523181 (accessed on 20 December 2025). The algorithms and methodologies are fully described to ensure reproducibility.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Shamsabadi, A.A.; Yadav, A.; Gadallah, Y.; Yanikomeroglu, H. Exploring the 6G potentials: Immersive, hyper-reliable, and low-latency communication. IEEE Veh. Technol. Mag. 2025, 20, 74–82. [Google Scholar] [CrossRef]
  2. International Telecommunication Union Radiocommunication Sector. Framework and Overall Objectives of the Future Development of IMT for 2030 and Beyond; ITU: Geneva, Switzerland, 2023. [Google Scholar]
  3. Chataut, R.; Nankya, M.; Akl, R. 6G Networks and the AI Revolution-Exploring Technologies, Applications, and Emerging Challenges. Sensors 2024, 24, 1888. [Google Scholar] [CrossRef] [PubMed]
  4. Mohammed, S.M.; Al-Barrak, A.; Mahmood, N.T. Enabling Technologies for Ultra-Low Latency and High-Reliability Communication in 6G Networks. Ing. Syst. d’Inf. 2024, 29, 1195–1208. [Google Scholar] [CrossRef]
  5. Gunasekaran, K.; Dhanasekaran, S.; Kumar, R.V.; Aswath, S. Advanced Beamforming and Multi-Access Edge Computing: Empowering Ultra-Reliable and Low-Latency Applications in 6G Networks. Int. J. Commun. Syst. 2024, 38, e6027. [Google Scholar] [CrossRef]
  6. Rodriguez, F.; Ahmad, I.; Huusko, J.; Seppänen, K. Towards dependable 6G networks. TechRxiv 2022. [Google Scholar] [CrossRef]
  7. Drampalou, S.F.; Uzunidis, D.; Vetsos, A.; Miridakis, N.I.; Karkazis, P. A User-Centric Perspective of 6G Networks: A Survey. IEEE Access 2024, 12, 190255–190294. [Google Scholar] [CrossRef]
  8. SAFE-6G. A Smart and Adaptive Framework for Enhancing Trust in 6G Networks. Smart Networks and Services Joint Undertaking (SNS JU) Under the European Union’s Horizon Europe Research and Innovation Programme Under Grant Agreement No 101139031. 2023. Available online: https://safe-6g.eu/ (accessed on 20 December 2025).
  9. Betzelos, C.; Uzunidis, D.; Karkazis, P.; Leligou, H.C. Enhancing Reliability in 6G Networks Based on User-Driven AI. In Proceedings of the 29th International Workshop on Computer Aided Modeling and Design of Communication Links and Networks (CAMAD), Athens, Greece, 21–23 October 2024. [Google Scholar]
  10. Taleghani, E.S.; Valencia, R.I.M.; Orozco, A.L.S.; Villalba, L.J.G. Trust Evaluation Techniques for 6G Networks: A Comprehensive Survey with Fuzzy Algorithm Approach. Electronics 2024, 13, 3013. [Google Scholar] [CrossRef]
  11. Ortin, J.; Serrano, P.; Garcia-Reinoso, J.; Banchs, A. Analysis of Scaling Policies for NFV Providing 5G/6G Reliability Levels with Fallible Servers. IEEE Trans. Netw. Serv. Manag. 2022, 19, 1287–1305. [Google Scholar] [CrossRef]
  12. SAFE-6G Consortium. D4.1-SAFE-6G Project Deliverable: Cognitive Coordinator, AI Agents and User-Centric Functions; SAFE-6G Project: European Union. 2025. Available online: https://safe-6g.eu/wp-content/uploads/2025/07/D4.1_SAFE-6G_v1.0.pdf (accessed on 20 December 2025).
  13. Trakadas, P.; Masip-Bruin, X.; Facca, F.M.; Spantideas, S.T.; Giannopoulos, A.E.; Kapsalis, N.C.; Martins, R.; Bosani, E.; Ramon, J.; Prats, R.G.; et al. A Reference Architecture for Cloud-Edge Meta-Operating Systems Enabling Cross-Domain, Data-Intensive, ML-Assisted Applications: Architectural Overview and Key Concepts. Sensors 2022, 22, 9003. [Google Scholar] [CrossRef] [PubMed]
  14. Rossini, R.; Velivassaki, T.H.; Voulkidis, A.; Zahariadis, T.; Karkazis, P.; Skias, D.; Montanera, E.P.P. Open Source in NExt Generation Meta Operating Systems (NEMO). In Proceedings of the 9th International Conference on Smart and Sustainable Technologies (SpliTech), Split, Croatia, 25–28 June 2024. [Google Scholar]
  15. aerOS EU HE Project: Autonomous, ScalablE, TRustworthy, Intelligent European Meta Operating System for the IoT Edge-Cloud Continuum. Available online: https://aeros-project.eu/ (accessed on 20 December 2025).
  16. Vaño, R.; Lacalle, I.; Sowiński, P.; S-Julián, R.; Palau, C.E. Cloud-Native Workload Orchestration at the Edge: A Deployment Review and Future Directions. Sensors 2023, 23, 2215. [Google Scholar] [CrossRef] [PubMed]
  17. Horn, G.; Verginadis, Y.; Ledakis, G.; Papageorgopoulos, N.; Veloudis, S. An Implemented Architecture for a Meta Operating System Managing Distributed Applications Across Heterogeneous Resources. In Advanced Information Networking and Applications. AINA 2025; Lecture Notes on Data Engineering and Communications Technologies; Barolli, L., Ed.; Springer: Cham, Switzerland, 2025; Volume 250. [Google Scholar]
  18. 3GPP. Common API Framework for 3GPP Northbound APIs; Technical Report (TR) 23.222 V19.1.0; 3GPP: Sophia Antipolis, France, 2017. [Google Scholar]
  19. Prometheus. Monitoring System & Time Series Database. Available online: https://prometheus.io/ (accessed on 20 December 2025).
  20. Thanos. Highly Available Prometheus Setup with Long Term Storage Capabilities. Available online: https://thanos.io/ (accessed on 20 December 2025).
  21. Kepler Project. Kepler. Available online: https://github.com/sustainable-computing-io/kepler (accessed on 20 December 2025).
  22. MediaMTX. The Modern Streaming Server. Available online: https://github.com/bluenviron/mediamtx (accessed on 20 December 2025).
  23. Schulzrinne, H.; Rao, A.; Lanphier, R. Real-Time Streaming Protocol (RTSP). IETF RFC 2326. 1998. Available online: https://datatracker.ietf.org/doc/html/rfc2326 (accessed on 20 December 2025).
  24. Drampalou, S.; Uzunidis, D.; Vetsos, A.; Karkazis, P.; Koumaras, H. An Open-Source Framework and Dataset for Multi-layer Monitoring and Predictive Autoscaling of 6G Video Streaming, Zenodo: Geneva, Switzerland, 2025. [CrossRef]
Figure 1. High-level overview of the reliability function (adapted with permission from [12]).
Figure 1. High-level overview of the reliability function (adapted with permission from [12]).
Telecom 07 00035 g001
Figure 2. Sequence diagram of the reliability function (adapted with permission from [12]).
Figure 2. Sequence diagram of the reliability function (adapted with permission from [12]).
Telecom 07 00035 g002
Figure 3. Architecture of the experimental setup.
Figure 3. Architecture of the experimental setup.
Telecom 07 00035 g003
Figure 4. Grouped bar chart comparing accuracy, precision, recall, and F1 score per model, including their standard deviations.
Figure 4. Grouped bar chart comparing accuracy, precision, recall, and F1 score per model, including their standard deviations.
Telecom 07 00035 g004
Figure 5. Confusion matrices per model for the aggregated results of all six scenarios: (a) SVC. (b) DT. (c) RF. (d) k-NN. (e) DNN.
Figure 5. Confusion matrices per model for the aggregated results of all six scenarios: (a) SVC. (b) DT. (c) RF. (d) k-NN. (e) DNN.
Telecom 07 00035 g005
Figure 6. Receiver Operating Characteristic (ROC) curve and Precision–Recall Curve (PRC) for the DNN, aggregated across all six scenarios: (a) ROC. (b) PRC.
Figure 6. Receiver Operating Characteristic (ROC) curve and Precision–Recall Curve (PRC) for the DNN, aggregated across all six scenarios: (a) ROC. (b) PRC.
Telecom 07 00035 g006
Figure 7. DNN’s predicted reliability labels compared to the actual ones, along with the corresponding average memory usage ratio per scenario: (a) One streamer with 100 MiB memory constraints. (b) One streamer with 200 MiB memory constraints. (c) Two streamers with 100 MiB memory constraints. (d) Two streamers with 200 MiB memory constraints. (e) Three streamers with 100 MiB memory constraints. (f) Three streamers with 200 MiB memory constraints.
Figure 7. DNN’s predicted reliability labels compared to the actual ones, along with the corresponding average memory usage ratio per scenario: (a) One streamer with 100 MiB memory constraints. (b) One streamer with 200 MiB memory constraints. (c) Two streamers with 100 MiB memory constraints. (d) Two streamers with 200 MiB memory constraints. (e) Three streamers with 100 MiB memory constraints. (f) Three streamers with 200 MiB memory constraints.
Telecom 07 00035 g007
Table 1. Mapping of LoR value to vApp selection, SLOs, inference budgets, and control policy parameters.
Table 1. Mapping of LoR value to vApp selection, SLOs, inference budgets, and control policy parameters.
LoR ValuevApp FlavorTarget SLO/KPIInference & Control Policy
1–40LowAcceptable service continuity under best-effort guaranteesRelaxed inference latency
Conservative control
Long cooldown
41–80MediumEnhanced service continuity with limited QoE degradation toleranceModerate inference latency
Balanced control
Medium cooldown
81–100HighStrict service continuity and minimal QoE degradationStrict inference latency
Aggressive control
Short cooldown
Table 2. AI Agent decision logic for LoR-driven vApp selection and runtime actions.
Table 2. AI Agent decision logic for LoR-driven vApp selection and runtime actions.
StepDescription
1Receive LoR value L
2Map L to a reliability tier using Tier ( L )
3Select the corresponding vApp flavor based on the tier
4Deploy the selected vApp and associated NetworkApp(s) via Meta-OS API
5while the service is active do
6    Perform reliability inference via the vApp
7    Update the reliability state history
8    if control policy P ( L ) = { M , S , C } is satisfied then
9       Trigger runtime mitigation action with step size S via Meta-OS API
10       Enforce cooldown period C
11    end if
12end while
Table 3. Class distribution across the six validation scenarios. Values are expressed as the number of samples, with percentages reported in parentheses.
Table 3. Class distribution across the six validation scenarios. Values are expressed as the number of samples, with percentages reported in parentheses.
ScenarioReliable, n (%)Unreliable, n (%)Total Samples, n
1 streamer—100 MiB759 (83.0)156 (17.0)915
2 streamers—100 MiB688 (73.0)255 (27.0)943
3 streamers—100 MiB716 (75.4)234 (24.6)950
1 streamer—200 MiB781 (87.0)117 (13.0)898
2 streamers—200 MiB792 (87.1)117 (12.9)909
3 streamers—200 MiB881 (91.9)78 (8.1)959
Table 4. Optimal hyperparameters for each ML model and scenario.
Table 4. Optimal hyperparameters for each ML model and scenario.
ModelHyperparameters
SVCkernel = poly, gamma = scale, C = 3
DTmin-samples-split = 10, min-samples-leaf = 5, max-depth = 7
RFn-estimators = 10, min-samples-split = 2, min-samples-leaf = 1, max-depth = 15
k-NNweights = distance, n-neighbors = 7, metric = manhattan
DNNhidden-layers = 3, units = 16, dropout-rate = 0.1, learning-rate = 0.001, batch-size = 32, epochs = 100
Table 5. Expected Calibration Error (ECE) for the DNN across validation scenarios.
Table 5. Expected Calibration Error (ECE) for the DNN across validation scenarios.
ScenarioECE
1 streamer—100 MiB0.0654
2 streamers—100 MiB0.0065
3 streamers—100 MiB0.0912
1 streamer—200 MiB0.0016
2 streamers—200 MiB0.0021
3 streamers—200 MiB0.0076
Average0.0290
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Betzelos, C.; Uzunidis, D.; Vetsos, A.; Karkazis, P.A. AI-Driven Reliability in 6G Networks: Enhancing QoE of Real-World Video Streaming. Telecom 2026, 7, 35. https://doi.org/10.3390/telecom7020035

AMA Style

Betzelos C, Uzunidis D, Vetsos A, Karkazis PA. AI-Driven Reliability in 6G Networks: Enhancing QoE of Real-World Video Streaming. Telecom. 2026; 7(2):35. https://doi.org/10.3390/telecom7020035

Chicago/Turabian Style

Betzelos, Christos, Dimitrios Uzunidis, Anastasios Vetsos, and Panagiotis A. Karkazis. 2026. "AI-Driven Reliability in 6G Networks: Enhancing QoE of Real-World Video Streaming" Telecom 7, no. 2: 35. https://doi.org/10.3390/telecom7020035

APA Style

Betzelos, C., Uzunidis, D., Vetsos, A., & Karkazis, P. A. (2026). AI-Driven Reliability in 6G Networks: Enhancing QoE of Real-World Video Streaming. Telecom, 7(2), 35. https://doi.org/10.3390/telecom7020035

Article Metrics

Back to TopTop