1. Introduction
In recent years, generative artificial intelligence and, in particular, large language models (LLMs) have rapidly transformed the landscape of data analysis, knowledge extraction, content generation, and intelligent decision support. The emergence of increasingly capable foundation models has accelerated progress in natural language processing and has also expanded the reach of AI into a wide variety of domains, including, among others, healthcare, education, cybersecurity, public administration, finance, accessibility, and social media analysis [
1,
2,
3]. Beyond their strong performance in text-generation and question-answering, these models are now being employed as general-purpose tools for summarization, reasoning, annotation, explanation, and interaction, often serving as the core computational component of more complex intelligent systems.
This fast-moving evolution has produced major opportunities, but it has also made visible a number of important gaps in current research and practice. Although the capabilities of LLMs continue to improve, several issues remain insufficiently addressed, including factual reliability, hallucination mitigation, ambiguity handling, explainability, reproducibility, bias control, and the robust integration of external knowledge sources [
4,
5,
6]. At the same time, there is a growing need for more systematic evidence on how these models behave in domain-specific and high-stakes settings, where performance alone is not enough and must be complemented by trustworthiness, transparency, safety, and meaningful human oversight. Recent studies have also shown the value of LLMs in supporting advanced interpretability and analytical tasks; for example, in the analysis of online opinions, reviews, and social media content [
7,
8]. These developments suggest that the field is entering a new phase, in which the focus is shifting from raw model capability to the effective, responsible, and context-aware deployment of generative AI.
This Special Issue, entitled “Generative AI and Large Language Models”, was conceived precisely in response to these challenges. Its objective was to collect contributions capable of advancing the field along multiple complementary directions: novel methods for retrieval and reasoning, rigorous approaches to evaluation and trustworthiness, analytical studies able to structure the state-of-the-art, and application-oriented investigations in socially relevant domains. The ten papers published in this Special Issue reflect the diversity and vitality of the field. Together, they offer a useful picture of how current research is addressing open questions related to grounded generation, agentic systems, domain adaptation, security, explainability, accessibility, and real-world adoption.
2. Overview of the Published Papers
The ten papers included in this Special Issue cover a broad set of research directions and application areas, highlighting both the methodological richness of the field and its growing practical relevance. The first cluster of contributions focuses on retrieval-augmented generation (RAG), contextual reasoning, and the evaluation of grounded responses. Ye et al. propose TRACE, a topical reasoning framework with adaptive contextual experts for long-text summarization, showing how document structure and semantic relationships can be exploited to improve structure-aware retrieval and multi-expert reasoning [
9]. Mansurova et al. investigate QA-RAG and study the extent to which LLMs effectively rely on external knowledge, emphasizing both the potential of retrieval-based pipelines and the challenges that still remain in integrating external truth in a reliable way [
10]. In their paper, Papageorgiou et al. examine faithfulness in agentic RAG systems for e-governance and introduce a modular framework based on LLM judges to analyze hallucination and redundancy across alternative retrieval pipelines [
11]. Collectively, these studies show that retrieval is not merely an auxiliary component of LLM-based systems, but a central design dimension whose quality must be considered together with rigorous evaluation of grounding, faithfulness, and attribution.
The second group of papers addresses the use of LLMs in human-centered and socially relevant contexts. Andruccioli et al. explore the role of LLMs in sustainable and inclusive web accessibility, showing how these models can support the identification of accessibility issues in dynamically generated web content that may be overlooked by conventional validation tools, while also discussing the risks associated with redundant or hallucinated warnings [
12]. Alostad investigates the use of LLMs as annotators for stance detection in the Kuwaiti dialect, demonstrating that carefully prompted open models can produce promising results even in a low-resource linguistic setting [
13]. In the educational domain, Mitroulias and Sioutas present a systematic review and bibliometric analysis of automated multiple-choice question generation, offering a structured perspective on the development of this research area and its intersection with recent LLM-based methods [
14]. These contributions illustrate the increasing ability of generative AI to support inclusion, education, and language technologies beyond high-resource scenarios, while also reinforcing the importance of prompt design, careful benchmarking, and domain-sensitive evaluation.
Healthcare is another area prominently represented in the Special Issue, reflecting both the promise and the sensitivity of LLM adoption in critical environments. In their contribution, Hamid and Brohi provide a review of LLMs in healthcare, discussing major application categories together with threats, vulnerabilities, and security frameworks needed to support safer deployment in real-world medical contexts [
15]. Karami et al. analyze ChatGPT (GPT-3) prompts shared through social media discourse to identify health-related uses of AI chatbots, thus offering an interesting empirical perspective on how users perceive and employ these systems in everyday practice [
16]. Taken together, these papers suggest that healthcare is one of the most promising domains for generative AI, but also one in which concerns about privacy, reliability, accountability, and security must remain central.
The Special Issue also includes studies focused on cybersecurity and risk-sensitive analytical scenarios. Daniel et al. compare machine learning models and LLMs for labeling network intrusion detection system rules with MITRE ATT&CK techniques, showing that LLMs provide interesting opportunities in terms of automation and explainability, while conventional machine learning approaches still maintain advantages in predictive accuracy [
17]. Roumeliotis et al. investigate LLMs and other NLP models for cryptocurrency sentiment analysis, comparing advanced language models for classifying the sentiment of crypto-related news and discussing their relevance for investment intelligence and risk management [
18]. These contributions demonstrate how generative AI is increasingly entering operational and high-impact settings, where effectiveness must be balanced with interpretability, robustness, and informed decision support.
3. Discussion
Although the ten papers address different problems, methods, and domains, several common themes emerge clearly from this collection. First, the Special Issue confirms that one of the most pressing gaps in current LLM research concerns the relationship between generative capability and grounded knowledge. The contributions on RAG, agentic reasoning, and evaluation frameworks show that future progress will depend not only on larger or more capable base models, but also on better integration with external information sources and on stronger methods for assessing factual consistency, transparency, and attribution [
9,
10,
11]. In this respect, the Special Issue contributes to an important ongoing shift in the field: from viewing LLMs as standalone text generators to treating them as components of broader knowledge-intensive systems.
Second, the papers collected here help address another key gap, namely the limited understanding of how LLMs behave in domain-specific, socially relevant, and high-stakes contexts. The studies on healthcare, accessibility, education, cybersecurity, finance, and low-resource language processing show that the value of generative AI cannot be measured exclusively in terms of benchmark performance [
12,
13,
14,
15,
16,
17,
18]. Instead, the practical usefulness of these systems depends on broader qualities such as trustworthiness, safety, interpretability, fairness, privacy protection, cost efficiency, and the degree to which human actors can remain meaningfully involved in the loop. By presenting evidence from multiple domains, this Special Issue helps broaden the discussion beyond purely technical performance and toward the conditions required for responsible real-world adoption.
Third, the Special Issue highlights the importance of analytical, comparative, and review-oriented studies in a field that is evolving extremely quickly. Some of the published papers do not simply propose new systems, but instead offer taxonomies, surveys, bibliometric analyses, and benchmarking frameworks that help organize current knowledge and clarify the state-of-the-art [
11,
14,
15]. This is especially valuable in generative AI research, where rapid innovation can easily outpace conceptual clarity. In this sense, the Special Issue addresses not only technical gaps, but also the need for more structured understanding of where the field currently stands and where it should move next.
4. Future Research Directions
The papers in this Special Issue make clear that future research on generative AI and LLMs should move decisively toward more trustworthy, grounded, and context-aware systems. One major direction concerns the improvement of retrieval and reasoning mechanisms, especially in environments where generated content must remain closely aligned with external evidence. More work is needed on hybrid architectures that combine LLMs with retrieval modules, knowledge graphs, domain repositories, and agentic workflows, as well as on evaluation protocols capable of measuring faithfulness, attribution, and robustness in a reproducible way.
A second priority concerns domain adaptation and responsible deployment in high-impact sectors. The growing use of LLMs in healthcare, education, accessibility, governance, finance, and cybersecurity shows that these systems are increasingly expected to support consequential decisions and sensitive interactions. Future studies should therefore pay greater attention to safety guarantees, privacy-preserving mechanisms, bias mitigation, human oversight strategies, and the design of interfaces that help users understand model limitations and uncertainties. In many cases, progress will likely depend not only on stronger models, but also on interdisciplinary collaboration capable of integrating technical, legal, organizational, and ethical perspectives.
A third important direction is the development of more transparent and interpretable forms of generative AI. As LLMs become more deeply embedded in research and professional workflows, it will be increasingly important to understand features such as how they produce outputs, how they respond to ambiguity, how they fail, and how their behavior can be made more controllable. Comparative studies, benchmark creation, and explanatory frameworks will continue to play a fundamental role in this process. More generally, future research should aim to ensure that generative AI systems are not only more powerful, but also more reliable, explainable, inclusive, and socially responsible.
5. Conclusions
The papers collected in this Special Issue provide a broad and timely perspective on current research in generative AI and large language models. Together, they show a field that is progressing rapidly, but also becoming more mature in the questions it asks and in the criteria by which success should be judged. Beyond raw performance, the studies published here underscore the importance of grounding, explainability, safety, transparency, and domain-aware evaluation.
Overall, this Special Issue contributes to addressing important gaps in the current literature by bringing together methodological advances, application-driven investigations, and analytical studies across a diverse set of contexts. At the same time, it makes evident that many open challenges remain. We believe that future research should focus increasingly on the trustworthy integration of LLMs into real-world systems, on rigorous and reproducible evaluation methodologies, and on approaches that combine technical innovation with human-centered and socially responsible design. We hope that this collection will stimulate further work in these directions and support the development of generative AI systems that are not only more capable, but also more dependable and beneficial.