Enhancing the Interpretability of NLI Models Using LLMs and Active Learning Algorithms
Abstract
1. Introduction
- This paper proposes a dual-model architecture, EGM–PM, tailored for low-resource scenarios. By replacing manual annotation with LLMs, this framework reduces annotation costs and jointly improves predictive performance and interpretability.
- This paper proposes ETCM, an explanation-enhanced active learning sampling method that progressively integrates explanation semantics to improve diversity-driven selection in the cold-start phase and introduces an explanation-discrepancy criterion to complement uncertainty signals, thereby increasing the semantic complementarity and information gain of the queried set.
- This paper constructs ExNLI, a cross-lingual NLI dataset enriched with LLM-generated natural language explanations. This dataset serves as a benchmark for evaluating our framework and facilitates research on explainable multilingual inference.
2. The Research Problem and Related Works
2.1. Related Works
2.1.1. Explainable Artificial Intelligence
2.1.2. Active Learning
2.1.3. Large Language Models and Prompt Engineering
2.2. The Research Problem and Research Questions
- RQ1, Annotation efficiency. How can we obtain scalable, low-cost natural language explanations while ensuring that their quality is controllable and that they are sufficient to serve as effective supervision signals?
- RQ2, Explanation-guided sample selection. Under annotation budgets, how can explanation semantics be integrated into active sample selection to approach full-data performance with fewer labeled instances?
- RQ3, Explanation-assisted inference. How can LLM-annotated explanations be leveraged as auxiliary signals to improve NLI inference compared with explanation-free baselines?
3. Materials and Methods
3.1. ETCM—Active Learning Sampling Algorithm
| Algorithm 1 TypiClust–Explanation |
|
| Algorithm 2 ETCM |
|
3.2. EGM–PM: A Dual Model Active Learning Framework
- Explanation-Generation Model
- Prediction Model
- Active Learning Sampler
3.3. Natural Language Explanation Generation
3.4. Construction of the ExNLI Dataset
3.4.1. Balanced Sampling, Translation, and Explanation Annotation
3.4.2. Data Selection and Splits
4. Experimental Setup and Results
4.1. Preliminary Experiment to Determine Candidate Pool Size
4.2. Impact of Shot Number and Example Selection
4.3. Hyperparameter Sensitivity Analysis Experiment
4.4. Comparative Experiment of Active Learning Sampling Algorithms
4.5. Dataset Quality Evaluation
4.5.1. Human Evaluation of Explanation Quality
- Does the explanation accurately reflect the logical relationship between the premise and hypothesis (Entailment, Contradiction, Neutral)?
- Is the language of the explanation clear and easy to understand?
- Is the explanation primarily grounded in the premise and hypothesis, and free of obviously fabricated information that contradicts the text?
- Is the explanation itself factually correct?
- Regardless of the AI’s correctness, does the explanation help you understand why the premise and hypothesis have the given relationship?
- Would you trust and use this AI in a real-world scenario?
4.5.2. Automated Evaluation
4.6. Cross-Task Transfer Experiment: Evaluation on Question Answering
4.7. Ablation Experiment and Model Comparison
4.7.1. Ablation Study on the Explanation Components in ETCM
- TCM/Text-Only (): clustering and selection are performed using only the text representation;
- Static Mix (): an equal-weight fusion of text and explanations is used throughout;
- ETCM/Dynamic Warm-up (Max = 0.5): a dynamic linear warm-up is applied, such that increases progressively with iterations and eventually stabilizes at .
4.7.2. Ablation Study on the ETCM and NLEs
- (a)
- PM (FLAN-T5 220 M): The prediction model is trained without active learning processing, without the ETCM sampling algorithm, and without supervision from natural language explanations generated by the EGM.
- (b)
- EGM-PM (Random 220 M): The EGM model is incorporated to generate natural language explanations, which are used as part of the input to the prediction model for explanation-based supervision. Active learning is employed with Random sampling to select data for annotation. This setting aims to evaluate the contribution of the EGM and the effectiveness of natural language explanations in the NLI task.
- (c)
- EGM-PM (ETCM 220 M): The EGM-generated explanations are similarly integrated into the prediction model as explanatory input, while the proposed ETCM algorithm is used for active learning sample selection. This variant is designed to assess the efficacy of the ETCM sampling strategy.
4.7.3. Comparison with Other Methods
5. Discussion
- Efficiency of Explanation Supervision in Low-Resource Settings. As shown in Table 10, under the same low-resource setting, EGM–PM with 220 M parameters in EGM-PM outperforms the substantially larger 780 M-parameter FLAN-T5. This result indicates that, when data are limited, the quality of the supervision signal plays a critical role. Standard NLI fine-tuning implicitly relies on the model to infer logical patterns from labels; in contrast, explanations generated by EGM provide explicit intermediate reasoning steps. This structure functions as a reasoning scaffold, enabling a smaller model to align with task logic more efficiently than a larger model trained with sparse label supervision alone.
- Dynamic Integration of Explanation Semantics. An ablation study on the fusion weight (Figure 9) shows that a progressive “dynamic warm-up” strategy outperforms static fusion. This observation highlights that explanation semantics are sensitive to the stage of learning and inference. During the early “cold-start” phase, EGM is trained on only a handful of samples, so the resulting explanation embeddings may be noisy or overly generic. Consequently, over-reliance on explanations for clustering in the initial rounds can be detrimental. ETCM derives its effectiveness from a dynamic adjustment mechanism: it prioritizes textual diversity when explanation quality is uncertain, and gradually incorporates explanation semantics as EGM stabilizes, thereby improving the trajectory of active learning.
- Complementarity between Semantic Discrepancy and Uncertainty. A sensitivity analysis of (Table 9) suggests that the explanation discrepancy () captures information that differs from probabilistic uncertainty. Uncertainty sampling typically targets instances near the decision boundary (e.g., low-confidence cases), whereas identifies instances that exhibit semantically novel reasoning patterns relative to the labeled set. The performance drop observed at high (e.g., ) indicates that these two signals should be balanced rather than treated as substitutes. Our findings suggest that NLEs can serve as an effective orthogonal signal that complements, rather than replaces, conventional uncertainty measures in active sampling.
6. Limitations
- Subjectivity and Scope of Human Evaluation: Human evaluation is inherently susceptible to cognitive biases. Although we implemented rigorous controls to mitigate these effects—including double-blind annotation, randomized ordering, and inter-rater agreement checks (ICC/Kappa)—subjectivity cannot be entirely eliminated. Annotators may still unconsciously favor fluent or verbose explanations while discounting concise ones. Moreover, due to resource constraints, our human evaluation was conducted on a stratified subset. This limited scope may fail to fully capture the model’s behavior on long-tail instances, motivating more comprehensive, large-scale evaluation protocols in future work.
- Logical Validity of Generated Explanations:Although EGM is conditioned on ground-truth labels, this does not fully guarantee that the generated explanations are always logically sound. In some cases, the LLM may rely on tautological paraphrasing or hallucinated details not supported by the text to justify a given label. Such “plausible but ungrounded” explanations constitute noisy supervision signals; while our current framework partially mitigates this issue through active selection, it does not explicitly filter them out.
- Computational Overhead in Sampling: Although EGM–PM reduces human annotation cost, ETCM introduces additional computational overhead. Computing explanation embeddings for the full candidate pool requires encoder inference, which is computationally more expensive than simple uncertainty sampling.
- Dependence on Base Model Capability: Our method assumes that the base LLM possesses a minimum level of reasoning capability. In extremely low-resource languages where the base model performs poorly, the generated explanations may be too noisy to provide effective guidance.
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Saarela, M.; Podgorelec, V. Recent applications of Explainable AI (XAI): A systematic literature review. Appl. Sci. 2024, 14, 8884. [Google Scholar] [CrossRef]
- Lipton, Z.C. The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue 2018, 16, 31–57. [Google Scholar] [CrossRef]
- Bowman, S.; Angeli, G.; Potts, C.; Manning, C.D. A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lisbon, Portugal, 17–21 September 2015; pp. 632–642. [Google Scholar] [CrossRef]
- Williams, A.; Nangia, N.; Bowman, S. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), Orleans, LA, USA, 1–6 June 2018; pp. 1112–1122. [Google Scholar] [CrossRef]
- Conneau, A.; Rinott, R.; Lample, G.; Schwenk, H.; Stoyanov, V.; Williams, A.; Bowman, S.R. XNLI: Evaluating cross-lingual sentence representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018, Brussels, Belgium, 31 October–4 November 2018; pp. 2475–2485. [Google Scholar] [CrossRef]
- Hu, H.; Richardson, K.; Xu, L.; Li, L.; Kübler, S.; Moss, L.S. OCNLI: Original Chinese Natural Language Inference. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16–20 November 2020; pp. 3512–3526. [Google Scholar]
- Camburu, O.M.; Rocktäschel, T.; Lukasiewicz, T.; Blunsom, P. e-snli: Natural language inference with natural language explanations. In Proceedings of the Advances in Neural Information Processing Systems 31, Montréal, QC, Canada, 3–8 December 2018. [Google Scholar]
- Zhang, S.; Gong, C.; Liu, X.; He, P.; Chen, W.; Zhou, M. ALLSH: Active Learning Guided by Local Sensitivity and Hardness. In Proceedings of the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Washington, DC, USA, 10–15 July 2022. [Google Scholar] [CrossRef]
- Barredo Arrieta, A.; Díaz-Rodríguez, N.; Del Ser, J.; Bennetot, A.; Tabik, S.; Barbado, A.; García, S.; Gil-López, S.; Molina, D.; Benjamins, R.; et al. Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef]
- Liu, G.; Zhang, J.; Chan, A.B.; Hsiao, J.H. Human attention guided explainable artificial intelligence for computer vision models. Neural Netw. 2024, 177, 106392. [Google Scholar] [CrossRef] [PubMed]
- Quan, X.; Valentino, M.; Dennis, L.; Freitas, A. Verification and Refinement of Natural Language Explanations through LLM-Symbolic Theorem Proving. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, 12–16 November 2024; pp. 2933–2958. [Google Scholar] [CrossRef]
- Gurrapu, S.; Kulkarni, A.; Huang, L.; Lourentzou, I.; Batarseh, F.A. Rationalization for explainable NLP: A survey. Front. Artif. Intell. 2023, 6, 1225093. [Google Scholar] [CrossRef] [PubMed]
- Popovič, N.; Färber, M. Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward Pass. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, 4–9 November 2025; pp. 31692–31705. [Google Scholar] [CrossRef]
- Hong, P.; Chen, B.; Peng, S.; de Marneffe, M.C.; Plank, B. LiTEx: A Linguistic Taxonomy of Explanations for Understanding Within-Label Variation in Natural Language Inference. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, 4–9 November 2025; pp. 34065–34085. [Google Scholar] [CrossRef]
- Zong, C.C.; Wang, Y.W.; Ning, K.P.; Ye, H.B.; Huang, S.J. Bidirectional Uncertainty-Based Active Learning for Open-Set Annotation. In Proceedings of the European Conference on Computer Vision (ECCV 2024), Milan, Italy, 29 September–4 October 2024; Part XXVIII. Springer: Cham, Switzerland, 2024; pp. 127–143. [Google Scholar] [CrossRef]
- Sener, O.; Savarese, S. Active Learning for Convolutional Neural Networks: A Core-Set Approach. arXiv 2017, arXiv:1706.03762. [Google Scholar]
- Hacohen, G.; Dekel, A.; Weinshall, D. Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets. In Proceedings of the 39th International Conference on Machine Learning (ICML 2022), Baltimore, MD USA, 17–23 July 2022; Chaudhuri, K., Jegelka, S., Song, L., Szepesvári, C., Niu, G., Sabato, S., Eds.; Proceedings of Machine Learning Research. PMLR: Cambridge, MA, USA, 2022; Volume 162, pp. 8175–8195. [Google Scholar]
- Doucet, P.; Estermann, B.; Aczel, T.; Wattenhofer, R. Bridging Diversity and Uncertainty in Active learning with Self-Supervised Pre-Training. In Proceedings of the 5th Workshop on Practical ML for Limited/Low Resource Settings (PML4LRS) @ ICLR 2024, Vienna, Austria, 11 May 2024; OpenReview: Alameda, CA, USA, 2024. [Google Scholar] [CrossRef]
- Yao, B.; Jindal, I.; Popa, L.; Katsis, Y.; Ghosh, S.; He, L.; Lu, Y.; Srivastava, S.; Hendler, J.A.; Wang, D. Beyond labels: Empowering human with natural language explanations through a novel active-learning architecture. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, 6–10 December 2023; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023. [Google Scholar] [CrossRef]
- Liu, P.; Yuan, W.; Fu, J.; Jiang, Z.; Hayashi, H.; Neubig, G. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput. Surv. 2023, 55, 1–35. [Google Scholar] [CrossRef]
- Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.H.; Le, Q.V.; Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. In Proceedings of the Advances in Neural Information Processing Systems 35 (NIPS ’22), Red Hook, NY, USA, 28 November–9 December 2022. [Google Scholar]
- Dhaini, M.; Vladika, J.; Erdogan, E.; Attaoui, Z.; Kasneci, G. Can LLM-Generated Textual Explanations Enhance Model Classification Performance? An Empirical Study. In Proceedings of the Artificial Neural Networks and Machine Learning—ICANN 2025, Kaunas, Lithuania, 9–12 September 2025; Senn, W., Sanguineti, M., Saudargiene, A., Tetko, I.V., Villa, A.E.P., Jirsa, V., Bengio, Y., Eds.; Lecture Notes in Computer Science. Springer: Cham, Switzerland, 2026; Volume 16070, pp. 192–204. [Google Scholar] [CrossRef]
- Pham, D.H.; Le, T.; Nguyen, H.T. How rationals boost textual entailment modeling: Insights from large language models. Comput. Electr. Eng. 2024, 119, 109517. [Google Scholar] [CrossRef]
- Heredia, M.; Etxaniz, J.; Zulaika, M.; Saralegi, X.; Barnes, J.; Soroa, A. XNLIeu: A dataset for cross-lingual NLI in Basque. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, 16–21 June 2024; pp. 4177–4188. [Google Scholar] [CrossRef]
- Htet, A.K.; Dras, M. Myanmar XNLI: Building a dataset and exploring low-resource approaches to natural language inference with Myanmar. Lang. Resour. Eval. 2025, 59, 3267–3310. [Google Scholar] [CrossRef]
- Reimers, N.; Gurevych, I. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, 3–7 November 2019; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 3982–3992. [Google Scholar]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 2020, 21, 5485–5551. [Google Scholar]
- Chung, H.W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al. Scaling Instruction-Finetuned Language Models. J. Mach. Learn. Res. 2024, 25, 3381–3433. [Google Scholar]
- Hui, B.; Yang, J.; Cui, Z.; Yang, J.; Liu, D.; Zhang, L.; Liu, T.; Zhang, J.; Yu, B.; Lu, K.; et al. Qwen2.5-coder technical report. arXiv 2024, arXiv:2409.12186. [Google Scholar] [CrossRef]
- Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X.; et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 2025, 645, 633–638. [Google Scholar] [CrossRef] [PubMed]
- Feng, F.; Yang, Y.; Cer, D.; Arivazhagan, N.; Wang, W. Language-agnostic BERT Sentence Embedding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, 22–27 May 2022; pp. 878–891. [Google Scholar] [CrossRef]
- Rajpurkar, P.; Zhang, J.; Lopyrev, K.; Liang, P. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, TX, USA, 1–4 November 2016; pp. 2383–2392. [Google Scholar] [CrossRef]
- Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
- Hsieh, C.Y.; Li, C.L.; Yeh, C.K.; Nakhost, H.; Fujii, Y.; Ratner, A.J.; Krishna, R.; Lee, C.Y.; Pfister, T. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. In Proceedings of the the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, ON, Canada, 9–14 July 2023. [Google Scholar] [CrossRef]









| Prompt Types | Content |
|---|---|
| Task Description | Please provide an explanation based on the given premise, hypothesis, and their relationship. The explanation should accurately reflect the logical connection between the premise-hypothesis pairs, be clear and easy to understand, avoid vagueness or redundancy, and stay faithful to logical reasoning without fabricating false reasons. |
| Background Knowledge | Natural Language Inference is a common task in the field of Natural Language Processing. It requires computers to understand the relationship between two sentences. Simply put, it asks the computer to determine whether a sentence can be reasonably inferred from another sentence. The possible logical relations between the two are: Entailment, Contradiction, or Neutral. |
| Example Data | Premise: The church choir is singing joyful songs inside the church to the crowd. Hypothesis: The church is very quiet. Label: contradiction Explanation: Since the choir is singing in the church, it cannot be quiet because there is singing. … (examples of the other two types of relationships) |
| Requirements | Avoid redundancy: Explanations should be concise and must not contain repetitive or irrelevant information. Prohibit anaphora: Do not use vague pronouns such as “it” or “this”, and do not directly quote the premise or hypothesis; the explanation itself must be a complete sentence. Maintain consistency: Explanations must remain logically consistent with the premise and hypothesis, avoiding contradictions. Faithful reasoning: Explanations should be based solely on the logical relationship between the premise and hypothesis, without introducing additional assumptions. Formal style: Responses should be written in complete natural language sentences, avoiding bullet points or colloquial expressions. |
| Dataset | Entailment | Contradiction | Neutral | Total |
|---|---|---|---|---|
| Dev Set | 825 | 825 | 840 | 2490 |
| Test Set | 1665 | 1665 | 1680 | 5010 |
| Total | 2500 | 2500 | 2500 | 7500 |
| Dataset | Entailment | Contradiction | Neutral | Total |
|---|---|---|---|---|
| Training Set | 16,891 | 16,622 | 16,487 | 50,000 |
| Dev Set | 1000 | 1000 | 1000 | 3000 |
| Total | 17,891 | 17,622 | 17,487 | 53,000 |
| Dataset | Entailment | Contradiction | Neutral | Total |
|---|---|---|---|---|
| Train Set | 2542 | 2542 | 2557 | 7641 |
| Dev Set | 283 | 283 | 283 | 849 |
| Test Set | 2665 | 2665 | 2680 | 8010 |
| Total | 5490 | 5490 | 5520 | 16,500 |
| Hyperparameter | Explanation Generation Model | Prediction Model |
|---|---|---|
| Learning Rate | 0.0001 | 0.0001 |
| Optimizer | Adam | AdamW |
| Training Epochs | 25 | 20 |
| Dropout | 0.1 | 0.1 |
| Batch Size | 4 | 4 |
| Seed | 42 | 42 |
| Max Input Length | 512 | 512 |
| Max Output Length | 128 | N/A |
| Decoding Strategy | Beam Search (beam size = 4) | N/A |
| Epoch-EGM | Epoch-PM | Accuracy |
|---|---|---|
| 10 | 20 | 0.756 |
| 20 | 20 | 0.789 |
| 25 | 20 | 0.819 |
| 30 | 20 | 0.816 |
| 25 | 100 | 0.818 |
| 10 | 150 | 0.761 |
| Question | Average Score 1 | ICC Score | Fleiss’ Kappa |
|---|---|---|---|
| Q1 | 91.93 | 0.92 | 0.90 |
| Q2 | 94.83 | 0.93 | 0.91 |
| Q3 | 94.89 | 0.94 | 0.93 |
| Q4 | 100.00 | 1.00 | 1.00 |
| Q5 | 92.68 | 0.89 | 0.88 |
| Q6 | 92.57 | 0.87 | 0.89 |
| Model | SQuAD1.1 | |
|---|---|---|
| EM | F1 | |
| PM | 85.44 | 92.08 |
| EGM-PM | 87.68 | 93.89 |
| Accuracy@Round15 | |
|---|---|
| 0 | 0.623 |
| 0.1 | 0.632 |
| 0.5 | 0.667 |
| 1.0 | 0.689 |
| 2.0 | 0.667 |
| Model | Architecture | Accuracy |
|---|---|---|
| RoBERTa (125 M) | Encoder-only | 44.08 |
| PM FLAN-T5 (220 M) | Encoder-Decoder | 56.32 |
| PM FLAN-T5 (780 M) | Encoder-Decoder | 57.29 |
| Distilling Step by Step | Encoder-Decoder | 75.38 |
| EGM-PM(Random) (220 M) | Encoder-Decoder | 82.13 |
| Pham et al. [23] | Encoder-Decoder | 83.93 |
| Yao et al. [19] | Encoder-Decoder | 87.02 |
| EGM-PM(ETCM) (220 M) | Encoder-Decoder | 88.89 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Q.; Liu, J. Enhancing the Interpretability of NLI Models Using LLMs and Active Learning Algorithms. Information 2026, 17, 119. https://doi.org/10.3390/info17020119
Wang Q, Liu J. Enhancing the Interpretability of NLI Models Using LLMs and Active Learning Algorithms. Information. 2026; 17(2):119. https://doi.org/10.3390/info17020119
Chicago/Turabian StyleWang, Qi, and Junqiang Liu. 2026. "Enhancing the Interpretability of NLI Models Using LLMs and Active Learning Algorithms" Information 17, no. 2: 119. https://doi.org/10.3390/info17020119
APA StyleWang, Q., & Liu, J. (2026). Enhancing the Interpretability of NLI Models Using LLMs and Active Learning Algorithms. Information, 17(2), 119. https://doi.org/10.3390/info17020119
