<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE ArticleSet PUBLIC "-//NLM//DTD PubMed 2.7//EN" "https://dtd.nlm.nih.gov/ncbi/pubmed/in/PubMed.dtd">
<ArticleSet>
<Article>
<Journal>
				<PublisherName>Interdisciplinary Quranic Studies Research Institute, Shahid Beheshti University
&amp;
Contemporary Thoughts Press</PublisherName>
				<JournalTitle>Journal of Interdisciplinary Qur'anic Studies</JournalTitle>
				<Issn>2976-6982</Issn>
				<Volume>5</Volume>
				<Issue>1</Issue>
				<PubDate PubStatus="epublish">
					<Year>2026</Year>
					<Month>03</Month>
					<Day>21</Day>
				</PubDate>
			</Journal>
<ArticleTitle>Enhancing Arabic Word Stemming Using the Char Stemmer Deep Learning Model</ArticleTitle>
<VernacularTitle>بهبود ریشه‌یابی کلمات عربی با استفاده از مدل یادگیری عمیق Char Stemmer</VernacularTitle>
			<FirstPage>5</FirstPage>
			<LastPage>25</LastPage>
			<ELocationID EIdType="pii">107215</ELocationID>
			
<ELocationID EIdType="doi">10.37264/JIQS.V5I1.6</ELocationID>
			
			<Language>EN</Language>
<AuthorList>
<Author>
					<FirstName>Azal</FirstName>
					<LastName>Alaswaad</LastName>
<Affiliation>PhD Candidate, Department of Computer Engineering, Iran University of Science and Technology, Tehran, Iran</Affiliation>

</Author>
<Author>
					<FirstName>Behrouz</FirstName>
					<LastName>Minaei-Bidgoli</LastName>
<Affiliation>Professor, Department of Computer Engineering, Iran University of Science and Technology, Tehran, Iran</Affiliation>
<Identifier Source="ORCID">0000-0002-9327-7345</Identifier>

</Author>
</AuthorList>
				<PublicationType>Journal Article</PublicationType>
			<History>
				<PubDate PubStatus="received">
					<Year>2026</Year>
					<Month>02</Month>
					<Day>17</Day>
				</PubDate>
			</History>
		<Abstract>The stemming of Arabic words is a crucial step in several text processing tasks, including text mining, information retrieval, and natural language processing. Arabic stemmers face many challenges, mainly due to the complex nature of Arabic words and their different writing styles. To address these challenges, this paper presents a novel approach integrated into a Python module for improving conventional Arabic stemming. The proposed approach employs neural network architectures to achieve high accuracy in Arabic text stemming. We utilize a Bi-LSTM architecture to extract Arabic stems and construct sequence-to-sequence stemming models. The proposed model is named Char Stemmer. To evaluate this model, we use a benchmark dataset of Qur’anic words and classical Arabic texts. Extensive experiments were conducted to evaluate the model, including comparisons with several well-established Arabic stemmers, namely Khoja, Light-10, P-Stemmer, Tashaphyne, and Al-Khalil. The experimental results show that the proposed model achieves a stem accuracy of 93.88%, substantially outperforming conventional rule-based and light stemming methods. The results demonstrate the effectiveness of deep learning sequence models for Arabic stemming and highlight the potential of character-level models for capturing complex morphological structures. The proposed model is generalizable, scales well, and can serve as a domain-independent solution for a wide range of Arabic NLP applications. Furthermore, the model can be directly integrated into Arabic information systems for tasks such as automated text normalization, intelligent search, and knowledge extraction in applied computing environments.</Abstract>
			<OtherAbstract Language="FA">ریشه‌یابی واژگان عربی یکی از مراحل اساسی در بسیاری از وظایف پردازش متن، از جمله متن‌کاوی، بازیابی اطلاعات و پردازش زبان طبیعی است. با این حال، ریشه‌یاب‌های زبان عربی به دلیل ساختار پیچیده واژگان این زبان و تنوع شیوه‌های نگارش آن با چالش‌های متعددی روبه‌رو هستند. برای غلبه بر این چالش‌ها و بهبود روش‌های متداول ریشه‌یابی واژگان عربی، این مقاله رویکردی نوین را که در قالب یک ماژول پایتون پیاده‌سازی شده است، ارائه می‌کند. در این رویکرد، از معماری‌های شبکه عصبی به‌منظور دستیابی به دقت بالاتر در ریشه‌یابی متون عربی استفاده شده است. بدین منظور، معماری Bi-LSTM برای استخراج ریشه واژگان عربی و ساخت مدل‌های ریشه‌یابی مبتنی بر توالی‌به‌توالی (Seq2Seq) به کار گرفته شده و مدل پیشنهادی Char Stemmer نام‌گذاری شده است. برای ارزیابی عملکرد این مدل، از یک مجموعه‌داده مرجع شامل واژگان قرآن کریم و متون عربی کلاسیک استفاده شد. همچنین آزمایش‌های گسترده‌ای با هدف مقایسه عملکرد مدل پیشنهادی با چندین ریشه‌یاب شناخته‌شده زبان عربی، شامل Khoja، Light-10، P-Stemmer، Tashaphyne و Al-Khalil انجام گرفت. نتایج نشان داد که مدل پیشنهادی با دستیابی به دقت 93.88 درصد در استخراج ریشه واژگان، عملکردی به‌مراتب بهتر از روش‌های متداول مبتنی بر قواعد و ریشه‌یابی سبک ارائه می‌دهد. یافته‌ها بیانگر کارایی بالای مدل‌های توالی مبتنی بر یادگیری عمیق در ریشه‌یابی زبان عربی و نیز توانایی مدل‌های مبتنی بر نویسه در شناسایی ساختارهای صرفی پیچیده این زبان است. افزون بر این، مدل پیشنهادی از قابلیت تعمیم‌پذیری و مقیاس‌پذیری مناسبی برخوردار است و می‌تواند به‌عنوان راهکاری مستقل از حوزه کاربرد، در طیف گسترده‌ای از سامانه‌های پردازش زبان طبیعی عربی مورد استفاده قرار گیرد. همچنین این مدل قابلیت ادغام مستقیم با سامانه‌های اطلاعاتی زبان عربی را برای کاربردهایی همچون هنجارسازی خودکار متون، جست‌وجوی هوشمند و استخراج دانش در محیط‌های محاسباتی داراست.&lt;br /&gt;&lt;br /&gt;</OtherAbstract>
		<ObjectList>
			<Object Type="keyword">
			<Param Name="value">Qur’anic Arabic</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Qur’anic corpus</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Computational Qur’anic Studies</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Arabic Natural Language Processing</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Seq2Seq Models</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Stemming</Param>
			</Object>
			<Object Type="keyword">
			<Param Name="value">Deep Learning Architectures</Param>
			</Object>
		</ObjectList>
<ArchiveCopySource DocType="pdf">https://iqs.sbu.ac.ir/article_107215_1fc6ad3a63a8c84a1b307a6d36df353b.pdf</ArchiveCopySource>
</Article>
</ArticleSet>
