<?xml version="1.0" encoding="UTF-8"?>
<ArticleSet>
  <Article>
    <Journal>
      <PublisherName></PublisherName>
      <JournalTitle>Journal of Artificial Intelligence, Applications and Innovations</JournalTitle>
      <Issn>3060-7124</Issn>
      <Volume>1</Volume>
      <Issue>Journal of Artificial Intelligence, Application and Inovations </Issue>
      <PubDate PubStatus="epublish">
        <Year>2024</Year>
        <Month>04</Month>
        <Day>02</Day>
      </PubDate>
    </Journal>
    <ArticleTitle>naab: A ready-to-use plug-and-play corpus for Farsi</ArticleTitle>
    <VernacularTitle>naab: A ready-to-use plug-and-play corpus for Farsi</VernacularTitle>
    <FirstPage>1</FirstPage>
    <LastPage>8</LastPage>
    <ELocationID EIdType="doi">10.61838/jaiai.1.2.1</ELocationID>
    <Language>EN</Language>
    <AuthorList>
      <Author>
        <FirstName></FirstName>
        <LastName></LastName>
        <Affiliation></Affiliation>
      </Author>
      <Author>
        <FirstName></FirstName>
        <LastName></LastName>
        <Affiliation></Affiliation>
      </Author>
      <Author>
        <FirstName></FirstName>
        <LastName></LastName>
        <Affiliation></Affiliation>
      </Author>
      <Author>
        <FirstName></FirstName>
        <LastName></LastName>
        <Affiliation></Affiliation>
      </Author>
    </AuthorList>
    <PublicationType>Journal Article</PublicationType>
    <History>
      <PubDate PubStatus="received">
        <Year>2023</Year>
        <Month>11</Month>
        <Day>10</Day>
      </PubDate>
    </History>
    <Abstract>&lt;p&gt;The rise of large language models (LLMs) has transformed numerous natural language processing (NLP) tasks, yet their performance in low and mid-resource languages, such as Farsi, still lags behind resource-rich languages like English. To address this gap, we introduce Naab, the largest publicly available, cleaned, and ready-to-use Farsi textual corpus. Naab consists of 130GB of data, comprising over 250 million paragraphs and 15 billion words. Named after the Farsi word ناب (meaning "pure" or "high-grade"), this corpus is openly accessible via Hugging Face, offering researchers a valuable resource for Farsi NLP tasks. In addition to naab, we provide naab-raw, an unprocessed version of the dataset, along with a pre-processing toolkit that allows users to clean their custom corpora. These resources empower NLP researchers and practitioners, particularly those focusing on low-resource languages, to improve the performance of LLMs in their respective domains and bridge the gap between resource-rich and resource-poor languages.&lt;/p&gt;</Abstract>
    <ObjectList>
      <Object Type="keyword">
        <Param Name="value">Natural Language Processing</Param>
      </Object>
      <Object Type="keyword">
        <Param Name="value">Low-resource Languages</Param>
      </Object>
      <Object Type="keyword">
        <Param Name="value">Large Language Models</Param>
      </Object>
      <Object Type="keyword">
        <Param Name="value">Textual Corpus</Param>
      </Object>
      <Object Type="keyword">
        <Param Name="value">Open-source Dataset</Param>
      </Object>
      <Object Type="keyword">
        <Param Name="value">Data Preprocessing</Param>
      </Object>
      <Object Type="keyword">
        <Param Name="value">Persian Language Resources</Param>
      </Object>
      <Object Type="keyword">
        <Param Name="value">Text Mining</Param>
      </Object>
    </ObjectList>
    <ArchiveCopySource DocType="pdf">https://www.journalaiai.com/index.php/aiai/article/download/7/7</ArchiveCopySource>
  </Article>
</ArticleSet>
