Photographing and storing Arabic books, newspapers and manuscripts in digital files is not enough to make their contents searchable or usable in artificial intelligence systems. To a computer, a scanned page remains an image, not text that can be indexed and analysed.
The gap in converting Arab archives into machine-readable data
While the SARD dataset, published in September 2026, reveals the scale of resources needed to teach machines to read Arabic pages, researchers warn that real historical documents present challenges extending beyond character recognition to reconstructing page structure, documenting sources and verifying the accuracy of extracted text. Arab libraries and archives hold a vast record of books, documents, newspapers and manuscripts that have survived damage or loss.
Some of these materials have emerged from isolation on library shelves after being scanned and stored digitally. But moving them from paper to screen does not necessarily mean that they have been brought into the digital environment in which search engines and computer models operate. A search engine cannot find a word inside an image of a page in the way it searches digital text, nor can an artificial intelligence model process the image as it would machine-readable content.
This creates a gap between preserving knowledge in digital files and converting it into data that systems can index, read, analyse and use. The gap widens when dealing with old manuscripts, damaged newspapers or documents written in multiple Arabic scripts.
OCR and HTR technologies for processing Arabic pages
In such cases, scanning does not complete the task. Technologies are needed to recognise text, understand page layout and correct errors, then link the extracted content to its historical context and metadata. Converting paper into material usable by computers involves optical character recognition, commonly abbreviated as OCR.
Documents and manuscripts written by hand require handwriting recognition technologies, or HTR. Training both types requires large and varied quantities of images and accurate reference texts. The SARD dataset illustrates the scale of these requirements.
Researchers from Prince Sultan University and Saudi Arabia’s Tuwaiq Academy published it in September 2026. It contains more than 2.6 million images of Arabic pages containing approximately 794.6 million words in total. The dataset was designed to train and test Arabic text-recognition systems using pages that imitate book layouts.
SARD is not a historical archive of manuscripts or documents that have been rescued and digitised, but a synthetic dataset. In compiling it, the researchers used digital Arabic texts from Alukah Library and converted them into images of pages, while retaining the original text as an accurate reference for training models and measuring their performance.
This method makes it possible to produce large numbers of page images paired with known reference texts, without manually transcribing every page. But it does not resolve the challenge of moving from pages generated in a controlled environment to the actual documents held in libraries and archives, with their damage, typographical variations and complex layouts.
SARD researchers explain the limits of synthetic data
Wadie Boualila, director of the Robotics and Internet of Things Laboratory at Prince Sultan University in Riyadh and one of the researchers involved in building SARD, said the main motivation for creating the dataset was to close a gap in the data needed to develop Arabic document-reading technologies.
He said earlier datasets had made important contributions, but many focused on individual characters, words or lines, or remained limited in size, typographical variety and representation of full pages. Boualila added that the dataset aims to provide data that helps models read complete pages resembling printed book pages, while accounting for differences in fonts and text density.
Synthetic generation made it possible to produce data on this scale, linking each image to the reference text used to create it. This reduces the cost of manual transcription and lays the groundwork for developing optical character recognition systems and models that combine vision and language capabilities.
Boualila nevertheless stressed that SARD does not represent the real Arab archive in all its complexity. It simulates book pages in controlled formats and does not include every form of complex layout or the defects found in old documents. Success on this dataset therefore cannot be regarded as sufficient evidence that a system can read actual historical materials.
According to Boualila, transferring these systems into real-world use requires further training and evaluation with data taken directly from genuine documents. He called for cooperation with libraries and archival centres to collect varied samples, prepare reference texts after specialist review, and adapt models to handle the characteristics of those materials.
Training can be strengthened by simulating the printing and imaging defects commonly found in archives, but Boualila said testing should also be conducted on genuine documents not included in the training data. Test materials should preferably come from multiple sources, to determine whether the model can handle new documents rather than only patterns it has previously learned.
In his view, the challenge is not limited to distinguishing characters and words. It also includes understanding page structure and column order, distinguishing headings, footnotes and main text, and linking each element to its position and function. Manuscripts require specialised data and expertise in handwriting and historical scripts, he said, stressing that “success in reading printed text does not automatically mean success in reading manuscripts”.
The Robotics and Internet of Things Laboratory team believes that developing this field requires combining artificial intelligence tools with the expertise of language and heritage specialists. This includes subjecting passages the system is uncertain about to human review, while always keeping the extracted text linked to the page image from which it was taken.
According to Boualila, the goal should not be confined to extracting text from millions of pages. It should extend to making Arab heritage available for research and knowledge without compromising the fidelity of its transmission or losing its historical details. This makes extraction quality and documentation an essential part of digitisation, not an optional additional stage. Actual old documents reveal another level of difficulty.
Damaged documents complicate the reading of structure and historical context
Hamdi Mubarak, principal engineer at the Qatar Computing Research Institute, said reading these documents was not merely a matter of recognising characters, but of attempting to reconstruct a complete document that may have lost some elements of its original structure during digitisation.
Mubarak explained that the condition of the material itself directly affects system performance, from poor scan quality and yellowing or damage to the paper, through ink stains and irregular pages, to connected Arabic letters, the absence of diacritics, and the variety of scripts and old printing styles. Each of these factors can lead to text-recognition errors. Newspapers present additional challenges because of their multiple columns, headlines, advertisements and footnotes, as well as the intermingling of images and text.
It is therefore not enough for a machine to recognise individual words. It must also determine the correct reading order and understand the relationship between page elements in order to reconstruct the content in a form close to the original. Errors are particularly consequential in Arabic, Mubarak said, because changing a single letter can turn a word into another grammatically valid Arabic word with an entirely different meaning.
Old texts may also use spellings and writing styles that do not match modern usage. The system may successfully read a word visually but misinterpret it or place it in the wrong historical context. Digitisation does not end with producing a searchable text file.
Mubarak stressed that text extraction is only the beginning of the process. Before content is entered into training data or information-retrieval systems, it must be cleaned, described, quality-controlled and checked to ensure that its elements are in the correct order. It is essential, he said, to preserve the link between the extracted text and the original page and to correct the reading order, particularly in newspapers with multiple columns.
Footnotes, headlines and captions should also be separated from the main body, while source data should be preserved, including the name of the book or newspaper, the author’s name, the date, edition and page number. Mubarak warned against treating OCR output as “fact” simply because it has been converted into searchable text.
He called for a confidence or quality level to be assigned to each page or passage, with low-quality samples subjected to human review, particularly if the texts are to be used to create datasets that will later train other models. The effects of OCR errors can extend beyond the original document.
A person’s name or a place name may be changed into a different word, figures and dates may be altered, columns may be mixed up, or an entire sentence may be omitted because of an error in analysing the page layout. Mubarak said that when such errors are repeated across millions of pages, they may become part of the data from which models learn.
In information-retrieval systems, an error in a name, place or term can prevent a document from appearing in search results, while generative models may learn the corrupted text and treat it as the correct form.
For this reason, Mubarak believes that the success of digitising Arab memory should not be measured only by the number of pages converted into text. It should also be measured by how closely the extracted text matches the source, whether it retains the document’s structure, context and creator information, and whether the confidence level of each extraction is known. At the same time, artificial intelligence can be used to verify results by detecting OCR errors, comparing extracted text with the page image and checking its consistency with the context.
But this automated review remains dependent on retaining the original image alongside the text, so that the results can be examined by people and verified later. The SARD experiment shows the amount of data needed to teach systems to read Arabic pages resembling printed books, while real documents show that extracting the text is not the end of the task.
Usable digitisation requires protecting page structure, documenting the source and context, measuring extraction quality and maintaining a direct link between digital text and its scanned original, so that Arab memory becomes searchable and analysable without losing its historical integrity.