Abstract
Owing to the significance of DNA storage technology in meeting exponential storage demands and longevity, the challenges caused by bio-molecular errors while reading/sequencing data from DNA molecules must be addressed. By reading redundant copies, data can be reconstructed but with associated cost of sequencing and decoding complexities. Hence, solutions for dealing with both errors and complexities are sought after. The main objective of this work is to study data reconstruction methods for processing sequence readouts at downstream stage of DNA data storage. We investigated applicability of three clustering tools -Starcode, Slidesort, MeShClust, and two algorithms - Majority Nucleotide Selection (MNS), Cooperative Sequence Clustering (CSC) by transforming them into suitable tools for storage application. We observed that for fixed redundancy of 6.3x to 8.6x based on the nature of the dataset, Starcode outperforms other tools with 1% to 40% higher recovery rate. However, it costs the highest decoding complexity whereas MNS and CSC provides the lowest decoding complexity. Moreover, the distribution of the cluster and clustering speed of each tool/method are compared. This is the first comparative analysis study of tools/methods for data reconstruction in DNA data storage.
Original language | English |
---|---|
Title of host publication | ICSEC 2022 - International Computer Science and Engineering Conference 2022 |
Publisher | Institute of Electrical and Electronics Engineers Inc. |
Pages | 269-274 |
Number of pages | 6 |
ISBN (Electronic) | 9781665491983 |
DOIs | |
Publication status | Published - 2022 |
Externally published | Yes |
Event | 26th International Computer Science and Engineering Conference, ICSEC 2022 - Sakon Nakhon, Thailand Duration: Dec 21 2022 → Dec 23 2022 |
Publication series
Name | ICSEC 2022 - International Computer Science and Engineering Conference 2022 |
---|
Conference
Conference | 26th International Computer Science and Engineering Conference, ICSEC 2022 |
---|---|
Country/Territory | Thailand |
City | Sakon Nakhon |
Period | 12/21/22 → 12/23/22 |
Bibliographical note
Publisher Copyright:© 2022 IEEE.
ASJC Scopus Subject Areas
- Artificial Intelligence
- Computer Networks and Communications
- Computer Science Applications
- Computer Vision and Pattern Recognition
- Information Systems and Management
- Control and Optimization
Keywords
- Clustering
- Data reconstruction
- DNA data storage
- Illumina sequencing