Comparative Analysis of Clustering Methodologies in DNA Storage

Subhasiny Sankar*, Yixin Wang, Zhang Jiayu, Nur Sabrina, Erry Gunawan, Yong Liang Guan, Noor A.Rahim Md, Chueh Loo Poh

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference contribution

2 Citations (Scopus)

Abstract

Owing to the significance of DNA storage technology in meeting exponential storage demands and longevity, the challenges caused by bio-molecular errors while reading/sequencing data from DNA molecules must be addressed. By reading redundant copies, data can be reconstructed but with associated cost of sequencing and decoding complexities. Hence, solutions for dealing with both errors and complexities are sought after. The main objective of this work is to study data reconstruction methods for processing sequence readouts at downstream stage of DNA data storage. We investigated applicability of three clustering tools -Starcode, Slidesort, MeShClust, and two algorithms - Majority Nucleotide Selection (MNS), Cooperative Sequence Clustering (CSC) by transforming them into suitable tools for storage application. We observed that for fixed redundancy of 6.3x to 8.6x based on the nature of the dataset, Starcode outperforms other tools with 1% to 40% higher recovery rate. However, it costs the highest decoding complexity whereas MNS and CSC provides the lowest decoding complexity. Moreover, the distribution of the cluster and clustering speed of each tool/method are compared. This is the first comparative analysis study of tools/methods for data reconstruction in DNA data storage.

Original languageEnglish
Title of host publicationICSEC 2022 - International Computer Science and Engineering Conference 2022
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages269-274
Number of pages6
ISBN (Electronic)9781665491983
DOIs
Publication statusPublished - 2022
Externally publishedYes
Event26th International Computer Science and Engineering Conference, ICSEC 2022 - Sakon Nakhon, Thailand
Duration: Dec 21 2022Dec 23 2022

Publication series

NameICSEC 2022 - International Computer Science and Engineering Conference 2022

Conference

Conference26th International Computer Science and Engineering Conference, ICSEC 2022
Country/TerritoryThailand
CitySakon Nakhon
Period12/21/2212/23/22

Bibliographical note

Publisher Copyright:
© 2022 IEEE.

ASJC Scopus Subject Areas

  • Artificial Intelligence
  • Computer Networks and Communications
  • Computer Science Applications
  • Computer Vision and Pattern Recognition
  • Information Systems and Management
  • Control and Optimization

Keywords

  • Clustering
  • Data reconstruction
  • DNA data storage
  • Illumina sequencing

Fingerprint

Dive into the research topics of 'Comparative Analysis of Clustering Methodologies in DNA Storage'. Together they form a unique fingerprint.

Cite this