Beyond the Alphabet: Deep Signal Embedding for Enhanced DNA Clustering

要約

DNA貯蔵の新たなフィールドは、DNA塩基（A/T/C/G）の鎖を、デジタル情報のストレージ媒体として使用して、大規模な密度と耐久性を可能にします。
DNAストレージパイプラインには、次のものが含まれます。（1）生データをDNA塩基のシーケンスにエンコードする。
（2）順序付けられていないセットとして時間の経過とともに保存されるDNA \ textIT {strands}としてシーケンスを合成する。
（3）DNA鎖をシーケンスしてDNA \ textIT {reads}を生成します。
（4）元のデータを推測します。
DNA合成とシーケンス段階はそれぞれ、各鎖のいくつかの独立したエラーが発生しやすい複製を生成し、最終段階で利用して元の鎖の最良の推定値を再構築します。
具体的には、読み取りは最初の\ textit {クラスター化}は、同じ鎖から発生する可能性のあるグループ（互いに類似性に基づいて）になり、各グループはそのグループの読み取りにつながった鎖に近似します。
この作業は、DNAシーケンスの一部として埋め込むことにより、DNAクラスタリング段階を改善します。
従来のDNA貯蔵ソリューションは、DNAシーケンスプロセスが離散DNA読み取り（A/T/C/G）を生成した後に始まりますが、ナノポアDNAシーケンスマシンによって生成された生の信号を使用する前に、塩基に離散化する前に未開発の可能性があることを特定します。
、\ textit {Basecalling}として知られるプロセス。これは、深いニューラルネットワークを使用して行われます。
これらの信号を直接閉じ込める深いニューラルネットワークを提案し、優れた精度を実証し、ベースコール後にクラスターする現在のアプローチと比較して計算時間を短縮します。

要約(オリジナル)

The emerging field of DNA storage employs strands of DNA bases (A/T/C/G) as a storage medium for digital information to enable massive density and durability. The DNA storage pipeline includes: (1) encoding the raw data into sequences of DNA bases; (2) synthesizing the sequences as DNA \textit{strands} that are stored over time as an unordered set; (3) sequencing the DNA strands to generate DNA \textit{reads}; and (4) deducing the original data. The DNA synthesis and sequencing stages each generate several independent error-prone duplicates of each strand which are then utilized in the final stage to reconstruct the best estimate for the original strand. Specifically, the reads are first \textit{clustered} into groups likely originating from the same strand (based on their similarity to each other), and then each group approximates the strand that led to the reads of that group. This work improves the DNA clustering stage by embedding it as part of the DNA sequencing. Traditional DNA storage solutions begin after the DNA sequencing process generates discrete DNA reads (A/T/C/G), yet we identify that there is untapped potential in using the raw signals generated by the Nanopore DNA sequencing machine before they are discretized into bases, a process known as \textit{basecalling}, which is done using a deep neural network. We propose a deep neural network that clusters these signals directly, demonstrating superior accuracy, and reduced computation times compared to current approaches that cluster after basecalling.

arxiv情報

著者	Hadas Abraham,Barak Gahtan,Adir Kobovich,Orian Leitersdorf,Alex M. Bronstein,Eitan Yaakobi
発行日	2025-01-27 16:32:53+00:00
arxivサイト	arxiv_id(pdf)

提供元, 利用サービス

arxiv.jp, Google

Beyond the Alphabet: Deep Signal Embedding for Enhanced DNA Clustering

要約

要約(オリジナル)

arxiv情報

提供元, 利用サービス

最近の投稿

最近のコメント

アーカイブ

カテゴリー