DBATNDA DataBase
mirRNA-lncRNA-disease repository

Introduction

Non-coding RNAs (ncRNAs) are a diverse family of RNA molecules that do not encode proteins but are involved in various critical biological processes in humans. Among them, long non-coding RNAs (lncRNAs), which exceed 200 nucleotides in length, are the largest group and have significant roles in processes such as transcription, translation, splicing, epigenetic modulation, immune responses, and regulating the cell cycle. LncRNAs like HOTAIR, PCA3, and UCA1 have been identified as valuable biomarkers in the recurrence of hepatocellular carcinoma, prostate cancer severity, and bladder cancer diagnosis. MicroRNAs (miRNAs) are short, endogenous, non-coding RNA molecules that typically regulate gene expression post-transcriptionally by interacting with the 3’-untranslated regions (UTRs) of target mRNAs. Growing research indicates that miRNAs play a significant role in regulating cellular activities such as development, proliferation, and differentiation.

DBATNDA is generated based on the combined relationships among three distinct datasets. If a disease appears in either the lncRNA-disease or miRNA-disease datasets, it is incorporated into DBATNDA. This approach also processes additional miRNAs and lncRNAs, significantly increasing the overall data size. As a result, over 43,869 pairs of associations were included in the final dataset, comprising 23,581 miRNA-disease associations, 18,250 miRNA-lncRNA associations, and 2,036 lncRNA-disease associations. The dataset includes a total of 1,596 miRNAs, 2,188 lncRNAs, and 1,297 diseases.

The framework of our method

The overview of DBATNDA. Following the construction of original graph and preprocess, IDG and corresponding feature representations are constructed. GCN extended with balanced topological augmentation strategy is applied to perform the predicting task, achieving classification results. And other experiments are conducted to evaluate DBATNDA.