DeepsmirUD: Prediction of Regulatory Effects on microRNA Expression Mediated by Small Molecules Using Deep Learning
Student13:48CCAI
paperi.ai
0:00 / 0:00
Jianfeng Sun, Jinlong Ru, Lorenzo Ramos‐Mucci, Fei Qi, Zihao Chen, Suyuan Chen, Adam P. Cribbs, Li Deng, Xia Wang
A drug may bind a microRNA, but binding alone does not tell you whether that microRNA goes up or down. DeepsmirUD tackles that missing direction using an ensemble of twelve deep-learning frameworks.
Aberrant miRNA expression has been associated with a large number of human diseases. Therefore, targeting miRNAs to regulate their expression levels has become an important therapy against diseases that stem from the dysfunction of pathways regulated by miRNAs. In recent years, small molecules have demonstrated enormous potential as drugs to regulate miRNA expression (i.e., SM-miR). A clear understanding of the mechanism of action of small molecules on the upregulation and downregulation of miRNA expression allows precise diagnosis and treatment of oncogenic pathways. However, outside of a slow and costly process of experimental determination, computational strategies to assist this on an ad hoc basis have yet to be formulated. In this work, we developed, to the best of our knowledge, the first cross-platform prediction tool, DeepsmirUD, to infer small-molecule-mediated regulatory effects on miRNA expression (i.e., upregulation or downregulation). This method is powered by 12 cutting-edge deep-learning frameworks and achieved AUC values of 0.843/0.984 and AUCPR values of 0.866/0.992 on two independent test datasets. With a complementarily constructed network inference approach based on similarity, we report a significantly improved accuracy of 0.813 in determining the regulatory effects of nearly 650 associated SM-miR relations, each formed with either novel small molecule or novel miRNA. By further integrating miRNA–cancer relationships, we established a database of potential pharmaceutical drugs from 1343 small molecules for 107 cancer diseases to understand the drug mechanisms of action and offer novel insight into drug repositioning. Furthermore, we have employed DeepsmirUD to predict the regulatory effects of a large number of high-confidence associated SM-miR relations. Taken together, our method shows promise to accelerate the development of potential miRNA targets and small molecule drugs.
Transcript
A drug may bind a microRNA, but binding alone does not tell you whether that microRNA goes up or down. DeepsmirUD tackles that missing direction using an ensemble of twelve deep-learning frameworks. Aberrant miRNA expression has been associated with a large number of human diseases.
Targeting miRNAs to regulate their expression levels has therefore become an important therapy against diseases stemming from dysfunctional pathways regulated by miRNAs. Small molecules have demonstrated enormous potential as drugs to regulate miRNA expression, but experimentally determining these effects is slow and costly, and computational strategies for this task have yet to be formulated.
DeepsmirUD is a cross-platform prediction tool for inferring small-molecule-mediated regulatory effects on miRNA expression: upregulation or downregulation. Experimentally determining whether binding relationships exist between small molecules and miRNAs is normally time-consuming and costly because it is challenging to study all possible combinations.
In the SM2miR database, built using data from more than two thousand publications, only one point one four percent of all possible pairs were experimentally verified across one thousand four hundred ninety-two unique miRNAs and two hundred twelve unique small molecules.
Binding prediction has been studied with similarity and machine-learning methods, but predicting whether a binding small molecule causes downregulation or upregulation has remained computationally unexplored. Advances in deep learning have created opportunities for biological applications and discoveries, including protein structural and functional prediction.
To maximize performance, the study examined convolutional neural network and recurrent neural network models from computer vision and speech recognition. These models allow broad comparison of performance from shallow to ultradeep layers, visually and semantically.
The method comprises twelve deep-learning frameworks for predicting small-molecule-mediated regulatory effects on miRNA expression. The models were evaluated using curated SM-miR relations and their biophysical and biochemical features. LSTMCNN reportedly outperformed the other models on experimentally resolved relations, while ResNet18, ResNet50, and SCAResNet18 were preferable for stable prediction over long training epochs.
Individual models achieved AUC values from zero point eight zero to zero point nine two, and the final ensemble gained up to about two percent in AUC and one to two percent in AUCPR. DeepsmirUD differs from association prediction by using a sequentially ordered regulatory-effect determination process after association determination.
Association determination indicates whether and how likely an SM-miR association exists, but it does not identify the regulation type. An experimentally evidenced or predicted associated pair is then passed to DeepsmirUD to predict its regulation type, meaning its regulatory effects.
A one thousand three hundred ninety-six-size feature vector characterizes the upregulation or downregulation relation between a small molecule and a miRNA. The vector represents positional, compositional, physicochemical, or structural properties, with three hundred sixty features for the small molecule and one thousand thirty-six for the miRNA.
For computer-vision methods, each feature vector is converted to a thirty-seven by thirty-seven matrix; for speech-recognition methods, thirty-seven features are sequentially cropped at each time step. The study constructed multiple test datasets to investigate the performance of its deep learning tools and to assess whether they gained generalization ability in regulatory-effect inference.
TestUniqMIR and TestUniqSM contained relations involving unique miRNAs or unique small molecules, while TestRptMIR and TestRptSM contained miRNAs or small molecules that appeared in remaining relations. The remaining one thousand five hundred seventy-five upregulation relations and one thousand two hundred fourteen downregulation relations were randomly split for training and testing in a nine-to-one ratio.
Overfitting lowers the generalization abilities of intelligent models on unseen SM-miR relations, especially when their small molecules or miRNAs are distinguishable from every training example. To avoid overfitting, the study used early stopping and selected sufficiently trained models before prediction performance showed a falling tendency on test or validation data.
Figure one tracks AUC across training epochs for twelve deep-learning models on the independent Test and TestSim datasets, with boxplots summarizing their AUC distributions. Black circles mark the final models selected through early stopping on Test, while red dots indicate average predictions; the panels also report r-squared values and statistically significant t-test results.
This matters because it examines whether model performance remains consistent across training and across distinct test data, supporting the selection of models for DeepsmirUD. The individual models showed good performance on the Train and Test datasets, and LSTMCNN showed the best AUC and AUCPR performance.
Combining the top-ranked models or all twelve models produced DeepsmirUD-top and DeepsmirUD-all, which outperformed the individual models with AUC values of zero point eight four zero and zero point eight four three, and AUCPR values of zero point eight six six and zero point eight six six.
All models achieved AUC above zero point seven seven zero and AUCPR above zero point eight one zero. ResNet-based methods produced two even prediction peaks for downregulation and upregulation, while RNN-based methods produced prediction values around zero point five.
Figure two evaluates the deep-learning models across Train, Test, and TestSim data using ROC and precision-recall curves, while panel b summarizes accuracy, MCC, F one score, and precision on Test. Panels c and d examine how the two leading models, LSTMCNN and DeepsmirUD, compare in their Test predictions and how prediction values—interpreted as regulatory effects—are distributed across models.
Together, the figure shows both classification performance and the consistency and spread of predicted SM–miR regulatory relations. Pairwise combinations of unique miRNAs and small molecules can create a very large new SM-miR relation space containing potential upregulation and downregulation relations.
The models were trained on samples mixed with SM-miR relations formed using guilt by association and then tested on TestSim. On TestSim, all models, especially DeepsmirUD-top and ResNet18, achieved extremely high predictive performance, with AUC and AUCPR values of up to one.
On TestRptSM, where small molecules appeared at least once in training relations, DeepsmirUD-top achieved the best performance, with AUC zero point eight zero seven and AUCPR zero point eight one four. On TestRptMIR, DeepsmirUD-top remained best, with AUC zero point nine three zero and AUCPR zero point nine three two.
Most deep-learning algorithms performed well with recurrent miRNAs or small molecules, which may imply acceptable power for screening small-molecule drugs or miRNA targets in practice. For relations formed with either a novel miRNA or a novel small molecule, performance on TestUniqSM and TestUniqMIR was unsatisfactory and more susceptible to novel small molecules than novel miRNAs.
Changing random seeds two to three times produced results similar to the observation above, excluding sample bias as an explanation. Almost all deep-learning methods were incapable of accurately predicting regulatory effects when the small molecules were novel, while novel miRNAs had a better AUC value of around zero point five.
A similarity-based network inference approach was constructed to assist the optimization-based learning algorithms and offer better predictive ability. For relations formed with novel miRNAs, the approach achieved accuracy values of zero point nine two nine for upregulation and zero point seven three three for downregulation.
For relations formed with novel small molecules, accuracy values were zero point eight one zero and zero point seven seven eight, with an average of zero point eight one three. The network approach does not always return a prediction because inference can stop when no or not enough similar miRNAs or small molecules exist in the networks.
Figure three evaluates regulatory-effect prediction across four test datasets. Panel a shows ROC and precision–recall curves for deep-learning methods, while panel b illustrates a similarity-network approach that uses known training interactions to infer upregulation or downregulation for novel small molecules and miRNAs.
Panels c and d quantify the available novel–training combinations, binding counts, and accuracy, including reported accuracies of zero point nine two nine and zero point seven three three for novel-miRNA relations, and zero point eight one zero and zero point seven seven eight for novel-small-molecule relations.
Using deep-learning methods, the study predicted regulatory effects for two hundred twenty-four indirectly linked SM-miR relations screened with miRNA pharmacogenomic data. Around two-thirds of these SM-miR relations were predicted as upregulating.
Small molecules that rectify abnormalities in disease miRNA perturbation profiles can potentially be used for disease treatment and can provide insight into drug mechanisms of action. Using miRNA-cancer associations from miRCancer and SM-miR relations predicted by DeepsmirUD, the study generated SM-disease relationships for drug discovery and repositioning.
Connectivity scores were calculated from similarity between small-molecule-mediated miRNA perturbation profiles and cancer-associated miRNA perturbation profiles. The final predicted associations involve one hundred seven cancers and one thousand three hundred forty-three small molecules.
A negative score suggests pharmaceutical potential, while a positive score suggests a similar perturbation profile between a small molecule and a cancer disease. Panel a illustrates how disease-associated microRNA perturbation signatures are compared with small-molecule profiles using a weighted Kolmogorov–Smirnov connectivity score.
Panel b summarizes these scores across one thousand three hundred forty-three small molecules and one hundred seven cancer types: red indicates similar perturbations, while blue indicates reversal of the disease signature. This matters because negative connectivity suggests small molecules with potential therapeutic relevance for cancer and supports drug discovery or repositioning.
Two case studies involved drug-like small molecules with tangible anticancer effects validated by other studies. The FDA-approved antibiotic sulfafurazole has been found to inhibit tumor growth and metastasis in breast cancer by targeting endothelin receptor A, while Meticrane has been verified to suppress squamous cell carcinoma.
Cervical cancer, lung cancer, and colon cancer can be inhibited by naringin, which was validated by the calculated negative connectivity scores. DeepsmirUD is a deep-learning tool that quantifies small-molecule-mediated regulatory effects on miRNA expression as either upregulation or downregulation.
It does this by training twelve deep-learning architectures together as an ensemble. Accumulated pharmaceutical studies might be enriched for evidenced SM-miR relations, creating a need for a publication-based database as an extensive repertoire of these relations.
Such a database could improve machine-learning performance by getting rid of underfitting, making database establishment a promising direction for future work. Another future direction is predicting small-molecule-mediated regulatory effects on other noncoding RNAs, including siRNAs and lncRNAs.
DeepsmirUD shows that deep learning can predict small-molecule-mediated microRNA upregulation and downregulation, but novel small molecules remain difficult; a similarity-based network raises average accuracy to zero point eight one three.
A derivative work by Paperi · AI-generated script, voice and captions
· pages and figures unaltered
Made with Paperi.
Drop in a research PDF — get a narrated video walkthrough like this one,
with highlights that follow the narration. Free to start.
A traditional herbal root contains a valued protective compound, but only a small amount is normally available. This study found that growing a food fungus on the root could make much more of that compound accessible after digestion.A medicinal fungus turns a hard-to-access Astragalus compound into something far more abundant—and the fermented material still shows stronger antioxidant activity after simulated digestion.
A dry spell can damage ginseng twice: it weakens the plant and threatens the valuable compounds in its roots. This study found that one natural plant signal may help with both problems.What if the same hormone that helps ginseng conserve water during drought also boosts its valuable medicinal compounds? This study tests that connection—and finds ABA strengthens both drought resistance and ginsenoside biosynthesis.
Byungju Kim, Jincheol Seol, Yoon Ki Kim, Jong‐Bong Lee
Inside every living cell, tiny messages are read to build proteins. Scientists thought many of those messages formed helpful loops—but watching them one at a time suggests the loop may be an illusion of motion.For years, mRNA circularization has been treated as a functional closed loop that helps translation. But single-molecule imaging reveals a striking possibility: translating mRNA can look compact without being physically connected at its ends.