1)College of Information Engineering, Northwest A&F University, Shaanxi 712100, China;2)South Australian immunoGENomics Cancer Institute (SAiGENCI), Adelaide University, South Australia 5005, Australia
Q811.4;TP18
This work was supported by grants from the National Key Research and Development Program of China (2022YFF1000100), The National Natural Science Foundation of China (62202388), and the Qin Chuangyuan Innovation and Entrepreneurship Talent Project (QCYRCXM-2022-230).
Objective Multi-omics integration and analysis remain a major challenge in biomedical research. These tasks often require extensive coding skills and specialised bioinformatics expertise, which many biological and medical researchers do not have. Deep learning has emerged as a powerful approach for predictive modelling and data-driven discovery. However, most existing platforms focus only on traditional statistical analysis or basic machine learning. They do not combine deep learning support with standard statistical workflows. Moreover, none of them provide intelligent assistance to help with parameter tuning, error diagnosis, or result interpretation. To address this gap, we introduce AutoMATA, a fully code-free platform that streamlines multi-omics analysis from expression data to predictive modelling. AutoMATA is designed to lower the technical barrier for experimental biologists and clinical researchers who want to use advanced deep learning methods but lack programming experience.Methods AutoMATA's architecture is built around three core modules. The data processing and normalisation module automates essential steps such as gene and protein ID conversion and data normalisation. This module also allows users to integrate multi-omics data. The statistical analysis and visualisation module offers important and commonly used functions including differential expression analysis, principal component analysis, correlation analysis, and pathway enrichment for both GO and KEGG. Users can adjust thresholds and choose output formats for publication. The deep learning module provides twelve neural network architectures covering supervised, unsupervised, and semi-supervised learning. These include Multilayer Perception (MLP), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Transformer, Autoencoder, Variational Autoencoder (VAE), Radial Basis Function Neural Network (RBFNN), Self-Organising Map (SOM), DeepCluster, Pseudo-Labeling, and Ladder Network. Users can customise key settings including epoch, regularisation method, regularisation weight, dropout rate, and feature selection method. Two training strategies (train-validation-test split and stratified K-fold) are available. In addition, two Artificial Intelligence (AI) agents, DeepSeek and Qwen, are integrated into the platform. These agents answer user questions about parameter suggestion, task failure diagnosis, and result interpretations based on the actual task context and platform logic.Results We demonstrate AutoMATA in three case studies using public datasets. First, for colon adenocarcinoma data, AutoMATA reproduces gold-standard differential expression and clustering results. It correctly identifies known upregulated genes such as CXCL3 and CXCL8, and visualises enriched pathways. This confirms that AutoMATA's statistical module works as reliably as existing tools. Second, using multi-omics data for bladder, pancreatic, and stomach cancers, AutoMATA's deep learning module predicts cancer recurrence with good performance. Recurrent neural networks achieve the best results, with accuracy above 79% for stomach cancer. The platform also reproduces known biomarkers such as CDK6 in bladder cancer. Third, for breast cancer subtype classification, AutoMATA achieves 88.2% accuracy using a tuned RNN model. This performance is better than traditional machine learning methods like logistic regression and random forest. The AI agents provide on-demand assistance, helping users understand why certain models perform better, suggest parameter adjustments, and explain the biological meaning of the output. This makes the platform especially useful for non-experts.Conclusion By offering advanced multi-omics integration, comprehensive statistical analysis, flexible deep learning modelling options, and built-in AI agents, AutoMATA empowers researchers to extract meaningful biological insights and build robust predictive models without writing any code. The platform and source code are freely accessible at
WANG Cong, JIA Pan, HAO Yi, RAN Zi-Xu, GUO Xu-Dong, BI Yue, LIU Ning, LI Fu-Yi. AutoMATA: an AI-enhanced Bioinformatics Platform for Multi-omics Data Processing, Exploration and Modelling[J]. Progress in Biochemistry and Biophysics,,():
Copy

Scan code to follow ® 2026 Website Copyright ICP:京ICP备05023138号-1 京公网安备 11010502031771号
