Bioinformatics & Systems Biology & Medical Informatics
I work as an Associate Professor of Computer Engineering at Istanbul Technical University. I build algorithms and databases that turn multi-omics data into clinical decision support at the meeting point of systems biology, machine learning, and medical informatics.
// a live map of the network below — click a node
Ali Çakmak, PhD
Associate Professor · Department of Computer Engineering, Istanbul Technical University
"My goal is to bridge the gap between computational discovery and patient care."
Over the past two decades, that's meant building metabolic network models that diagnose disease from blood chemistry, high-performance algorithms for integrating large-scale multi-omics datasets, and statistical models that predict which patients are at risk of various clinical conditions.
Before returning to academia, I spent four years on Oracle's Query Optimizer team in Redwood Shores, where two of my histogram designs went into Oracle Database 12c.
I currently coordinate the dual-degree BSc program in Information Systems Engineering between ITU and SUNY Binghamton, and consult for the Turkish Genome Project and National Health Data Sharing Portal at the Health Institutes of Türkiye (TÜSEB). I serve as an Associate Editor for IEEE Transactions on Computational Biology and Bioinformatics. Moreover, I actively lead the computational work packages in two different EU consortia on neurodegenerative disease biomarker discovery and cardiovascular disease risk.
My lab's near-term goal is an individualized clinical decision system: one that simulates how blocking a specific drug target reshapes a patient's metabolic network — and recommends treatment accordingly.
A comprehensive metabolism-oriented integrated multi-omics analysis method structurally designed to accommodate genomics, transcriptomics, proteomics, and metabolomics datasets. It centers on the construction of an integrated multi-omic interaction network that incorporates a wide range of biological interactions, including gene expression, translation, transcription factor activity, and post-transcriptional regulation via microRNAs. To capture the cascading effects of molecular changes, it utilizes information diffusion models, such as Linear Threshold Diffusion, to propagate fold-changes throughout the system.
Managing computational analysis for a transnational EU consortium on neurodegenerative disease, concluding April 2027, supported by Biomark-X ↗ and Biomark ↗ for biomarker analysis — part of a broader push to move computational insight into clinical practice. Our goal is to discover novel non-invasive biomarkers and therapeutic targets for Alzheimer’s disease using omics and data science approaches.
An algorithm for systems-level analysis of biochemical network changes, applied to breast cancer, Crohn's disease, and colorectal cancer with over 90% average diagnostic accuracy. Filed as an international patent for low-cost, non-invasive clinical decision support; made accessible through the web-based MetaboliticsDB ↗.
Database-enabled tools that integrate metabolic pathway data with systems biology models for simulation, visualization, and querying — most recently extended to model genome-wide viral mutation probabilities with CovMutEx ↗.
Public metabolomics databases offer a large number of datasets. However, most datasets include measurements for only a very small fraction of the known metabolites. Hence, simply putting together these studies leads to very sparse datasets, which do not lend themselves well to training machine learning models. We developed novel approaches for vertical dataset merging across datasets. Variational autoencoders (VAE) are heavily employed for imputation model training. The developed approaches can readily be applied to other multi-omics integration challenges.
Flux Variability Analysis (FVA) is the gold standard for computing reaction flux intervals. However, its reliance on linear programming makes it computationally intensive, often requiring hours or days for large cohorts on complex genome-scale metabolic network models. mFLIP is a machine learning-based framework for predicting flux intervals across metabolic pathways. mFLIP provides a fast and accurate alternative to traditional FVA-based approaches for metabolic flux interval estimation. By leveraging supervised learning on FVA-derived data, it enables scalable analysis of large cohorts while maintaining high predictive performance, making it a practical tool for large-scale metabolic studies.
The main aim of this project is to develop a metabolism-based pipeline that will allow a computational investigation of whether the currently approved drugs for any disease are suitable for use in treating diseases other than their intended target. In particular, we develop tools to simulate the effect of each drug on patient metabolism, and then assess whether the drug can push a diseased metabolic profile back toward a healthy one. This approach can also surface entirely new drug targets that are not on the market.
Readmission of patients after a short period of time from their discharge costs billions of dollars to governments every year. This research aims to develop an integrated readmission risk monitoring system that will continuously monitor patients, and act as an auxiliary decision support system to provide clinicians with the risk of readmission during the entire period of a patient's stay, as well as after the discharge through the periodic remote measurements taken by the patient at home.
Chronic diseases such as diabetes and heart failure often suffer from late diagnoses, which can lead to preventable complications, inefficient care management, and increased healthcare costs. This project presents a timeline-aware machine learning framework that estimates disease probability at each hospital visit per-patient and compares these predictions with actual diagnosis date to identify missed opportunities.
Static record-keeping in the form of electronic health records (EHRs) is increasingly inadequate in an era defined by multimodal data and foundation models. Computational Patient is a continuously evolving, general-purpose digital representation that integrates heterogeneous inputs, including EHRs, wearable device data, and multi-omics profiles, into a unified latent state-space. Unlike task-specific digital twins, this framework supports cyclic feedback loops, uncertainty quantification, and simulation-based reasoning.
Cancer susceptibility is driven by a complex interplay of genetic predispositions and environmental exposures, yet traditional analyses frequently evaluate these factors in isolation. We investigate the combinatorial impact of targeted Single Nucleotide Polymorphisms (SNPs), clinical demographics, and behavioral risks on lung and colorectal cancer etiology.
Understanding how genomic variability is translated into stable or dysfunctional metabolic phenotypes remains a central challenge in systems biology. We develop a multi-omic, information-theoretic framework that maps transcriptomic variation onto metabolic network activity using deterministic Gene-Protein-Reaction (GPR) rules applied to the Recon3D human metabolic network. By integrating paired transcriptomic and metabolomic data across cancer and neurodegenerative cohorts, we quantify pathway-level entropy, variance buffering (canalization), and topological information flow.
Breast cancer–related lymphedema (BCRL) is a frequent and clinically significant complication of breast cancer treatment. Early identification of high-risk patients may facilitate personalized surveillance and preventive interventions. We develop an interpretable machine learning framework to enable accurate BCRL risk stratification and support personalized surveillance and survivorship care in breast cancer patients. We also designed a web-based clinical decision support system.
Clinical artificial intelligence (AI) systems are predominantly trained under the assumption that each patient corresponds to a single, well-defined ground truth label. However, increasing evidence from diagnostic variability, noisy label learning, and uncertainty quantification suggests that medical labels are inherently probabilistic, fluid, and context-dependent. We develop a Causal Probabilistic Clinical State Framework (C-PCSF). In this paradigm, disease is modeled not as a discrete label, but as a posterior distribution over a latent physiological manifold, coupled with intervention-aware causal dynamics and utility-driven decision policies.
Single-cell RNA sequencing provides high-resolution characterization of the tumor microenvironment, yet its limited scalability, high cost, and technical variability restrict widespread clinical adoption. Consequently, bulk tissue multi-omics remains the dominant profiling strategy in clinical oncology, despite its inherent limitation in averaging heterogeneous cellular signals and obscuring cell-type–specific states. We develop a topology-guided representation learning framework that integrates paired bulk transcriptomic and metabolomic data to construct high-resolution pseudo-single-cell–like representations.
Complex genetic diseases result from intricate interactions among genetic variants and environmental factors. While genome-wide association studies (GWAS) have identified numerous risk loci, their limited mechanistic insight underscores the need for integrative, systems-level approaches. We develop a novel computational framework that integrates polygenic risk scores with genome-scale metabolic models to contextualize genetic susceptibility within a condition-specific metabolic landscape.
Where this is heading
Expanding Metabolitics to integrate transcriptome, proteome, and metabolome data into one individualized clinical decision system — computing optimal treatments by simulating drug-target blocks on a patient's own metabolic network.
A metabolism-based pipeline to test whether approved drugs — alone or in combination — can push a diseased metabolic profile back toward a healthy one, and to surface entirely new drug targets no therapy currently touches.
A discharge decision-support system that tracks patients during their stay and after, predicting readmission risk for costly chronic conditions like myocardial infarction, diabetes, and Alzheimer's disease.
A web-based database of metabolomics analyses built on Metabolitics — lets researchers store, compare, and query network-based flux analysis results across studies to track how a disease progresses.
Visit MetaboliticsDB ↗The original single-omics biomarker discovery tool — combines robust feature selection, classification, and SHAP/LIME model explainability to help researchers uncover patient subgroups from a single dataset. Since extended into Biomark-X.
Visit Biomark ↗A comprehensively expanded successor to Biomark: native multi-omics integration (with an optional MOGONET graph pipeline), SMOTE/ADASYN handling of imbalanced clinical cohorts, survival analysis (Kaplan-Meier, Cox PH), automated KEGG/GO pathway enrichment, and external biomarker validation against Open Targets, EWAS Atlas, and JensenLab DISEASES — all backed by a stateful platform that lets researchers pause and resume long analyses.
Visit Biomark-X ↗An extensible software framework for exploring SARS-CoV-2 genome-wide mutation probabilities — built on the same systems biology platform used for metabolic pathway modeling.
Visit CovMutEx ↗A generalized mutation-exploration platform extending the CovMutEx framework beyond SARS-CoV-2 to viral genomes more broadly.
Visit VirMutEx ↗LoopLegends is a dedicated community platform where die-cast model car collectors can create personal digital garages to share photos, videos, and stories of their collections with fellow enthusiasts.
Visit LoopLegends ↗"Students are the ultimate decision-makers who determine whether an educational effort successfully translates into tangible learning outcomes."
I treat students as the primary actors in the classroom rather than a passive audience, and design courses to flex around the varied expectations students arrive with. That same philosophy carries into mentorship when moving doctoral and graduate researchers from structured coursework into independent, hypothesis-driven work on real omics data, thesis planning, and manuscript revision, and into curriculum leadership, where I coordinate the ITU–SUNY Binghamton dual-degree program in Information Systems Engineering.
Istanbul Technical University
IEEE Transactions on Computational Biology and Bioinformatics
Health Institutes of Türkiye
Istanbul Technical University
Istanbul Technical University
Istanbul Şehir University
Oracle, Inc., Redwood Shores, CA
Case Western Reserve University
Case Western Reserve University — PhD, Computer Science
Case Western Reserve University, Cleveland, OH
Bilkent University, Ankara, Türkiye
Horizon Europe (International collaboration with partners from Spain, Netherlands, Denmark, Belgium, Lithuania, France, England, and Sweden) · Co-PI · 2026–2031
JPND · EU Joint Funding Program (International collaboration with partners from Ireland, Italy, and Poland) · Co-PI · 2024–2027
TÜBİTAK ARDEB 1001 · Co-PI · 2026–2029
TÜBİTAK ARDEB 1001 · Co-PI · 2025–2028
ITU BAP GAP · PI · 2025–2027
TUSEB ARGE-02 · Co-PI · 2023–2026
ITU BAP HIZDEP · PI · 2025
ITU BAP HIZDEP · PI · 2024
TUSEB ARGE-02 · Co-PI · 2023–2024
TUSEB TA-01 · PI · 2020–2023 (awarded, but not started due to institutional move)
TUSEB TA-01 · Co-PI · 2020–2023
TÜBİTAK ARDEB 1001 · PI · 2017–2019
TÜBİTAK ARDEB 3501 (CAREER Grant) · PI · 2015–2017
TÜBİTAK BIDEB 2232 (Returning Researcher Grant) · PI · 2014–2015
I'm interested in photography, though I wouldn't say it's gone much beyond that. Below are a few photographs taken in Cleveland, Antalya, and Istanbul.









