Bioinformatics & Systems Biology & Medical Informatics

Tracing the biological wiring of diseases, in code.

I work as an Associate Professor of Computer Engineering at Istanbul Technical University. I build algorithms and databases that turn multi-omics data into clinical decision support at the meeting point of systems biology, machine learning, and medical informatics.

Istanbul Technical University · Dept. of Computer Engineering · PhD, Case Western Reserve '08

// a live map of the network below — click a node

About

01 / Profile
Portrait of Ali Çakmak

Ali Çakmak, PhD

Associate Professor · Department of Computer Engineering, Istanbul Technical University

"My goal is to bridge the gap between computational discovery and patient care."

Over the past two decades, that's meant building metabolic network models that diagnose disease from blood chemistry, high-performance algorithms for integrating large-scale multi-omics datasets, and statistical models that predict which patients are at risk of various clinical conditions.

Before returning to academia, I spent four years on Oracle's Query Optimizer team in Redwood Shores, where two of my histogram designs went into Oracle Database 12c.

I currently coordinate the dual-degree BSc program in Information Systems Engineering between ITU and SUNY Binghamton, and consult for the Turkish Genome Project and National Health Data Sharing Portal at the Health Institutes of Türkiye (TÜSEB). I serve as an Associate Editor for IEEE Transactions on Computational Biology and Bioinformatics. Moreover, I actively lead the computational work packages in two different EU consortia on neurodegenerative disease biomarker discovery and cardiovascular disease risk.

My lab's near-term goal is an individualized clinical decision system: one that simulates how blocking a specific drug target reshapes a patient's metabolic network — and recommends treatment accordingly.

Research

02 / Current work

MetabOmicsmulti-omics

A comprehensive metabolism-oriented integrated multi-omics analysis method structurally designed to accommodate genomics, transcriptomics, proteomics, and metabolomics datasets. It centers on the construction of an integrated multi-omic interaction network that incorporates a wide range of biological interactions, including gene expression, translation, transcription factor activity, and post-transcriptional regulation via microRNAs. To capture the cascading effects of molecular changes, it utilizes information diffusion models, such as Linear Threshold Diffusion, to propagate fold-changes throughout the system.

Biomarker and Therapeutic Target Discovery for Alzheimer’s Disease: neurodegeneration

Managing computational analysis for a transnational EU consortium on neurodegenerative disease, concluding April 2027, supported by Biomark-X ↗ and Biomark ↗ for biomarker analysis — part of a broader push to move computational insight into clinical practice. Our goal is to discover novel non-invasive biomarkers and therapeutic targets for Alzheimer’s disease using omics and data science approaches.

Metabolitics multi-omics

An algorithm for systems-level analysis of biochemical network changes, applied to breast cancer, Crohn's disease, and colorectal cancer with over 90% average diagnostic accuracy. Filed as an international patent for low-cost, non-invasive clinical decision support; made accessible through the web-based MetaboliticsDB ↗.

Systems biology platforms infrastructure

Database-enabled tools that integrate metabolic pathway data with systems biology models for simulation, visualization, and querying — most recently extended to model genome-wide viral mutation probabilities with CovMutEx ↗.

Vertical Omics Data Integration data integration

Public metabolomics databases offer a large number of datasets. However, most datasets include measurements for only a very small fraction of the known metabolites. Hence, simply putting together these studies leads to very sparse datasets, which do not lend themselves well to training machine learning models. We developed novel approaches for vertical dataset merging across datasets. Variational autoencoders (VAE) are heavily employed for imputation model training. The developed approaches can readily be applied to other multi-omics integration challenges.

Metabolic Flux Interval Prediction ML

Flux Variability Analysis (FVA) is the gold standard for computing reaction flux intervals. However, its reliance on linear programming makes it computationally intensive, often requiring hours or days for large cohorts on complex genome-scale metabolic network models. mFLIP is a machine learning-based framework for predicting flux intervals across metabolic pathways. mFLIP provides a fast and accurate alternative to traditional FVA-based approaches for metabolic flux interval estimation. By leveraging supervised learning on FVA-derived data, it enables scalable analysis of large cohorts while maintaining high predictive performance, making it a practical tool for large-scale metabolic studies.

Drug Repositioning ML

The main aim of this project is to develop a metabolism-based pipeline that will allow a computational investigation of whether the currently approved drugs for any disease are suitable for use in treating diseases other than their intended target. In particular, we develop tools to simulate the effect of each drug on patient metabolism, and then assess whether the drug can push a diseased metabolic profile back toward a healthy one. This approach can also surface entirely new drug targets that are not on the market.

Readmission Risk Prediction medical informatics

Readmission of patients after a short period of time from their discharge costs billions of dollars to governments every year. This research aims to develop an integrated readmission risk monitoring system that will continuously monitor patients, and act as an auxiliary decision support system to provide clinicians with the risk of readmission during the entire period of a patient's stay, as well as after the discharge through the periodic remote measurements taken by the patient at home.

Identifying Missed Opportunities in Chronic Diseasesmedical informatics

Chronic diseases such as diabetes and heart failure often suffer from late diagnoses, which can lead to preventable complications, inefficient care management, and increased healthcare costs. This project presents a timeline-aware machine learning framework that estimates disease probability at each hospital visit per-patient and compares these predictions with actual diagnosis date to identify missed opportunities.

Computational Patient medical informatics

Static record-keeping in the form of electronic health records (EHRs) is increasingly inadequate in an era defined by multimodal data and foundation models. Computational Patient is a continuously evolving, general-purpose digital representation that integrates heterogeneous inputs, including EHRs, wearable device data, and multi-omics profiles, into a unified latent state-space. Unlike task-specific digital twins, this framework supports cyclic feedback loops, uncertainty quantification, and simulation-based reasoning.

Cancer Risk Prediction genomics, clinical informatics

Cancer susceptibility is driven by a complex interplay of genetic predispositions and environmental exposures, yet traditional analyses frequently evaluate these factors in isolation. We investigate the combinatorial impact of targeted Single Nucleotide Polymorphisms (SNPs), clinical demographics, and behavioral risks on lung and colorectal cancer etiology.

Signatures of Biological Information Flow multi-omics, network analysis, systems biology

Understanding how genomic variability is translated into stable or dysfunctional metabolic phenotypes remains a central challenge in systems biology. We develop a multi-omic, information-theoretic framework that maps transcriptomic variation onto metabolic network activity using deterministic Gene-Protein-Reaction (GPR) rules applied to the Recon3D human metabolic network. By integrating paired transcriptomic and metabolomic data across cancer and neurodegenerative cohorts, we quantify pathway-level entropy, variance buffering (canalization), and topological information flow.

Risk Stratification of Lymphedema medical informatics

Breast cancer–related lymphedema (BCRL) is a frequent and clinically significant complication of breast cancer treatment. Early identification of high-risk patients may facilitate personalized surveillance and preventive interventions. We develop an interpretable machine learning framework to enable accurate BCRL risk stratification and support personalized surveillance and survivorship care in breast cancer patients. We also designed a web-based clinical decision support system.

Causal Probabilistic Clinical State Framework medical informatics, AI

Clinical artificial intelligence (AI) systems are predominantly trained under the assumption that each patient corresponds to a single, well-defined ground truth label. However, increasing evidence from diagnostic variability, noisy label learning, and uncertainty quantification suggests that medical labels are inherently probabilistic, fluid, and context-dependent. We develop a Causal Probabilistic Clinical State Framework (C-PCSF). In this paradigm, disease is modeled not as a discrete label, but as a posterior distribution over a latent physiological manifold, coupled with intervention-aware causal dynamics and utility-driven decision policies.

Metabolic Deconvolution of Bulk Omics systems biology, multi-omics

Single-cell RNA sequencing provides high-resolution characterization of the tumor microenvironment, yet its limited scalability, high cost, and technical variability restrict widespread clinical adoption. Consequently, bulk tissue multi-omics remains the dominant profiling strategy in clinical oncology, despite its inherent limitation in averaging heterogeneous cellular signals and obscuring cell-type–specific states. We develop a topology-guided representation learning framework that integrates paired bulk transcriptomic and metabolomic data to construct high-resolution pseudo-single-cell–like representations.

Fluxomics Meets Polygenic Risk Scores systems biology, multi-omics

Complex genetic diseases result from intricate interactions among genetic variants and environmental factors. While genome-wide association studies (GWAS) have identified numerous risk loci, their limited mechanistic insight underscores the need for integrative, systems-level approaches. We develop a novel computational framework that integrates polygenic risk scores with genome-scale metabolic models to contextualize genetic susceptibility within a condition-specific metabolic landscape.

Where this is heading

TREATMENT

Expanding Metabolitics to integrate transcriptome, proteome, and metabolome data into one individualized clinical decision system — computing optimal treatments by simulating drug-target blocks on a patient's own metabolic network.

REPOSITIONING

A metabolism-based pipeline to test whether approved drugs — alone or in combination — can push a diseased metabolic profile back toward a healthy one, and to surface entirely new drug targets no therapy currently touches.

READMISSION

A discharge decision-support system that tracks patients during their stay and after, predicting readmission risk for costly chronic conditions like myocardial infarction, diabetes, and Alzheimer's disease.

Tools

03 / Software built in the lab

MetaboliticsDB metabolomics

A web-based database of metabolomics analyses built on Metabolitics — lets researchers store, compare, and query network-based flux analysis results across studies to track how a disease progresses.

Visit MetaboliticsDB ↗

Biomark biomarker analysis

The original single-omics biomarker discovery tool — combines robust feature selection, classification, and SHAP/LIME model explainability to help researchers uncover patient subgroups from a single dataset. Since extended into Biomark-X.

Visit Biomark ↗

Biomark-X multi-omics · biomarker validation

A comprehensively expanded successor to Biomark: native multi-omics integration (with an optional MOGONET graph pipeline), SMOTE/ADASYN handling of imbalanced clinical cohorts, survival analysis (Kaplan-Meier, Cox PH), automated KEGG/GO pathway enrichment, and external biomarker validation against Open Targets, EWAS Atlas, and JensenLab DISEASES — all backed by a stateful platform that lets researchers pause and resume long analyses.

Visit Biomark-X ↗

CovMutEx viral genomics

An extensible software framework for exploring SARS-CoV-2 genome-wide mutation probabilities — built on the same systems biology platform used for metabolic pathway modeling.

Visit CovMutEx ↗

VirMutEx (coming soon) viral genomics

A generalized mutation-exploration platform extending the CovMutEx framework beyond SARS-CoV-2 to viral genomes more broadly.

Visit VirMutEx ↗

LoopLegends courtesy of our youngest trainee in residence (my son 😀) - built completely with vibecoding

LoopLegends is a dedicated community platform where die-cast model car collectors can create personal digital garages to share photos, videos, and stories of their collections with fellow enthusiasts.

Visit LoopLegends ↗

Publications

04 / Publications

Teaching

05 / Courses

"Students are the ultimate decision-makers who determine whether an educational effort successfully translates into tangible learning outcomes."

I treat students as the primary actors in the classroom rather than a passive audience, and design courses to flex around the varied expectations students arrive with. That same philosophy carries into mentorship when moving doctoral and graduate researchers from structured coursework into independent, hypothesis-driven work on real omics data, thesis planning, and manuscript revision, and into curriculum leadership, where I coordinate the ITU–SUNY Binghamton dual-degree program in Information Systems Engineering.

Istanbul Technical University

  • Bioinformatics Algorithms (Graduate)
  • Advanced Database Systems (Graduate)
  • Introduction to Bioinformatics (Junior)
  • Database Systems (Junior)
  • Analysis of Algorithms (Junior)
  • IT Systems Analysis & Design (Junior)
  • Introduction to Information Systems & Computer Engineering (Freshman)
  • Introduction to Programming (Freshman)

Istanbul Şehir University

  • Bioinformatics (Graduate)
  • Networks Modeling (Graduate)
  • Database Systems (Junior)
  • Introduction to Programming (Freshman)
  • Programming Practice (Freshman)

Case Western Reserve University

  • Data Structures (Sophomore)

Experience

06 / Positions
2025—

Program Coordinator, SUNY Binghamton–ITU Dual Degree

Istanbul Technical University

2025—

Associate Editor

IEEE Transactions on Computational Biology and Bioinformatics

2022—

Consulting Researcher, Turkish Genome Project

Health Institutes of Türkiye

2021, 2024

Vice Chair, Department of Computer Engineering

Istanbul Technical University

2020—

Associate Professor

Istanbul Technical University

2013–2020

Assistant → Associate Professor

Istanbul Şehir University

2009–2013

Senior Member of Technical Staff, Query Optimizer Team

Oracle, Inc., Redwood Shores, CA

2008–2009

Instructor & Postdoctoral Research Associate

Case Western Reserve University

2003–2008

Graduate Research Assistant, Bioinformatics & Database Group

Case Western Reserve University — PhD, Computer Science

Education

07 / Degrees
2003–2008

Ph.D., Computer Science

Case Western Reserve University, Cleveland, OH

1999–2003

B.Sc., Computer Engineering

Bilkent University, Ankara, Türkiye

Grants

08 / Research Funding

Understanding Sex-specific Omic Patterns Leading to Cardiovascular Disease Risk - UNSOLVED-CARDS

Horizon Europe (International collaboration with partners from Spain, Netherlands, Denmark, Belgium, Lithuania, France, England, and Sweden) · Co-PI · 2026–2031

€6,999,542

(Micro)RNA & informatics approaches for diagnosis, prognosis and treatment of Alzheimer's disease and Dementia

JPND · EU Joint Funding Program (International collaboration with partners from Ireland, Italy, and Poland) · Co-PI · 2024–2027

€1,104,433

Investigation of the Genetic and Metabolic Mechanisms of Increased Proliferation in Breast Cancer Cells Under Aspartate Deprivation

TÜBİTAK ARDEB 1001 · Co-PI · 2026–2029

~2,700,000 TL

Investigation of the Interaction Between Arginine and Cholesterol Metabolism in Triple Negative Breast Cancer Cells

TÜBİTAK ARDEB 1001 · Co-PI · 2025–2028

~2,400,000 TL

New microRNA Biomarkers for Alzheimer's Disease

ITU BAP GAP · PI · 2025–2027

300,000 TL

Türkiye Genome Project — Population Genomics Analysis

TUSEB ARGE-02 · Co-PI · 2023–2026

~5,000,000 TL

Predicting Student Course Grades — A Clustering-Based Hierarchical Approach

ITU BAP HIZDEP · PI · 2025

90,000 TL

Predicting Future Covid-19 Mutations using Deep Learning Methods

ITU BAP HIZDEP · PI · 2024

60,000 TL

Personal Medicine Approach in Bladder Cancer Treatment: Genomic/Transcriptomic Analysis and Tumor Modeling

TUSEB ARGE-02 · Co-PI · 2023–2024

~2,200,000 TL

Tools and Algorithms for Multi-omics Supported Personalized Treatment Recommendation, Drug Repositioning, and Drug Target Discovery

TUSEB TA-01 · PI · 2020–2023 (awarded, but not started due to institutional move)

~349,000 TL

A Flexible and Easy-to-use Pipeline for Next Generation Sequencing Analysis of Cancer Samples

TUSEB TA-01 · Co-PI · 2020–2023

~348,000 TL

New Methods and Algorithms to Estimate the Selectivity of SQL LIKE Queries

TÜBİTAK ARDEB 1001 · PI · 2017–2019

~121,000 TL

Algorithms and Tools for Computational Modeling of Metabolomics Data

TÜBİTAK ARDEB 3501 (CAREER Grant) · PI · 2015–2017

~223,000 TL

Data Mining Techniques for Fast Query Optimization in Relational Database Systems

TÜBİTAK BIDEB 2232 (Returning Researcher Grant) · PI · 2014–2015

~95,000 TL

Photography

09 / An interest, not a second career

I'm interested in photography, though I wouldn't say it's gone much beyond that. Below are a few photographs taken in Cleveland, Antalya, and Istanbul.

Aspendos ancient theatre, Antalya
Aspendos, Antalya
Mosque, Afyonkarahisar
Paşa Camii, Afyonkarahisar
Istanbul
Çengelköy, Istanbul
Sunset
Sunset on Lake Erie
Moon
Moon, Cleveland
Fisherman, Lake Erie
Fisherman, Lake Erie
Cleveland
Cleveland
Natural History Museum, Cleveland
Natural History Museum, Cleveland
Natural History Museum, Cleveland
Natural History Museum, Cleveland
Natural History Museum, Cleveland
Natural History Museum, Cleveland