Before any new drug reaches a large population of patients, the central question is not only whether it works, but whether it will cause harm. For inflammatory bowel disease, a group of chronic conditions that includes Crohn’s disease and ulcerative colitis, the therapeutic landscape has expanded dramatically in recent years, with biologics and small molecules targeting interleukins, integrins, Janus kinases and tumour necrosis factor pathways. Yet each of these targets carries a potential shadow: the possibility that suppressing or modulating a molecule central to gut inflammation may also disturb immune surveillance, cardiovascular balance or thrombotic risk elsewhere in the body. A new study published in the Journal of Translational Medicine offers a systematic, genetics-first way to anticipate those shadows before broad clinical exposure, using naturally occurring human genetic variation as a large-scale, decades-long natural experiment.
The research, led by Xin Gao and Xiaofeng Guo of Shanxi Provincial People’s Hospital together with Shaopeng Yang of Zhejiang Cancer Hospital and corresponding author Qi Zhang of Shandong Cancer Hospital and Institute, describes a protein-prioritised, colocalisation-informed framework for generating target-outcome safety hypotheses. Rather than waiting for post-marketing surveillance to reveal unexpected adverse events, the team asked whether human genetics could flag plausible safety concerns for eighteen genes that represent established or emerging inflammatory bowel disease therapeutic targets. The logic is elegant: if genetically lowered levels or expression of a drug’s target protein are associated with increased risk of a particular disease outcome in population data, then pharmacologically inhibiting that same target deserves closer scrutiny for a similar effect.
The technical architecture of the study is where much of its novelty lies. The authors assembled twenty-five sets of genetic instruments, drawn either from cis-protein quantitative trait loci, which are variants that alter circulating protein levels, or from cis-expression quantitative trait loci, which are variants that alter gene expression in blood. These proxies covered the eighteen therapeutic genes and were tested against seventy-six prespecified safety outcomes drawn from the FinnGen R13 biobank release, a Finnish resource combining genetic data with national health registry records covering conditions from autoimmune disease to venous thromboembolism. In total, 1,900 tasks were planned, of which 1,824 proved analysable, generating a matrix of target-outcome hypotheses that could be evaluated in a statistically disciplined way.
A key methodological challenge was sparsity. Stringent selection of cis-acting variants, which are preferred because they are less prone to confounding by distant genetic effects, often leaves only one or a handful of usable instruments per gene. Conventional Mendelian randomisation methods that require many variants are therefore poorly suited to this setting. The workflow adapted by combining single-variant or few-variant Wald ratios with inverse-variance-weighted estimates where multiple variants were available, and then applying rigorous false-discovery-rate control. A global Benjamini-Hochberg correction across all 1,824 estimates served as the primary filter, with secondary corrections applied within forty narrower outcome domains. This two-tier approach balanced statistical conservatism against the risk of missing genuine signals buried in the multiple-testing burden.
The criteria for advancing a hypothesis were deliberately strict. An estimate had to pass the global false-discovery threshold, point in a risk-increasing direction under simulated target lowering, and lie outside the major histocompatibility complex, a densely variable genomic region notorious for producing misleading associations because of its extreme linkage disequilibrium. Signals meeting these conditions entered an objective follow-up phase involving sensitivity analyses, instrument-centred colocalisation and variant-level annotation. Colocalisation is the critical step: it asks whether the genetic variant associated with the molecular proxy, such as a protein level, and the variant associated with the disease outcome are in fact the same causal variant, rather than two nearby variants inherited together. The team used Bayesian colocalisation methods in windows of plus or minus 250 kilobases centred on each instrument, requiring a posterior probability of a shared causal variant above 0.5.
From the initial screen, one hundred estimates passed global false-discovery control and twenty more met only the narrow-domain threshold, producing fifteen tasks representing thirteen distinct target-outcome pairs for follow-up. The final classification stratified hypotheses by molecular layer, distinguishing protein-level P1 and P2 categories from expression-level E1 and E2 categories depending on whether colocalisation support exceeded the threshold. This stratification matters because a hypothesis supported at the protein level, where the drug will actually act, carries different translational weight than one supported only at the messenger RNA level, where the relationship between expression and protein abundance is often imperfect.
Two protein-level P1 hypotheses emerged as the strongest safety signals. The first linked lower ITGB7, the integrin targeted by vedolizumab, a widely used gut-selective biologic, with increased risk of multiple sclerosis and demyelinating disease, showing an odds ratio of 2.35 with a global q-value of 0.003 and colocalisation support of 0.935. The second connected TNFRSF1A, a tumour necrosis factor receptor, with myocardial infarction, an association replicated across two independent proteomic resources, the UK Biobank Pharma Proteomics Project and deCODE, with odds ratios of 1.56 and 1.37 respectively, both driven by the same variant, rs4149584, which encodes the R92Q amino-acid substitution. The convergence of two independent protein datasets on the same variant and the same outcome lends considerable credibility to the signal.
At the expression level, two E1 hypotheses stood out, both involving psoriasis, an outcome of particular interest because psoriasis shares inflammatory pathways with inflammatory bowel disease and is itself a comorbidity of the condition. Reduced IL23A expression was associated with dramatically increased psoriasis risk, with an odds ratio of 4.13 and an exceptionally strong q-value of 7.47 times ten to the minus nine, while reduced TYK2 expression showed an odds ratio of 1.53 with a q-value of 1.10 times ten to the minus eleven and near-perfect colocalisation support of 0.984. These findings are biologically coherent: IL-23 and TYK2 are the molecular foundations of some of the most successful recent therapies in both psoriasis and inflammatory bowel disease, and the genetic data suggest that their pathways are deeply intertwined with psoriasis susceptibility in ways that clinicians managing patients with both conditions should note.
Not every candidate signal survived scrutiny, and the framework’s treatment of these failures is as informative as its successes. IL12B-psoriasis and TNFSF15-lower-extremity deep-vein thrombosis were classified as P2 hypotheses because colocalisation posterior probabilities fell below the 0.5 threshold, meaning the apparent associations could not be confidently attributed to a shared causal variant. This is precisely the kind of auditable, transparent triage that distinguishes a disciplined hypothesis-generation framework from a fishing expedition. Every step, from the 1,900 planned tasks through the 2,584 SNP-level rows, heterogeneity analyses, weighted-median estimates and leave-one-out sensitivity checks, is documented in supplementary workbooks, allowing independent researchers to audit the provenance of each conclusion.
The broader significance of this work lies in its positioning within the drug development pipeline. Safety signals that emerge from human genetics are not deterministic predictions; they are prioritised hypotheses that can direct mechanistic validation in experimental systems and targeted evaluation in clinical safety datasets before large populations are exposed. By stratifying evidence by molecular proxy type, enforcing locus specificity through colocalisation, and excluding the confounded major histocompatibility complex region, the framework produces a short, defensible list of target-outcome pairs that merit attention. For a therapeutic area where biologics have transformed patient outcomes but questions remain about long-term risks such as demyelination, infection and cardiovascular events, the ability to generate such lists computationally, using only publicly available summary-level data, represents a meaningful step toward safer and faster drug development. The authors’ emphasis on auditability and reproducibility suggests the framework could be extended to other disease areas, turning the human genome into a routine preclinical screening tool for pharmacovigilance.
Subject of Research: Genetic prediction of safety outcomes for inflammatory bowel disease therapeutic targets using Mendelian randomisation and colocalisation
Article Title: A protein-prioritised, colocalisation-informed genetic framework for safety hypothesis generation across inflammatory bowel disease therapeutic pathways
Article References: Gao, X., Guo, X., Yang, S., & Zhang, Q. (2026). A protein-prioritised, colocalisation-informed genetic framework for safety hypothesis generation across inflammatory bowel disease therapeutic pathways. Journal of Translational Medicine. https://doi.org/10.1186/s12967-026-09068-z
Image Credits: AI Generated
DOI: 10.1186/s12967-026-09068-z
Keywords: inflammatory bowel disease, Mendelian randomisation, colocalisation, pQTL, eQTL, drug safety, FinnGen, ITGB7, TNFRSF1A, TYK2, IL23A, pharmacovigilance
News Source: Juliet Wilcox. (October 6, 2026). Genetic Framework Flags Safety Risks for Inflammatory Bowel Disease Drug Targets. Scienmag.



