The drug target discovery problem
Drug discovery and development is one of the most resource-intensive endeavors in research and innovation. Bringing a new drug therapy from early research to clinical use typically takes 10 to 15 years and costs on average $1 - to 2 billion per approved drug [1]. Despite these efforts, the rate of success is low. Approximately 90% of drug candidates that enter clinical trials ultimately fail during development or the approval process. This figure only includes candidates that reach phase I trials; many more drug candidates fail during the preclinical stage, meaning the overall failure rate across the drug discovery pipeline is even higher [1].
The process of creating a new therapeutic requires a long sequence of steps (Figure 1), starting from understanding the disease in question and ending in a new drug that treats patients. The first step is to gain sufficient insight into a disease so that a drug target can be identified. By understanding the disease and its effects mechanistically, an intervention can be envisioned that is subsequently tested and refined. Many drugs modulate the function of specific genes or proteins in the body; these are called the "drug targets." The discovery of these targets, i.e., drug target discovery, is the first step in drug discovery and development. After identifying the target, a new intervention can be designed. Many classes of biological intervention molecules can now be utilized to create a new drug, including small molecules, antibodies and antibody derivatives, CAR-T cells and other cell-based therapies, and gene therapy. All of these methods aim to interact with their target with high specificity so that the drug in development intervenes precisely and minimizes side effects. As drug development progresses, the costs per phase increase, in particular due to large-scale clinical trials.

The drug discovery process is frequently interrupted and therefore is generally seen as a funnel: one starts with many drug targets, and at every developmental stage, candidates drop out to the point where many drug discovery campaigns are unsuccessful (Figure 2). A large proportion of development costs are incurred due to these failures; dropped-out candidates never become approved drugs. Improving the ability to identify promising drug targets and molecules early in the discovery process is therefore critical to curb drug discovery costs. Approaches that help distinguish drug candidates that are more likely to succeed from those that are likely to fail can substantially improve drug development efficacy and ultimately increase the number of new medicines reaching patients #ref2">[2]

There are several reasons for the high attrition rates during drug discovery and development. Analyses of clinical trial outcomes between 2010 and 2017 suggest that the main causes of failure are, first, a lack of clinical efficacy (40-50%); second, unmanageable toxicity (around 30%); third, poor drug-like properties of the compound (10-15%); and, fourth, limited commercial viability or poor strategic planning (around 10%) [1]. Although these reasons are multifactorial, gains can be made by taking these factors into account from the earliest phases of drug target discovery onwards. Specifically, a key driver of the high new therapy attrition rate is the disconnect between the biological evidence supporting a drug target and the confidence with which that target enters clinical development. Retrospective analyses of clinical trial outcomes, including research from Open Targets, show that clinical trials are significantly more likely to succeed when the link between a disease and its drug target is supported by evidence from human genetics and genomics studies, such as genome-wide association studies (GWAS) [3]. Yet historically, much of this evidence has been fragmented across databases, inconsistently formatted, and difficult to integrate at scale.
This fragmentation has practical consequences for target identification in drug discovery and target validation in drug discovery. Improving drug target selection and prioritization is an active area of research, and many computational tools, datasets, and evidence frameworks have been developed to support the early phases of drug discovery and development. However, these resources are generally focused on specific aspects of drug discovery and do not integrate the various perspectives and data silos that are informative. In addition, drug discovery pipelines and platforms often require significant technical expertise; easy-to-use alternatives are usually based on black-box technology. This can create a disconnect between the dry-lab developers and maintainers of these systems and the wet-lab researchers who ultimately rely on insights generated by these methods. To fully leverage the potential of these tools, datasets, and frameworks, a more systematic and unified approach to integration is needed.
The Open Targets Platform was designed to address this gap precisely. It aims to bring together datasets and frameworks to support the early phases of drug target discovery and prioritization, making diverse evidence streams comparable, traceable, and actionable.
What is Open Targets?
The Open Targets Platform is an open-source, comprehensive tool for the systematic identification and prioritization of therapeutic drug targets. It integrates diverse, publicly available datasets, including those generated by the Open Targets Consortium, to build and score target-disease associations. The platform also provides critical annotations on targets, diseases or phenotypes, variants, studies, and drugs.
The platform has a strong focus on genetic evidence, including genome-wide association studies (GWAS) and molecular quantitative trait loci (molQTL), reflecting the established correlation between human genetic support and drug discovery campaign success. These data sources are complemented by more than 20 additional datasets spanning clinical evidence, somatic variation, pathway data, literature, RNA expression, and mouse models. All evidence is combined into a single association score that links diseases to genes as potential drug targets, with full data provenance visible for traceable, reproducible evaluation.
The Open Targets Platform is built by the Open Targets Consortium, a public-private partnership focused on integrating human genetics, genomics, and experimental data to improve drug target discovery. As of 2025, the consortium consists of partner institutions from both industry and academia, including EMBL-EBI, Wellcome Sanger Institute, Genentech, GSK, MSD, Pfizer, and Sanofi (Figure 3). Additionally, there is a wider community of collaborators and Open Targets users that regularly contribute ideas, features and code to the Open Targets Project, often through the Open Targets Community Forum.

This collaborative, pre-competitive model reflects a broader trend in biomedical research: as challenges grow more complex, scientific progress increasingly depends on aggregating expertise, data, and resources across organizational boundaries. Public-private partnerships such as Open Targets address limitations of isolated research by enabling larger-scale studies, shared risk and cost, and improved agility [4]. This is particularly important in computational drug target discovery, where individual organizations often have limited capacity to fully support the high costs of R&D, with many achieving only a few drug approvals per year [4].
The rationale for building the Open Targets Platform as open source is twofold. Pragmatically, the scientific and computational challenges of target identification in drug discovery and target validation in drug discovery is substantial undertaking for any single organization; these challenges are far easier to tackle collaboratively. By embracing open science, organizations can leverage the global bioinformatics community to build and maintain core infrastructure, allowing R&D teams to focus resources on analyzing results and advancing research.
Methodologically, open-source infrastructure establishes a transparent common ground for drug target validation, a standard or common reference of sorts. Historically, clinical failure rates were increased by different labs interpreting the same biological data through disparate, closed-source algorithms - a dynamic that has contributed to what 52% of researchers identify as a "significant reproducibility crisis" in the field [4]. Open-source methodologies allow researchers to validate, reproduce, and scrutinize each other's findings, thereby improving auditability for regulatory authorities, easing data delivery to authorities through open standards, and sharing general insights and standards that benefit the field as a whole. These collective benefits support more effective drug discovery and, ultimately, greater impact on patient wellbeing.
To ensure the scientific community always works with the most current information, the Open Targets Platform is actively maintained with quarterly updates. Data are accessible via a web interface, API, or full download. The processing pipeline for the integrated Open Targets data, called Gentropy, is also available as open-source software. By fostering openness and collaboration, Open Targets helps create a more efficient and transparent research ecosystem, ultimately supporting the discovery of better drug targets and enabling development of better treatments for patients.
The Open Targets Platform in Practice
The Open Targets Platform web interface enables researchers to investigate evidence for associations between diseases, targets, variants, and drugs through interactive evidence views and association rankings. It supports hypothesis generation, literature contextualization, and early-stage target assessment. The Platform consists of several key pages that offer specific insights into a topic of interest.
The Associations On The Fly page
The Platform's core functionality is the Associations On The Fly (AOTF) page (Figure 4). This page presents target-disease associations. Evidence from multiple data types, such as clinical evidence, genetic associations, pathways, RNA expression, mouse models, and literature-mining pipelines, is aggregated into association scores that help researchers identify biologically relevant and potentially druggable targets. Researchers can prioritize or exclude specific data sources or data types to fine-tune the overall score. The figure below shows the AOTF page for inflammatory bowel disease. Columns represent data sources (e.g., GWAS associations and ClinVar), and rows correspond to gene targets. To view details for a specific evidence item, researchers can click a circle to expand it and display individual evidence records linking the disease to the target.

Target prioritization page
In addition, researchers can use the Target Prioritization Factors overview to investigate favorable and unfavorable therapeutic characteristics of a target (Figure 5). This enables teams to narrow broad hypothesis spaces into smaller sets of high-priority candidates for downstream validation, supporting target validation in drug discovery.

Profile page
Another key strength of the Platform is its extensive visualization of annotation data sources that provide biological and clinical context for targets, diseases, and drugs (Figure 6). These annotations include protein function, pathway membership, subcellular localization, known drug mechanisms, disease classifications, safety liabilities, tractability assessments, and clinical development status. Rather than contributing directly to association scores, these datasets help researchers interpret and contextualize targets within broader translational and therapeutic workflows.

Exploring targets through the web interface and API
In addition to its web interface, the Platform supports programmatic access through a GraphQL API and bulk data downloads (Figure 7). These interfaces allow teams to integrate Open Targets data into internal pipelines, dashboards, notebooks, and machine learning workflows. Consequently, the Platform is used not only by academic researchers but also by biotechnology and pharmaceutical organizations building scalable drug target discovery pipelines.
Researchers can also run queries via the web interface, explore the data, and read documentation on the GraphQL schema.

Typical drug target discovery workflows
Genetics data plays a central role in many Open Targets workflows. The Platform integrates genome-wide association study (GWAS) results, fine-mapping results, credible sets, quantitative trait loci (QTL) data, and variant-to-gene mapping to support causal target identification. This is particularly valuable for interpreting non-coding variants and moving beyond simple nearest-gene assumptions. The following sections describe two common applications.
Workflow 1: exploring gene-disease relationships and prioritizing therapeutic targets
A common workflow begins with a disease or phenotype of interest (Figure 8). Researchers identify associated targets, inspect supporting evidence, and compare how strongly different targets are supported across evidence categories. By integrating evidence across genetics, expression, pathways, literature, and known drug mechanisms, the Platform provides a clear overview of the evidence landscape for a disease area, supporting target identification in drug discovery.
Workflow 2: drug repurposing and indication expansion
By integrating drug-target relationships and clinical evidence, the Platform also supports repurposing analyses and indication expansion (Figure 8). Researchers can investigate whether targets associated with one disease are already modulated by approved or investigational compounds in other therapeutic areas. The Platform also provides drug- or compound-specific information, including mechanisms of action, drug warnings, and pharmacovigilance. These data help identify shared biological mechanisms and generate hypotheses for therapeutic repositioning strategies.

Proprietary and customized Open Targets Platform deployments
While the Open Targets Platform provides a powerful foundation for translational research, the publicly available version of this Platform has a number of limitations. The available data is preselected and may not contain all available information for a particular therapeutic area. Also, data that does not have a permissive license is not available in the public Open Targets Platform. Larger or highly specialized organizations frequently have requirements that involve changes to the Open Targets Platform. Organizations generally deploy a proprietary Open Targets Platform instance for the following reasons:
- Add proprietary or organization-specific data. The public platform is limited to openly available datasets. Internal studies, proprietary data, and unpublished experimental results are not represented. As a result, organizations often complement the platform with internal evidence layers and proprietary analyses.
- Customize pipelines and scoring. The platform relies on standardized evidence integration and scoring pipelines designed to support broad usability and reproducibility. However, different organizations may require disease-specific prioritization strategies, custom weighting schemes, or alternative evidence models that are not fully captured by the Gentropy pipeline that is used for the public releases.
- Set up organization-specific workflows. Many research organizations also require integrations with internal infrastructure, knowledge graphs, artificial intelligence (AI) systems and AI-assisted prioritization tools, or secure collaborative environments that extend beyond the scope of the public platform.
- Incorporate company branding. Branding turns it from a tool researchers can use into infrastructure they rely on, accelerating adoption and standardizing target evaluation across the organization.
The setup, integration, and customization of proprietary Open Targets Platform instances is facilitated by the permissively licensed data and the fact that this is an open-source software project. The permissive Apache 2.0 license allows for changes and customization to the code that can either be shared with the wider Open Targets community or can be kept private.
The open character of the software not only allows for customization, but also promotes auditability, and ensures any proprietary data can always be moved out (i.e., no vendor lock-in). The availability of APIs and the use of open standards form a good starting point for integrating the Open Targets Platform into research departments' digital infrastructure to make the best use of its capabilities.
Examples of Platform customization
The Open Targets Platform can be customized in many directions: the user interface (UI) and visuals can be adapted, new datasets can be added, new evidence types can be integrated on the evidence page, or entirely new entities can be created. On the more technical side, connections to external databases or tools can also be established. For an example of what a customized instance could look like, take a look at our Open Targets demo platform where a selection of customizations can be explored.
Extending the data
Some organizations focus on a specific disease area and maintain their own datasets for associating diseases with genes. In such cases, these datasets can be added as evidence directly to the platform. One example is the integration of cBioPortal, for which a custom scoring logic was developed to support prioritization of gene targets. cBioPortal is an open-source tool focused on multidimensional cancer genomics datasets. Combining its comprehensive cancer type information with the drug target evidence in Open Targets brings together two complementary sources of insight, accelerating hypothesis generation and drug development. Adding cBioPortal particularly boosted the available copy number variation data as compared to the standard Open Targets oncology data sources. More information on this customization can be found here.
It is also possible to extend the platform with an entirely new data type, allowing organizations to make full use of the platform's potential. In this article, we show how we extended the platform with cell type data and how this enhances its utility (Figure 9). Different cell types have very different gene expression patterns and rely on different cell signaling pathways, meaning that drugs are not always effective across all cell types. Since diseases are often linked to a specific cell type or organ, this information can be used to improve drug selectivity and predict off-target effects.

Tailoring user experience
Updates to the UI are important for streamlining the research process and improving usability. On the demo server, we added dynamic filter buttons that allow users to include or exclude sections from the profile page, reducing visual overload and helping researchers focus on their primary area of interest (Figure 10).

Another custom view is the Metrics Page that provides insights into data coverage across different evidence types and compares data coverage between the current and previous Platform releases (Figure 11). It highlights key metrics such as the number of targets, diseases, and evidence types available in the platform.

Integrating in-house solutions and external tools
When deploying a private Open Targets instance, references and links to internal databases or platforms can be made available. Via an API, the data can be dynamically added and shown in the platform. Additionally, new tools can be layered on top of the platform, for example, Matomo for user tracking and analytics, or Keycloak, an open-source authentication provider that can integrate with an organization's existing identity provider. When aligned, this enables single sign-on access, allowing researchers to log in with their existing organizational credentials. In summary, Open Targets is a versatile drug discovery platform that integrates a large body of data that is relevant to drug target discovery. Due to its open nature, it fosters collaborations and allows the drug target discovery community to move ahead through shared innovation.
Glossary of Specialized Terms
| Term | Definition |
|---|---|
| API (Application Programming Interface) | A programmatic interface allowing software applications to communicate; the Open Targets Platform provides a GraphQL API for data access |
| Associations On The Fly (AOTF) | The core functionality of the Open Targets Platform that presents target-disease associations by aggregating evidence from multiple data types into association scores |
| Cell Type Data | Information about different cell types and their gene expression patterns, used to improve drug selectivity and predict off-target effects |
| ClinVar | A public archive of reports of the relationships among human variations and phenotypes, used as a data source in Open Targets |
| Credible Sets | Statistical sets of genetic variants identified through fine-mapping that are likely to contain the causal variant for a given association |
| Drug Repurposing | The process of identifying new therapeutic uses for existing or investigational drugs |
| Drug Target | A molecule (typically a protein or gene) in the body that a drug interacts with to produce its therapeutic effect |
| Drug Target Discovery | The process of identifying and validating biological molecules that can be modulated by drugs to treat disease |
| Data Provenance | Documentation of the origin and history of data, ensuring traceability and reproducibility |
| Evidence-Based Medicine | Medical practice informed by the best available evidence from systematic research |
| Evidence Datasets | Collections of data providing support for target-disease associations; Open Targets integrates more than 20 such datasets |
| Evidence Integration | The combination of multiple data types (genetics, expression, pathways, literature, etc.) to support target-disease associations |
| Fine-Mapping | A statistical method to narrow down genomic regions associated with traits to identify the most likely causal variants |
| Fine-Mapping Results | Output from statistical analyses that localize genetic associations to specific variants |
| Genome-Wide Association Study (GWAS) | A study design that examines many genetic variants across the genome to identify associations with traits or diseases |
| Genomics | The branch of molecular biology concerned with the structure, function, evolution, and mapping of genomes |
| Genetics Data | Information derived from the study of genes and heredity, including GWAS results, QTL data, and variant information |
| Gentropy | The platform's data harmonization and scoring pipeline |
| GraphQL | A query language for APIs that allows clients to request exactly the data they need |
| Human Genetics | The study of inherited characteristics and genetic variation in humans |
| Indication Expansion | The process of identifying new disease indications for existing drugs beyond their originally approved use |
| Inflammatory Bowel Disease (IBD) | An example disease used in the document to illustrate Open Targets Platform functionality |
| Keycloak | An open-source authentication provider that can integrate with organizational identity providers for single sign-on |
| Literature Mining | The automated extraction of information from scientific literature using computational methods |
| Matomo | An open-source analytics platform for user tracking |
| Metrics Page | A custom interface showing data coverage statistics across different evidence types and platform releases |
| Molecular Quantitative Trait Loci (molQTL) | Genetic variants associated with molecular phenotypes such as gene expression, protein levels, or metabolite concentrations |
| Multi-Omics Data | Data integrating multiple types of omics measurements (genomics, transcriptomics, proteomics, etc.) |
| NOD2 | Example gene/protein used to illustrate the Open Targets Profile page; a nucleotide-binding oligomerization domain-containing protein |
| Open Targets | A public-private consortium and platform integrating human genetics, genomics, and experimental data for drug target identification and prioritization |
| Open Targets Consortium | The partnership of industry and academic institutions that builds and maintains the Open Targets Platform |
| Open Targets Community Forum | A platform for users and collaborators to contribute ideas, features, and code to the Open Targets Project |
| Open Targets Platform | The free, open-source, open-data comprehensive tool for systematic identification and prioritization of therapeutic drug targets |
| Off-Target Effects | Unintended biological effects of a drug on targets other than the intended therapeutic target |
| Open Data Open Source | The philosophy of making software and data freely available for use, modification, and distribution |
| Pharmacogenetics | The study of how genetic variation affects individual response to drugs |
| Pharmacovigilance | The science and activities relating to the detection, assessment, understanding, and prevention of adverse effects or other drug-related problems |
| Phenotype | Observable characteristics or traits of an organism resulting from the interaction of its genotype with the environment |
| Private Deployment | Local installation of the Open Targets Platform allowing customization with proprietary data and integration with internal tools |
| Profile Page | A feature of the Open Targets Platform providing extensive visualization of annotation data for targets, diseases, and drugs |
| Public-Private Partnership | Collaboration between government/academic and commercial entities, as exemplified by the Open Targets Consortium |
| Quantitative Trait Loci (QTL) | Genomic regions containing variants that influence quantitative traits; includes expression QTL (eQTL), protein QTL (pQTL), which are generalized to molecular QTLs (molQTL) |
| RNA Expression | The process by which information from a gene is used to synthesize functional gene products (RNA), measured as a data type in Open Targets |
| Repurposing Analyses | Computational or experimental studies to identify new uses for existing drugs |
| Safety Liabilities | Potential adverse effects or toxicities associated with a drug target |
| Single Sign-On (SSO) | An authentication scheme allowing users to log in with a single ID to multiple related systems |
| Somatic Variation | Genetic alterations present in somatic (non-germline) cells, often relevant in cancer |
| Subcellular Localization | The specific location of a protein or molecule within a cell |
| Target Prioritization | The process of ranking potential drug targets based on multiple criteria to identify the most promising candidates for validation |
| Target Prioritization Factors | Criteria used to evaluate therapeutic characteristics of targets, including favorable and unfavorable properties |
| Target Validation | The process of confirming that a biological target is relevant to a disease and that modulating it will have therapeutic benefit |
| Target-Disease Association | A scored relationship linking a potential drug target to a disease based on integrated evidence |
| Target-Disease Association Scoring | The computational framework combining multiple evidence types into a single quantitative score |
| Tractability Assessment | Evaluation of how "druggable" a target is—whether it can be modulated by small molecules, antibodies, or other therapeutic modalities |
| Translational Research | Research that aims to "translate" findings from basic science into practical applications to improve human health |
| Variant-to-Gene Mapping | Computational methods linking genetic variants to their likely causal genes, particularly important for interpreting non-coding variants |
| Vendor Lock-in | Dependency on a single supplier; avoided by Open Targets through its open-source nature |
[1] Singh N, Vayer P, Tanwar S, Poyet JL, Tsaioun K, Villoutreix BO. Drug discovery and development: introduction to the general public and patient groups. Front Drug Discov. 2023 May 24;3:1201419. doi:10.3389/fddsv.2023.1201419
[2] Sun D, Gao W, Hu H, Zhou S. Why 90% of clinical drug development fails and how to improve it? Acta Pharm Sin B. 2022 Jul;12(7):3049–62. doi:10.1016/j.apsb.2022.02.002
[3] Tsepilov YA, Suveges D, Considine D, Szyszkowski S, Ge XJ, Santiago L, et al. The Human Pleiotropic Map of GWAS Associations and Therapeutic Implications.
[4] Hinkson IV, Madej B, Stahlberg EA. Accelerating Therapeutics for Opportunities in Medicine: A Paradigm Shift in Drug Discovery. Front Pharmacol. 2020 Jun 30;11:770. doi:10.3389/fphar.2020.00770