AI-Generated OQL: Valid Syntax, Different Meaning

In an earlier article, The Hyve introduced the AI OQL-Helper, now called the OQL Gene Assistant. It is integrated into cBioPortal, and translates natural-language requests into Onco Query Language (OQL). This can make advanced genomic queries more approachable for new users and speed up repetitive query formulation for experienced users.

Individual examples demonstrated the potential of this approach. The next milestone was to test it systematically: how reliably can different language models translate a broad range of cancer-genomics requests into OQL, where do errors occur, and how should generated queries be evaluated before they are used to define patient cohorts?

This article looks back at the completed benchmark project, developed as part of a master's research project at The Hyve. It outlines what was built, the scope of the evaluation, the main baseline findings and the practical lessons for the OQL Gene Assistant. A separate follow-up article will examine how prompt design and temperature settings affected performance.

From an experimental feature to a systematic benchmark

A few successful demonstrations cannot show whether a system is reliable across the full range of requests users may make. The project therefore created a reusable benchmark for natural-language-to-OQL generation rather than testing only a small collection of showcase examples.

The benchmark contains 94 natural-language requests paired with reference OQL queries. The set combines researcher-style questions, examples derived from the OQL documentation and additions from domain experts. It covers mutations, copy-number alterations, expression, protein abundance and fusions, as well as more detailed constructions involving labels, exclusions, ranges and gene-specific constraints. Figure 2 shows the complete baseline comparison across all 25 models, including the variation observed across the five replicate runs.
 

https://www.thehyve.nl/uploads/Screenshot-2026-09-29-at-16.30.01.png
Figure 2. Baseline syntactic validity and semantic exact-match performance across 25 language models. Small symbols show the five individual runs for each model, while diamonds and error bars show the mean and standard deviation. Boxes summarize the middle 50% of runs, with the central line marking the median. Models are ordered by mean semantic exact match.

In the main comparison, 25 language models from seven providers were evaluated using the same baseline prompt. Each main configuration was run five times, allowing the project to measure not only average performance but also variation between repeated outputs. Latency and cost were recorded alongside quality, although this article focuses on reliability.

Evaluating more than whether a query runs

Every generated query passed through a layered evaluation pipeline. First, the cBioPortal OQL parser checked whether the output was syntactically valid. This answers an important but limited question: is the query well formed and executable?

Next, a semantic evaluator compared the biological constraints in the generated query with those in the reference query. Equivalent formatting differences were normalized, while omitted, added or reassigned constraints remained errors. This made it possible to distinguish harmless surface variation from a change that could alter the biological question or selected cohort.

Finally, two case studies examined the consequences at cohort level. Generated queries for KRAS G12C in colorectal cancer and BRAF V600E in melanoma were executed against cBioPortal data and compared with the intended drug-eligibility cohorts.

https://www.thehyve.nl/uploads/Screenshot-2026-09-24-at-16.01.23.png
Figure 1. Overview of the OQL benchmark evaluation pipeline. Model outputs were first checked using the cBioPortal OQL parser. Syntactically valid queries were converted into structured fields and compared with the reference OQL. For selected case studies, reference and generated queries were executed in cBioPortal and the returned cohorts were compared. Cohort symbols are illustrative; sample counts are not to scale.

What the baseline comparison showed

The 25 models showed a wide range of reliability profiles. Under the shared baseline prompt, mean syntactic validity ranged from 81.1% to 98.7%. Semantic exact match - requiring all normalized biological constraints to be preserved - ranged from 61.1% to 82.6%. No model was perfect, and the leading semantic models formed a close cluster with overlapping confidence intervals, so small differences at the top should not be interpreted as a definitive ranking.

The more informative result was how differently syntax and meaning could behave. DeepSeek V3.1 and Ministral 3B produced valid OQL at nearly the same rate: 94.5% and 93.4%, respectively. Their semantic exact-match scores were much further apart at 79.4% and 61.3%. A syntax-only evaluation would have hidden an 18.1-percentage-point difference in how often the models preserved the requested constraints.

When a small query difference changes the cohort

Consider two valid OQL queries:

BRAF: MUT=V600E
BRAF: MUT

Both queries parse correctly, but they do not ask the same biological question. The first selects the specific BRAF V600E alteration; the second selects all BRAF mutations. The difference is visually small, yet it can substantially change the returned patient group.

The cohort case studies showed the same issue in practice. In the KRAS example, one model broadened the requested single-variant cohort into a multi-gene pathway query in four of five runs, increasing the cohort four- to fivefold. In the BRAF example, another model repeatedly produced a non-existent alteration token and returned no patients. An empty or plausible-looking result could be misread as a biological finding when it is actually a translation error. Across the two case studies, most observed failures broadened rather than narrowed the intended cohort.

What this project milestone adds

The project produced more than a leaderboard. It delivered a curated OQL benchmark, a repeatable evaluation pipeline and a way to inspect errors at syntax, meaning and cohort level. Together, these provide a foundation for comparing future models and configurations using the same criteria.

The results also clarify what reliable assistance requires. A parser remains valuable because malformed output should be rejected immediately. During development, semantic checks and representative cohort tests are needed to detect valid-looking errors. When a request is underspecified, the safer response may be to ask a clarifying question rather than guess. In the user interface, generated OQL should remain visible and editable so that the researcher can review it before execution.

A separate question: how should the configuration be optimized?

The baseline comparison established that models have different reliability profiles. The project also tested how prompt design and inference settings, including temperature, changed those profiles. Those experiments answer a separate practical question - how to improve a model configuration once the evaluation framework is in place - and will be covered in a follow-up article.

For this milestone, the relevant lesson is that a model name alone does not define system performance. Natural-language-to-OQL generation should be evaluated as a complete configuration, using the prompt and settings intended for deployment.

From promising output to reviewed assistance

The benchmark confirms that large language models can generate useful OQL across a broad and technically varied set of requests. It also provides a clearer picture of the remaining reliability challenge: outputs that fail syntax are easy to catch, while outputs that preserve the grammar but change the meaning can be harder to recognize.

The practical question is therefore not only, 'Does this query run?' It is also, 'Does this query still mean what I asked?' Measuring both gives the OQL Gene Assistant a stronger foundation for continued development and responsible use in cBioPortal.

Try the OQL Gene Assistant

Explore the OQL Gene Assistant on The Hyve's cBioPortal demo. Treat each generated query as a proposal: review the OQL and confirm that it preserves your intended cohort before running it. Interested in reliable AI-assisted querying or other cBioPortal extensions? Contact The Hyve to discuss your use case.

Let's start collaborating

The Hyve provides services to develop, extend, and improve features in cBioPortal. If you have a desired feature in mind to improve your cancer genomics data exploration with cBioPortal, feel free to contact us.

Fill in the form and we will get in touch