Benchmarking LLM query generation across SQL, Cypher and TypeQL
Explore our new LLM benchmark comparing MySQL, Neo4j, and TypeDB. Learn how context and retry budgets help AI agents self-correct and improve performance.

A report on the Claude Sonnet 5 / DeepSeek V4 Pro run of db-llm-bench
Summary
We ran a benchmark to test LLM query generation across multiple different database query languages. This benchmark attempts to understand, for a fixed task, what happens when you vary the LLM used, the database used, as well as information available in the prompt and how many times we allow a retry in the face of an error.
To do this, we used LLMs to generate queries based on natural-language questions. We chose an existing professionally curated dataset that exists in relational and graph models, Reactome, and also implemented it in TypeDB.
We scored the correctness of queries generated by Claude Sonnet 5 and DeepSeek V4 Pro, in 12 different configurations for a total of 3,024 runs. These configurations are designed to assess baseline query generation capabilities, the effect of in-context learning, and the benefits of including agentic loops.
With the right resources, TypeQL is able to match, or even beat SQL and Cypher for query generation. This stems from two general results:
In-context learning can almost entirely compensate for baseline training data deficit.
Neither model has substantial prior experience with or knowledge of TypeQL – with no in-context learning and no retries, Claude Sonnet 5 scores 62% and 58% on SQL and Cypher respectively, and 25% on TypeQL. DeepSeek V4 Pro was more stark – 56% and 62% on SQL and Cypher, and 0% on TypeQL. Adding in-context resources – skills and examples for all databases (still without retries) – sends Cypher up to 84% and 74% on Sonnet and DeepSeek respectively, SQL up to 82% and 78%, and TypeQL up to 75% and 62%.
Retrying pushes through errors, with better results coming from better errors.
Retries allow an LLM to self-correct errors, which yields significant accuracy improvements – in particular when combined with in-context learning. This signal depends on a verifier to return meaningful errors. The more powerful the verifier, the more effective the retry loop.
Allowing Sonnet 5 with in-context resources to retry errors up to 4 times pushes SQL up from 82% to 83%, Cypher from 74% to 82%, and TypeQL from 75% up to 92%. TypeDB benefits the most here – reaching the overall best result out of any configuration when provided with in-context resources and four retries – because it is a more powerful verifier: its schema also flags semantic failures (manifested as type check errors).
This blog does a deep dive into the benchmark and the data. It’s structured with section 1 describing the benchmark and section 2 its limitations. Section 3 gives the data behind each finding, section 4 our interpretation, section 5 how to reproduce the run, and section 6 what we recommend. The appendices hold additional data from the benchmark that supports the findings or may otherwise be of interest.
1. The benchmark
Reactome is a professionally curated biological pathway database. Its publishers ship each release as both a MySQL dump of the curation database and a Neo4j graph dump.We build the TypeDB database ourselves based on the Neo4j graph. The three query languages under test are SQL for MySQL, Cypher for Neo4j and TypeQL for TypeDB. The schema is large (the MySQL DDL alone is ~62KB, or ~16k tokens) and is placed in every prompt in full.
Each run works as follows. The model receives a prompt containing the schema for its target language, optional resources,and one question. It answers with a single query. We execute the query and compare the result to the expected answer for that question. A run is accurate only if the result matches. Reference queries in all three languages, and the expected answers, were authored by us for this dataset.
A retry is triggered only by a failure the harness can observe: a syntax or execution error, a timeout, a malformed result shape, or a response containing no query. The error is fed back to the model with the conversation so far, up to the retry budget. A query that executes and returns a wrong answer is terminal: the harness does not know the answer is wrong, any more than a real application would.
Questions
The 42 questions are split into tiers. Three are labelled by difficulty; six are labelled by the query construct they exercise; one tier measures whether the model recognises an unanswerable question and says so rather than guessing.
| tier | questions | exercises |
|---|---|---|
| easy | 4 | lookups and counts |
| medium | 5 | joins/traversals with filters |
| expert | 11 | multi-hop structure, negation, subqueries |
| recursion | 4 | transitive closure over the pathway hierarchy |
| reification | 1 | n-ary facts constrained on several roles |
| argmax | 1 | per-group extremes (with a tie in the answer) |
| aggregation | 2 | stacked aggregates |
| polymorphism | 11 | queries through the class hierarchy via a supertype |
| unanswerable | 3 | declining to answer |
The 3 unanswerable questions were correctly flagged as unanswerable in every tested configuration, and are excluded from our results. Increasing the complexity or plausibility of these questions is an interesting avenue to explore in the future: refusing to answer at the right time is a valuable property.
Due to this limitation, we’ve reported accuracy figures for the rest of the report discounting the unanswerable questions.
Configurations
Each run of the benchmark used one configuration over model, skills, examples, and retry budget:
- Models: Claude Sonnet 5 and DeepSeek V4 Pro.
- Skill on/off: whether a query-writing skill (a markdown document teaching the language) is included in the prompt. We use each language’s official or best publicly available skill: TypeDB’s official TypeQL skill (~33KB), the Neo4j contrib Cypher skill (~21KB), and Anthropic’s general SQL skill (~11KB); no database-specific SQL skill comparable to the other two exists. See §2 on how the skills compare.
- Few-shot examples: 0 or 5 worked question-and-query examples written for this dataset. See Appendix D for details.
- Retry budget: 0, 2 or 4.
Each query generation task is done 3 times per configuration to take into account LLM nondeterminism.
Per model and language, that is 12 configurations × 42 questions × 3 repetitions = 1,512 scored runs.
2. Limitations
- The dataset favours Cypher and SQL. Reactome designed its database for the relational and graph stores it publishes; our TypeDB database is a reconstruction of the Neo4j graph, so the structure is optimized for Cypher queries. A schema designed for TypeDB first would make fuller use of its hypergraph model.
- The question mix is weighted towards polymorphism. 11 of the 42 questions test polymorphic queries as many as the expert tier. TypeDB’s data model is designed to offer polymorphic queries and we wanted to measure LLMs understanding of itWe kept the weighting moderate so the set still explores the languages broadly, but the mix is our choice, and a different mix would shift the overall numbers. This is one reason we focus not on the absolute numbers, but rather the trends and effects observed.
- The skills are not equivalent. We used the best available skill for each language rather than degrading any to match the weakest. The TypeQL skill is the strongest of the three (official, current, and written for query generation), so part of TypeQL’s skill-enabled gain may reflect skill quality rather than the language’s teachability. The Cypher skill targets Cypher 25 and instructs a CYPHER 25 preamble that the benchmark’s Neo4j 5.26 rejects; we include the target Cypher version into the prompts to the LLM. The SQL skill is Anthropic’s general-purpose one; it teaches analytical warehouse dialects (PostgreSQL, Snowflake, BigQuery, Redshift, Databricks) but does not cover MySQL.
- Model selection is narrow. We paired a frontier-lab model with a strong open-weights model because we did not want to test only top-of-the-line frontier-lab models; no OpenAI or Google model is included, and results may differ across model families.
- The unanswerable tier is saturated (100% everywhere) and provides no discrimination in this run.
- Two tiers have one question each (argmax, reification). We report them with counts and treat them as indicative.
3. Results
3.1 Effect of in-context learning
First-attempt accuracy, with no retries, as the prompt gains a skill, five examples, or both:


Full data
| Claude Sonnet 5 | SQL | Cypher | TypeQL |
|---|---|---|---|
| no resources | 62.4% | 58.1% | 24.8% |
| skill | 71.8% | 1.7% | 62.4% |
| examples | 76.1% | 79.5% | 63.2% |
| skill + examples | 82.1% | 74.4% | 75.2% |
| DeepSeek V4 Pro | SQL | Cypher | TypeQL |
|---|---|---|---|
| no resources | 56.4% | 61.5% | 0.0% |
| skill | 60.7% | 63.2% | 48.7% |
| examples | 76.9% | 76.1% | 58.1% |
| skill + examples | 77.8% | 83.8% | 62.4% |
The data shows several interesting results.
Firstly, we observe the baseline performance for each model – this lines up with what we would expect considering the relative amounts of training data available for each language. That is to say, TypeQL greatly trails behind SQL and Cypher for both DeepSeek and Sonnet
Secondly, the skills have different effects. TypeQL benefits by far the most from the addition of a skill, gaining 38 percentage points on Sonnet and 49 on DeepSeek. Cypher and SQL benefit far less – with Cypher substantially dropping points on Sonnet, due to a versioning issue with the skill (discussed in §2). Later on, we will explore the addition of a retry loop that will resolve this issue.
Interestingly, domain-specific examples alone are worth about as much as the skill alone for TypeQL, and more than the skill for SQL and Cypher, where they are the larger of the two effects. This is likely down to the skill being more general-purpose – which the models already have knowledge on, while the examples provide domain-specific knowledge, which is where the LLM’s knowledge gaps for SQL and Cypher are.
Lastly, combining the general skill and domain specific examples is consistently helpful in every configuration. It benefits TypeQL the most, with it sitting 7 points behind SQL and 1 point ahead of Cypher for Sonnet, and 15 and 21 points behind them for DeepSeek, from bare-prompt deficits of 38 and 33 for Sonnet and 56 and 62 for DeepSeek.
3.2 Effect of a retry loop
We can also allow the LLM to retry query generation based on errors. Up to a variable limit, if the query produces an error such as a syntax error, we can feed the error message back into the LLM and allow it to try again. Here’s how that changes things for an otherwise bare run (that is, without any skills or examples):

Full data
| no resources | SQL | Cypher | TypeQL |
|---|---|---|---|
| Sonnet, 0 retries | 62.4% | 58.1% | 24.8% |
| Sonnet, 2 retries | 66.7% | 60.7% | 36.8% |
| Sonnet, 4 retries | 66.7% | 60.7% | 41.9% |
| DeepSeek, 0 retries | 56.4% | 61.5% | 0.0% |
| DeepSeek, 2 retries | 72.6% | 66.7% | 7.7% |
| DeepSeek, 4 retries | 76.9% | 68.4% | 16.2% |
We can combine the retry loop with the previously-discussed addition of skills and examples, to see what a fully-resourced run would look like.

Full data
| skill + examples | SQL | Cypher | TypeQL |
|---|---|---|---|
| Sonnet, 0 retries | 82.1% | 74.4% | 75.2% |
| Sonnet, 2 retries | 82.9% | 81.2% | 88.0% |
| Sonnet, 4 retries | 82.9% | 82.1% | 92.3% |
| DeepSeek, 0 retries | 77.8% | 83.8% | 62.4% |
| DeepSeek, 2 retries | 82.9% | 88.0% | 74.4% |
| DeepSeek, 4 retries | 84.6% | 88.9% | 85.5% |
We can see from this data that retries help less when a language is unknown – DeepSeek, which knows the language so little that it scores 0% without any assistance, only makes its way up to 16% even with 4 retries. This makes intuitive sense, since guessing at syntax tweaks is hard without a baseline understanding of a language. SQL on the other hand goes from 56% bare to 77% just from retries, almost matching the 78% from just adding resources.
We also observe that TypeQL gains greater benefit than SQL and Cypher as you continue allowing more retries. At full resources, for both Sonnet and DeepSeek, SQL and Cypher never gain more than 2 percentage points going from 2 retries to 4, while the same change in TypeQL provides 4 percentage points in Sonnet and 12 in DeepSeek.
3.3 Why retries help TypeQL more: failure visibility
Since the retry loop kicks in when the database returns an actual error, we’d expect to see that the systems that benefit the most from retries are those that are the strongest verifiers, i.e. can understand and enforce the most structure.
Let’s look at the data from Claude Sonnet 5 with skill and examples, share of the 117 runs per cell:
| Sonnet, first attempt | SQL | Cypher | TypeQL |
|---|---|---|---|
| correct | 82.1% | 74.4% | 75.2% |
| failed with a visible error | 0.9% | 10.3% | 23.1% |
| ran, returned a wrong answer | 17.1% | 15.4% | 1.7% |
| Sonnet, after 4 retries | SQL | Cypher | TypeQL |
|---|---|---|---|
| correct | 82.9% | 82.1% | 92.3% |
| failed with a visible error | 0.0% | 1.7% | 2.6% |
| ran, returned a wrong answer | 17.1% | 16.2% | 5.1% |
The same split for DeepSeek is in Appendix C and shows the same shape.
When a language provides errors, we see similar rates of recovery from retries. With zero retries, a fully-resourced Sonnet 5 produces 27 visible TypeQL failures, 12 visible Cypher failures, and only 1 visible SQL failure. Adding four retries resolves 20 of those TypeQL failures, 9 of the Cypher failures, and the 1 SQL failure.
While these rates of recovery are similar, the rates of error are not. TypeQL’s type system allows more failures to be errored failures, rather than simply incorrect values. Since TypeDB produces more errors overall, it can in turn benefit more from being able to retry those errors.
This is where we see the value of TypeQL’s type system come out. Since valid TypeQL queries are more restricted in what you’re allowed to write, we can error more frequently instead of accepting a semantically incorrect query. This then allows the LLM to correct its error, rather than getting an incorrect value back.
3.4 Token usage
When using LLMs, token usage is always a concern. Below, we see the mean output tokens per run with skill and examples, both at no retries and at four, over all 42 questions. In parentheses is the mean number of attempts made:
| model | retries | SQL | Cypher | TypeQL |
|---|---|---|---|---|
| Sonnet | 0 | 1,222 | 957 | 1,714 |
| Sonnet | 4 | 1,242 (1.01) | 1,434 (1.17) | 2,785 (1.39) |
| DeepSeek | 0 | 6,107 | 5,561 | 7,223 |
| DeepSeek | 4 | 7,967 (1.15) | 7,267 (1.20) | 17,029 (1.83) |
We see that TypeQL generally uses more tokens than SQL and Cypher, which we attribute to two different causes:
- Additional reasoning. Without retries, TypeQL costs 1.2 to 1.4 times the output of SQL and 1.3 to 1.8 times that of Cypher. Every run makes exactly one call here, so the difference is the length of the reply, which we attribute to the models reasoning longer over a language they have seen less of. DeepSeek’s totals are dominated by its reasoning trace, which is billed as output, so its ratio is smaller.
- Gaining accuracy through retries. TypeQL runs average 1.4 to 1.8 calls against 1.0 to 1.2 for SQL and 1.2 for Cypher, because TypeQL’s failures are visible and so trigger retries, while the others’ mostly are not. The extra calls buy the 17 to 23 accuracy points in §3.2. A SQL or Cypher run that returns a wrong answer costs one call and stays wrong.
We expect the single-shot gap to narrow as TypeQL 3.x enters training corpora, and fine-tuning on TypeQL would narrow it further; the numbers here come from two models that had to learn the language from a 33KB document in the prompt. The retry gap comes from where each language catches mistakes, and we expect it to persist.
4. Interpretation
With this data in mind, we can now return to the main points we highlighted at the start of this piece, as well as a few extra points of interest.
Firstly, we’ve observed the effect of in-context learning, particularly for languages with small training corpora. A skill and a few examples is enough to close the majority of the knowledge gap on a single-shot query generation task. For a team choosing a database to put behind an LLM, how well the language can be taught in context matters as much if not more than how much of it the model has already seen.
We’ve also viewed the way retry loops are made more powerful by better errors, and more frequent errors. A more heavily structured language like TypeQL leaves a model more likely to produce an incorrect query with an explicit, fixable, error than it is to produce a query that silently produces an incorrect result with no way to identify it as such.
We should also note here that this has all been done in a well-curated dataset with clear and consistent naming conventions used throughout. On a messier dataset we’d expect many of these gaps to widen between a language like TypeQL, that will outright reject an incorrectly named node type, and a language like Cypher that will accept it.
5. Recommendations
We believe this data shows a few important recommendations, particularly for anyone attempting to generate in a language that’s under-represented in training data.
- Add general advice on your language. The TypeQL skill on its own took Sonnet’s single-shot accuracy from 25% to 62% and cut syntax errors from a third of attempts to 2%.
- Add a handful of worked examples specific to your use-case. Five examples were worth as much again as the skill, and the two together reached 75% single-shot.
- Run generation inside a loop that executes the query and feeds any error back. Two retries were worth 13 more points and four were worth 17, because TypeDB rejects most wrong queries before they run. Account for these retries in your token budgets: TypeQL used on average 1.4 calls per query for Sonnet.
All these steps are most effective for a language like TypeQL. The first two provide dramatically greater benefits for languages that aren’t already known, and the third depends on having good errors in a structured language that avoids silent failures where possible.
6. Reproducibility
The benchmark runner, dataset build scripts, prompts, vendored skills, questions with per-language reference queries, the full results file for this run, the analysis scripts that produce every table above, and the script that renders the charts are in the db-llm-bench repository.
Please feel free to reach out to community@typedb.com for questions or comments, or hop into our Discord server.
Appendices
We include appendices of further data below.
Appendix A. Full accuracy matrix
Every configuration for each model, answerable questions, 117 runs per cell.
| Claude Sonnet 5 | retries | SQL | Cypher | TypeQL |
|---|---|---|---|---|
| no resources | 0 | 62.4% | 58.1% | 24.8% |
| no resources | 2 | 66.7% | 60.7% | 36.8% |
| no resources | 4 | 66.7% | 60.7% | 41.9% |
| skill | 0 | 71.8% | 1.7% | 62.4% |
| skill | 2 | 74.4% | 67.5% | 72.6% |
| skill | 4 | 74.4% | 67.5% | 74.4% |
| examples | 0 | 76.1% | 79.5% | 63.2% |
| examples | 2 | 76.9% | 81.2% | 75.2% |
| examples | 4 | 76.9% | 81.2% | 82.1% |
| skill + examples | 0 | 82.1% | 74.4% | 75.2% |
| skill + examples | 2 | 82.9% | 81.2% | 88.0% |
| skill + examples | 4 | 82.9% | 82.1% | 92.3% |
| DeepSeek V4 Pro | retries | SQL | Cypher | TypeQL |
|---|---|---|---|---|
| no resources | 0 | 56.4% | 61.5% | 0.0% |
| no resources | 2 | 72.6% | 66.7% | 7.7% |
| no resources | 4 | 76.9% | 68.4% | 16.2% |
| skill | 0 | 60.7% | 63.2% | 48.7% |
| skill | 2 | 76.9% | 75.2% | 59.0% |
| skill | 4 | 79.5% | 75.2% | 65.0% |
| examples | 0 | 76.9% | 76.1% | 58.1% |
| examples | 2 | 84.6% | 82.9% | 71.8% |
| examples | 4 | 84.6% | 83.8% | 82.1% |
| skill + examples | 0 | 77.8% | 83.8% | 62.4% |
| skill + examples | 2 | 82.9% | 88.0% | 74.4% |
| skill + examples | 4 | 84.6% | 88.9% | 85.5% |
Including the three unanswerable questions, which every configuration answered correctly, the fully resourced four-retry figures are Sonnet 84.1% / 83.3% / 92.9% and DeepSeek 85.7% / 89.7% / 86.5%.
Appendix B. Claude Sonnet 5 by question tier
Claude Sonnet 5 on each tier, comparing the bare prompt with no retries against the fully resourced prompt with four retries. Tiers are ordered as in §1; the counts are in the table below.

Observations:
- Fully resourced, every language reaches 100% on easy, medium, recursion, reification and aggregation. The recursion tier, where bare TypeQL scored 0 of 12 because neither model knew TypeQL’s recursive functions, shows a construct that in-context resources teach completely.
- Polymorphism is the tier with the largest gap between the languages, with TypeQL ahead: 88% against 55% for both SQL and Cypher. These questions range over Reactome’s class hierarchy through a supertype. In TypeQL the supertype is a schema object ($x has name $n matches every subtype of name), so the query can be derived from the schema; in SQL the same question requires enumerating the per-class tables. On the widest such question (objects with any kind of name containing “PIK3” whose display name does not), the SQL reference query is a seventeen-branch UNION ALL, and no model produced a correct SQL or Cypher answer in any fully resourced run, against 4 of 6 for TypeQL.
- Argmax (one question; small sample) is TypeQL’s worst tier at 67% against 100%. The expected answer is a two-way tie, and a tie-safe per-group extreme requires re-deriving the pipeline twice in TypeQL or writing a user-defined function. Models instead generate sort … limit 1 and return one of the two correct people. SQL window functions and Cypher’s collect express the tie-safe version directly. This is an area of future language development for TypeQL.
- Expert is the only tier where the fully resourced languages finish a few points apart: TypeQL 88%, SQL 85%, Cypher 82%.
“Bare” is no resources and no retries; “full” is skill, examples and four retries.
| tier | SQL bare | SQL full | Cypher bare | Cypher full | TypeQL bare | TypeQL full |
|---|---|---|---|---|---|---|
| easy | 75.0% (9/12) | 100% (12/12) | 100% (12/12) | 100% (12/12) | 50.0% (6/12) | 100% (12/12) |
| medium | 80.0% (12/15) | 100% (15/15) | 86.7% (13/15) | 100% (15/15) | 46.7% (7/15) | 100% (15/15) |
| expert | 60.6% (20/33) | 84.8% (28/33) | 45.5% (15/33) | 81.8% (27/33) | 18.2% (6/33) | 87.9% (29/33) |
| recursion | 58.3% (7/12) | 100% (12/12) | 66.7% (8/12) | 100% (12/12) | 0.0% (0/12) | 100% (12/12) |
| reification | 66.7% (2/3) | 100% (3/3) | 100% (3/3) | 100% (3/3) | 66.7% (2/3) | 100% (3/3) |
| argmax | 100% (3/3) | 100% (3/3) | 66.7% (2/3) | 100% (3/3) | 0.0% (0/3) | 66.7% (2/3) |
| aggregation | 100% (6/6) | 100% (6/6) | 50.0% (3/6) | 100% (6/6) | 0.0% (0/6) | 100% (6/6) |
| polymorphism | 42.4% (14/33) | 54.5% (18/33) | 36.4% (12/33) | 54.5% (18/33) | 24.2% (8/33) | 87.9% (29/33) |
Appendix C. DeepSeek V4 Pro failure modes
The §3.3 split for DeepSeek, skill and examples present, 117 runs per cell:
| DeepSeek, first attempt | SQL | Cypher | TypeQL |
|---|---|---|---|
| correct | 77.8% | 83.8% | 62.4% |
| failed with a visible error | 11.1% | 12.8% | 33.3% |
| ran, returned a wrong answer | 11.1% | 3.4% | 4.3% |
| DeepSeek, after 4 retries | SQL | Cypher | TypeQL |
|---|---|---|---|
| correct | 84.6% | 88.9% | 85.5% |
| failed with a visible error | 0.9% | 4.3% | 6.0% |
| ran, returned a wrong answer | 14.5% | 6.8% | 8.5% |
Retries recovered 27 of DeepSeek’s 39 visible TypeQL failures, 6 of 15 in Cypher and 8 of 13 in SQL, and none of its 22 silent failures. Retries help DeepSeek’s SQL more than Sonnet’s (7 points against 1) because more of its SQL failures are visible; in Cypher the two models gain similar amounts (5 and 8 points), and in TypeQL 23 and 17.
Appendix D. Few-shot Examples
Each of the 5 few-shot examples is a question/answer pair, that is to say a question and the query (in the appropriate language) that answers it. We use the same five questions for all languages, for example one of the questions is “How many pathways are assigned to Homo sapiens?”. Which is answered by the following queries:
SQL:
SELECT COUNT(DISTINCT es.DB_ID)
FROM Event_2_species es
JOIN Pathway p ON p.DB_ID = es.DB_ID
JOIN DatabaseObject d ON d.DB_ID = es.species
WHERE d._displayName = 'Homo sapiens';
Cypher:
MATCH
(p:Pathway)-[:species]->(s:Species {displayName:'Homo sapiens'}) RETURN count(DISTINCT p)
TypeQL:
match
$p isa pathway;
species-assignment (classified-thing: $p, species: $s);
$s has display-name "Homo sapiens";
select $p; distinct;
reduce $count = count;
When examples are enabled, we would provide all five question/answer pairs into the LLM’s prompt.
