← Research building on SQLancer

Dawei Li, Yuxiao Guo, Qifan Liu, Jie Liang, Zhiyong Wu, Jingzhou Fu, Chi Zhang, Yu Jiang. 2025. International Conference on Automated Software Engineering.

Read the paper · doi:10.1109/ase63991.2025.00151

What this paper does with SQLancer

ARG integrates SQLancer to construct dynamic database schemas and populate them, and then measures itself against it, reporting 1017% more rewrite rules triggered and 15 more bugs in 24 hours. ARG fuzzes query rewriters, the components that transform a query into a faster but semantically equivalent form. The authors argue that general DBMS testing tools cover only a limited subset of rewrite scenarios given the diversity of rewrite rules. ARG uses abstract rules -- a unified representation of the AST patterns and constraints that trigger a rewrite, plus the resulting transformation -- as coverage feedback, steering generation toward patterns not yet exercised. Testing Apache Calcite, WeTune, SQLSolver and LearnedRewrite it found 38 previously unknown bugs. Written by claude-opus-5 from the 21 places this paper refers to SQLancer. The quotations below are the paper's own words, stored verbatim when the text was extracted.

How it was classified

uses infrastructure — yes (generator)

M5 states SQLancer was integrated to construct dynamic database schemas and populate them, so a generation component is reused while the rewrite-rule fuzzing is ARG's own.

To generate DDL statements and enhance abstract rule extraction, we integrated SQLancer to construct dynamic database schemas and populate test data. M5 · IV IMPLEMENTATION · page 6

extends technique — no

Abstract rule guided fuzzing is ARG's contribution; no SQLancer oracle is generalised.

compares with — yes

ARG reports triggering 1017% more rewrite rules and finding 15 more bugs than SQLancer over 24 hours.

In 24 hours, ARG triggered 76% and 1017% more written rules, triggered 13 and 15 more bugs than SQLsmith and SQLancer, respectively. M1 · page 1
In 24 hour experiment on these query rewriters, ARG triggered 76% and 1017% more written rules, triggered 13 and 15 more bugs than SQLsmith and SQLancer, respectively, In summary, we make the following contributions: •We observe that while query rewriters are widely used in DBMSs, bugs persist and can cause serious issues, yet effective testing tools remain lacking. M3 · I INTRODUCTION · page 2

describes as state of the art — no

M4 calls SQLancer a traditional generation-based fuzzer and argues it struggles on rewrite rules.

Where this differs from the pattern checks

The regular expressions that scan for these relationships are advisory. Where the reading above contradicts one, the reason is recorded.


SQLancer publications it cites (3)

Bibliography entries that resolved to a SQLancer publication, or to a paper by one of the project's authors. A sentence citing one of these numbers is a reference to SQLancer even when it never writes the name.

#EntryMatched as
18 RIGGER, M., ANDSU, Z. Detecting optimization bugs in database engines via non-optimizing reference engine construction. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symp... sqlancer publication · NOREC
19 RIGGER, M., ANDSU, Z. Finding bugs in database systems via query partitioning. Proc. ACM Program. Lang. 4, OOPSLA (Nov. 2020). sqlancer publication · TLP
22 SQLANCER .https://github .com/sqlancer/sqlancer .git, 2025. sqlancer publication

Every place it refers to SQLancer (21)

21 sentences, each stored verbatim from the extracted text with where it was found and how. “Citation marker” means the sentence names no tool at all and was reached through a reference number that resolved to a SQLancer publication.

Id Sentence Found by Where
M1 In 24 hours, ARG triggered 76% and 1017% more written rules, triggered 13 and 15 more bugs than SQLsmith and SQLancer, respectively. name
result comparison
page 1
M2 Many existing DBMS testing tools [ 22], [23], [31] effectively uncover logic bugs and crashes, but they are primarily designed for broad bug detection rather than targeting query rewrites. citation marker
background
I INTRODUCTION
page 1
M3 In 24 hour experiment on these query rewriters, ARG triggered 76% and 1017% more written rules, triggered 13 and 15 more bugs than SQLsmith and SQLancer, respectively, In summary, we make the following contributions: •We observe that while query rewriters are widely used in DBMSs, bugs persist and can cause serious ... name
result comparison
I INTRODUCTION
page 2
M4 Traditional generation-based fuzzers like SQLancer and SQLsmith create diverse SQL statements based on syntax rules but struggle to systematically activate specific rewriting rule combinations due to their random nature. name
motivation
II BACKGROUND ANDMOTIVATION
page 3
M5 To generate DDL statements and enhance abstract rule extraction, we integrated SQLancer to construct dynamic database schemas and populate test data. name
reuse component
IV IMPLEMENTATION
page 6
M6 This combined approach mitigates the limitations of each tool: SQLancer’s native FUZZER lacks sufficient coverage for complex rule patterns, while SQLsmith, despite its strength in generating intricate query structures, does not support adaptive schema generation. name IV IMPLEMENTATION
page 6
M7 Therefore, to evaluate the effectiveness of our approach, we compared ARG with two state-of-the-art SQL generators, SQLsmith and SQLancer, both widely used in the industry for generating large volumes of SQL queries that are fed into rewriters for transformation. name C Comparison with Other Techniques
page 8
M8 For a fair comparison, after SQLancer 1814 and SQLsmith generated SQL queries, we sent these queries to the target rewriter to collect coverage data and bug counts. name C Comparison with Other Techniques
page 8
M9 ARG significantly outperforms SQLsmith and SQLancer in terms of rule coverage across all four systems. name C Comparison with Other Techniques
page 9
M10 Specifically, ARG covered 18% and 15% more branches in total compared to SQLsmith and SQLancer, respectively. name C Comparison with Other Techniques
page 9
M11 TABLE IIINUMBER OFBRANCHES COVERED BYEACH TOOL IN 24HOURS Rewriter SQLsmith SQLancer ARG Calcite 7208 7191 9703 WeTune 2617 2727 2810 SQLSolver 2869 3045 3119 LearnedRewrite 8721 8995 9661 Total 21415 21958 25293 Improvement 18% ↑ 15%↑ The primary reason for ARG’s improved code coverage is its ability to cover more ... name C Comparison with Other Techniques
page 9
M12 Specifically, ARG covers a total of 76% and 1017% more abstract rules than SQLsmith and SQLancer, respectively. name C Comparison with Other Techniques
page 9
M13 ARG covers a total of 53% and 476% more unique rewriting rule combinations than SQLsmith and SQLancer, respectively. name C Comparison with Other Techniques
page 9
M14 TABLE IVNUMBER OFUNIQUE REWRITING RULECOMBINATIONS INREWRITER BYEACH TOOL IN 24HOURS Rewriter SQLsmith SQLancer ARG Calcite 7548 1824 11544 WeTune 289 44 324 SQLSolver 370 56 402 LearnedRewrite 814 469 1524 Total 9021 2393 13794 Improvement 53% ↑ 476%↑ Triggered Bugs. name C Comparison with Other Techniques
page 9
M15 As shown in the table, ARG detects 13 and 15 more bugs than SQLsmith and SQLancer, respectively. name C Comparison with Other Techniques
page 9
M16 We tracked the growth in the number of unique abstract rules extracted by SQLancer, SQLsmith, and ARG during a 24hour experiment. name C Comparison with Other Techniques
page 9
M17 Similarly, SQLancer focuses on generating SQL queries guided by predefined oracles and grammar constraints, making it difficult to explore the full spectrum of rewrite logic, especially those involving equivalence-preserving transformations. name C Comparison with Other Techniques
page 9
M18 TABLE VNUMBER OFTRIGGER BUGS BY EACHTOOL IN 24HOURS Rewriter SQLsmith SQLancer ARG Calcite 9 6 12 WeTune 5 8 11 SQLSolver 6 4 8 LearnedRewrite 5 5 7 Total 25 23 38 Increment 13 ↑ 15↑ D. name C Comparison with Other Techniques
page 9
M19 SQLancer [ 22] leverages predefined syntax tree models to produce grammatically valid SQL 1816 queries, while also creating schemas and populating the database with data. name C Comparison with Other Techniques
page 10
M20 NoREC [ 18] converts an optimizable query into a non-optimizable form of the query, then identifies logic bugs by comparing their execution results. technique C Comparison with Other Techniques
page 11
M21 Similarly, TLP [ 19] employs ternary logic to generate three equivalent queries combined via UNION operations. technique C Comparison with Other Techniques
page 11

This page is rendered from _data/papers/paper_doi_10_1109_ase63991_2025_00151.json, extracted from supplied pdf. 12 pages, 34 references parsed.