quack-rs: DuckDB extensions in Rust
A Rust SDK for building DuckDB loadable extensions — no C++ required.
What is quack-rs?
quack-rs is a Rust SDK for building DuckDB loadable extensions
without C++ or CMake. It wraps DuckDB's C Extension API — the C interface DuckDB exposes
to loadable extensions — in safe builders and RAII types, and guards against the FFI
pitfalls documented in the Pitfall Catalog, so you can focus on
extension logic.
DuckDB's own documentation acknowledges the gap:
"Writing a Rust-based DuckDB extension requires writing glue code in C++ and will force you to build through DuckDB's CMake & C++ based extension template. We understand that this is not ideal and acknowledge the fact that Rust developers prefer to work on pure Rust codebases."
quack-rs closes that gap. No C++. No CMake. No glue code.
What you can build
| Extension type | quack-rs support |
|---|---|
| Scalar functions | ✅ ScalarFunctionBuilder |
| Overloaded scalars | ✅ ScalarFunctionSetBuilder |
| Aggregate functions | ✅ AggregateFunctionBuilder |
| Overloaded aggregates | ✅ AggregateFunctionSetBuilder (per-overload return types) |
| Table functions | ✅ TableFunctionBuilder (raw) + TypedTableFunctionBuilder<S> (closure-based, typed scan state) |
| Cast / TRY_CAST functions | ✅ CastFunctionBuilder |
| Replacement scans | ✅ ReplacementScanBuilder |
| SQL macros (scalar) | ✅ SqlMacro::scalar |
| SQL macros (table) | ✅ SqlMacro::table |
Copy functions (COPY TO / COPY FROM) | ✅ CopyFunctionBuilder (requires duckdb-1-5) |
Note: Window functions have no counterpart in DuckDB's public C Extension API and cannot be implemented from Rust (or any language) via that API. See Known Limitations.
Why does this exist?
quack-rs was extracted from
duckdb-behavioral, a production DuckDB
community extension. Building that extension revealed 16 undocumented pitfalls in DuckDB's
Rust FFI surface — struct layouts, callback contracts, and initialization sequences that
aren't covered anywhere in the DuckDB documentation or libduckdb-sys docs.
Three of those pitfalls caused extension-breaking bugs that passed 435 unit tests before being caught by end-to-end tests:
- A SEGFAULT on load (wrong entry point sequence)
- 6 of 7 functions silently not registered (undocumented function-set naming rule)
- Wrong aggregate results under parallel plans (combine callback not propagating configuration fields to fresh target states)
quack-rs rules out the first two: entry_point! performs the correct initialization
sequence, and the function-set builders name every member. The third lives in your own
combine callback, so no API can prevent it; AggregateTestHarness::combine lets you test
for it without DuckDB. The full list is in the Pitfall Catalog.
Key features
- Zero C++ — no
CMakeLists.txt, no header files, no glue code - Every function kind the C API can register — scalar, aggregate, table, cast, replacement scan, copy function (
duckdb-1-5) — plus SQL macros - Panic-safe FFI — the entry point and the callbacks quack-rs generates catch panics and report them as SQL errors; registration errors surface via
Result - RAII memory management —
LogicalTypeandFfiState<T>prevent leaks and double-frees - Type-safe builders —
ScalarFunctionBuilder,AggregateFunctionBuilder,TableFunctionBuilder,CastFunctionBuilder,ReplacementScanBuilder - SQL macros — register
CREATE MACROstatements without any FFI callbacks - Testable state —
AggregateTestHarness<T>tests aggregate logic without a live DuckDB - Scaffold generator — generates a complete community extension project (
Cargo.toml,Makefile, CI,description.yml, tests) from one function call - 31 pitfalls documented — every known DuckDB Rust FFI pitfall, with symptom, root cause and fix
Navigation
New to DuckDB extensions? → Start with Quick Start
Adding quack-rs to an existing project? → See Installation
Writing your first function? → See Scalar Functions or Aggregate Functions
Want SQL macros without FFI callbacks? → See SQL Macros
Submitting a community extension? → See Community Extensions
Something broke? → See Pitfall Catalog
Quick Start
This page takes you from an empty crate to a DuckDB loadable extension written in Rust, loaded into the DuckDB CLI, in three steps.
Prerequisites
Step 1 — Add quack-rs to your extension
In your extension's Cargo.toml:
[dependencies]
quack-rs = "0.18"
libduckdb-sys = { version = ">=1.4.4, <2", features = ["loadable-extension"] }
[lib]
name = "my_extension" # must match your extension name — see Pitfall P1
crate-type = ["cdylib", "rlib"]
[profile.release]
panic = "unwind" # required — quack-rs's panic guards need unwinding (see Installation)
lto = true
opt-level = 3
codegen-units = 1
strip = true
Starting from scratch? The scaffold generator generates a complete community extension project, including this
Cargo.toml.
Step 2 — Write the extension
#![allow(unused)] fn main() { // src/lib.rs use quack_rs::entry_point; use quack_rs::error::ExtensionError; use quack_rs::scalar::ScalarFunctionBuilder; use quack_rs::types::TypeId; use quack_rs::vector::{VectorReader, VectorWriter}; use libduckdb_sys::{duckdb_connection, duckdb_function_info, duckdb_data_chunk, duckdb_vector}; /// Scalar function: double_it(BIGINT) → BIGINT unsafe extern "C" fn double_it( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { // SAFETY: input is a valid data chunk provided by DuckDB. let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; let row_count = reader.row_count(); for row in 0..row_count { if unsafe { !reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } let value = unsafe { reader.read_i64(row) }; unsafe { writer.write_i64(row, value * 2) }; } } fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { ScalarFunctionBuilder::new("double_it") .param(TypeId::BigInt) .returns(TypeId::BigInt) .function(double_it) .register(con)?; } Ok(()) } entry_point!(my_extension_init_c_api, |con| register(con)); }
Step 3 — Build and test
# Build the extension
cargo build --release
# DuckDB refuses to LOAD a bare .so: append the metadata footer first.
# quack-rs ships the `append_metadata` binary for this step.
cargo install quack-rs --bin append_metadata
append_metadata target/release/libmy_extension.so my_extension.duckdb_extension \
--abi-type C_STRUCT --extension-version v0.1.0 \
--duckdb-version v1.2.0 --platform linux_amd64
# Load it; -unsigned allows a locally built, unsigned extension.
duckdb -unsigned -c "LOAD './my_extension.duckdb_extension'; SELECT double_it(21);"
# ┌───────────────┐
# │ double_it(21) │
# │ int64 │
# ├───────────────┤
# │ 42 │
# └───────────────┘
macOS: the library is
libmy_extension.dyliband the platformosx_arm64(orosx_amd64). Windows:my_extension.dllandwindows_amd64.
What's next?
- Learn how DuckDB calls your extension: Extension Anatomy
- Add an aggregate function: Aggregate Functions
- Add SQL macros without any callbacks: SQL Macros
- Generate a complete community extension project: Project Scaffold
Installation
This page covers the Cargo.toml dependencies and release-profile settings a DuckDB
loadable extension built with quack-rs needs, the minimum Rust version, and the optional
test features.
Adding quack-rs to an existing extension
Add the following to your extension's Cargo.toml:
[dependencies]
quack-rs = "0.18"
libduckdb-sys = { version = ">=1.4.4, <2", features = ["loadable-extension"] }
Why
>=1.4.4, <2? Every DuckDB 1.4.x and 1.5.x release loads extensions built for C API versionv1.2.0(1.5.6 declaresv1.5.6and accepts every earlier one), soquack-rssupports both with a single bounded range. The<2upper bound prevents silent adoption of a future major release whose C API may change in breaking ways — making any such upgrade an explicit, auditable decision. See Extension Anatomy.
Required Cargo.toml settings
Every DuckDB extension requires specific Cargo settings to link and behave correctly:
[lib]
name = "my_extension" # ← must match extension name exactly (Pitfall P1)
crate-type = ["cdylib", "rlib"]
# ^^^^^^ cdylib produces the .so/.dylib/.dll DuckDB loads
# rlib optional: lets doctests, examples and tests/ link the crate
[profile.release]
panic = "unwind" # REQUIRED — quack-rs catches panics at every FFI boundary;
# "abort" makes that impossible (see below)
lto = true # recommended — reduces binary size, improves performance
opt-level = 3 # recommended
codegen-units = 1 # recommended — better optimisation, slower build
strip = true # recommended — reduces binary size
Why panic = "unwind", not "abort"?
Every callback quack-rs generates (the *_callback! macros, the closure-based builders,
FfiState's callbacks), and every entry point, runs your code inside
std::panic::catch_unwind and turns a panic into an ordinary SQL error that DuckDB
reports to the user. catch_unwind can only catch a panic that unwinds: under
panic = "abort" the process terminates at the panic site, before any guard runs, taking
the user's whole DuckDB session with it.
A raw unsafe extern "C" fn that you pass to a builder yourself (such as double_it in the
Quick Start) is installed as written, with no guard. Define it with the
matching macro (scalar_callback!, aggregate_update_callback!, …), which reports a panic
as a SQL error, or wrap its body in quack_rs::callback::catch_ffi_panic and report the
Err it returns yourself.
(A panic that escapes an extern "C" function without being caught is not undefined
behaviour on Rust ≥ 1.81 — the runtime aborts the process — but that is exactly the
outcome the guards exist to prevent.)
validate_release_profile rejects panic = "abort", and the scaffold generator emits
panic = "unwind".
Minimum Supported Rust Version
quack-rs requires Rust ≥ 1.86.0.
1.86.0 is a ceiling as much as a floor. DuckDB's community-extension build
workflow (_extension_distribution.yml in duckdb/extension-ci-tools) pins
Rust 1.86.0 for its WebAssembly jobs, so an extension — and therefore
quack-rs — must build on it; CI's msrv-vs-duckdb-ci job re-derives that pin
and fails if the MSRV rises above it. The msrv job checks the crate with
cargo +1.86.0 check, and the benchmark dev-dependency (criterion) needs
1.86 as well.
Install or update via:
rustup update stable
rustup default stable
Verify:
rustc --version # must be ≥ 1.86.0
Development dependencies
To run SQL against your functions inside cargo test, enable one of quack-rs's test
features as a dev-dependency:
[dev-dependencies]
# Compiles DuckDB from C++ source: no setup, slow cold build.
quack-rs = { version = "0.18", features = ["bundled-test"] }
# ...or link a prebuilt libduckdb instead (set DUCKDB_DOWNLOAD_LIB=1 or DUCKDB_LIB_DIR):
# quack-rs = { version = "0.18", features = ["bundled-test-prebuilt"] }
Either one initialises the loadable-extension dispatch table from the linked DuckDB, so
testing::InMemoryDb works and the whole C API — including your own registration code — can
be exercised in a test. Without them, any duckdb_* call in a cargo test process panics,
because nothing has filled the dispatch table. See the Testing Guide.
Starting a new extension from scratch
Use the scaffold generator to produce a complete project with these settings, the build files and CI already in place.
Your First Extension
This page builds a DuckDB extension in Rust step by step, using hello-ext, the example
extension bundled with quack-rs. The table lists four of the functions it registers, one of
each major kind (the full list is in its
README); this page walks
through the aggregate and the scalar. The table function and the cast are covered in
Table Functions and
Cast Functions.
| SQL | Kind | Signature |
|---|---|---|
word_count(text) | Aggregate | VARCHAR → BIGINT |
first_word(text) | Scalar | VARCHAR → VARCHAR |
generate_series_ext(n) | Table | BIGINT → TABLE(value BIGINT) |
CAST(VARCHAR AS INTEGER) | Cast | VARCHAR → INTEGER |
Full source: examples/hello-ext/src/lib.rs
Build and try it
cargo build --release --manifest-path examples/hello-ext/Cargo.toml
# DuckDB only loads files ending in `.duckdb_extension` that carry its
# 512-byte metadata footer; a bare `.so` is refused. Append the footer:
cargo run --bin append_metadata -- \
examples/hello-ext/target/release/libhello_ext.so \
hello_ext.duckdb_extension \
--abi-type C_STRUCT --extension-version v0.1.0 \
--duckdb-version v1.2.0 --platform linux_amd64
# An unsigned local build needs -unsigned (it must be a startup flag).
duckdb -unsigned
Then in the DuckDB CLI:
LOAD './hello_ext.duckdb_extension';
-- Aggregate: total words across all rows
SELECT word_count(sentence) FROM (
VALUES ('hello world'), ('one two three'), (NULL)
) t(sentence);
-- → 5 (2 + 3; NULL contributes 0)
-- Scalar: first word of each row
SELECT first_word(sentence) FROM (
VALUES ('hello world'), (' padded '), (''), (NULL)
) t(sentence);
-- → 'hello', 'padded', '', NULL
Overview
An extension has four parts:
- State struct — holds data accumulated during aggregation (aggregate only)
- Callbacks —
update,combine,finalize(aggregate;ffi_state::<T>()suppliesstate_size,state_initandstate_destroy) or a single function callback (scalar) - Registration — wire callbacks to DuckDB via
AggregateFunctionBuilder/ScalarFunctionBuilder - Entry point — DuckDB's initialization hook, generated by
entry_point!
Part 1 — Aggregate function: word_count
An aggregate function accumulates state across many rows and emits one result per group.
1a. The state struct
#![allow(unused)] fn main() { use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64, } impl AggregateState for WordCountState {} }
AggregateState is a marker trait with no methods; its supertraits require the state to be
Default + Send + Sync + 'static.
FfiState<WordCountState> stores it in the bytes DuckDB allocates for each group
(boxing it only when it is large or over-aligned) and manages its lifecycle
(size, init, destroy).
1b. state_size, state_init and state_destroy
These three callbacks are always identical boilerplate, so you do not write them:
.ffi_state::<WordCountState>() on the builder (see Part 3)
installs FfiState::<WordCountState>::size_callback, init_callback and
destroy_callback together. Installing all three for one type in a single call means
DuckDB never allocates space sized for one state and has it initialised as another.
size_callback returns FfiState::<WordCountState>::size() — one usize tag word followed
by the state. A T aligned no more strictly than usize and at most 256 bytes (like
WordCountState) is stored inline in the bytes DuckDB allocates per group; a larger or more
strictly aligned T is boxed, and the slot holds the Box pointer. init_callback writes
WordCountState::default() into the slot (or boxes it), then sets the tag — a value salted per
state type, so a destructor only drops states that were initialised for its own T.
destroy_callback drops the T in each state (and its box, if the state is too
large to be stored inline), clearing the state's tag first so a second call is a
no-op.
1c. update — accumulate one batch
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 } unsafe extern "C" fn wc_update( _info: duckdb_function_info, input: duckdb_data_chunk, states: *mut duckdb_aggregate_state, ) { let reader = unsafe { VectorReader::new(input, 0) }; let row_count = reader.row_count(); for row in 0..row_count { if !unsafe { reader.is_valid(row) } { continue; // NULL input → skip (contributes 0 words) } let s = unsafe { reader.read_str(row) }; let words = count_words(s); let state_ptr = unsafe { *states.add(row) }; if let Some(st) = unsafe { FfiState::<WordCountState>::with_state_mut(state_ptr) } { st.count += words; } } } }
Key points:
- Check
is_valid(row)before reading — never dereference an invalid (NULL) row VectorReader::new(chunk, col)gives columncolfrom the chunkcount_wordsis pure Rust — no unsafe, easy to unit-test separately
1d. combine — merge parallel results
Pitfall L1: DuckDB creates fresh target states before calling
combine, set up bystate_init(hereWordCountState::default()), not copies of the source. You must copy all fields — not just the result field. In an aggregate with config fields (e.g., a histogram with abin_width) you must also copy those, or results will be silently corrupted.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} unsafe extern "C" fn wc_combine( _info: duckdb_function_info, source: *mut duckdb_aggregate_state, target: *mut duckdb_aggregate_state, count: idx_t, ) { for i in 0..count as usize { let src_ptr = unsafe { *source.add(i) }; let tgt_ptr = unsafe { *target.add(i) }; let src = unsafe { FfiState::<WordCountState>::with_state(src_ptr) }; let tgt = unsafe { FfiState::<WordCountState>::with_state_mut(tgt_ptr) }; if let (Some(s), Some(t)) = (src, tgt) { t.count += s.count; // If you add fields to WordCountState, combine them here too. } } } }
1e. finalize — write output
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} unsafe extern "C" fn wc_finalize( _info: duckdb_function_info, source: *mut duckdb_aggregate_state, result: duckdb_vector, count: idx_t, offset: idx_t, ) { let mut writer = unsafe { VectorWriter::new(result) }; for i in 0..count as usize { let state_ptr = unsafe { *source.add(i) }; match unsafe { FfiState::<WordCountState>::with_state(state_ptr) } { Some(st) => unsafe { writer.write_i64(offset as usize + i, st.count) }, None => unsafe { writer.set_null(offset as usize + i) }, } } } }
offset is DuckDB's output row offset — always use offset as usize + i, not just i.
Part 2 — Scalar function: first_word
A scalar function processes one data chunk and returns one output value per row. The callback receives the full chunk and an output vector (not per-row state pointers).
Key rule: always propagate NULL
If the input row is NULL, write NULL to output — never read from an invalid row.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } unsafe extern "C" fn first_word_scalar( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; let row_count = reader.row_count(); for row in 0..row_count { if !unsafe { reader.is_valid(row) } { unsafe { writer.set_null(row) }; // NULL in → NULL out continue; } let s = unsafe { reader.read_str(row) }; unsafe { writer.write_varchar(row, first_word(s)) }; } } }
The pure logic:
#![allow(unused)] fn main() { pub fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } }
Note: set_null internally calls duckdb_vector_ensure_validity_writable before writing
the null flag — this is required by DuckDB and handled for you by VectorWriter.
Part 3 — Registration
The snippets below register through the raw builders and entry_point!, which take a
duckdb_connection. hello-ext itself uses the equivalent entry_point_v2!, whose closure
receives a &Connection and registers each builder through the Registrar trait
(con.register_aggregate(...), con.register_scalar(...)); see
The Entry Point.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 } fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } unsafe extern "C" fn wc_update(_info: duckdb_function_info, input: duckdb_data_chunk, states: *mut duckdb_aggregate_state) { let reader = unsafe { VectorReader::new(input, 0) }; for row in 0..reader.row_count() { if !unsafe { reader.is_valid(row) } { continue; } let words = count_words(unsafe { reader.read_str(row) }); if let Some(st) = unsafe { FfiState::<WordCountState>::with_state_mut(*states.add(row)) } { st.count += words; } } } unsafe extern "C" fn wc_combine(_info: duckdb_function_info, source: *mut duckdb_aggregate_state, target: *mut duckdb_aggregate_state, count: idx_t) { for i in 0..count as usize { let src = unsafe { FfiState::<WordCountState>::with_state(*source.add(i)) }; let tgt = unsafe { FfiState::<WordCountState>::with_state_mut(*target.add(i)) }; if let (Some(s), Some(t)) = (src, tgt) { t.count += s.count; } } } unsafe extern "C" fn wc_finalize(_info: duckdb_function_info, source: *mut duckdb_aggregate_state, result: duckdb_vector, count: idx_t, offset: idx_t) { let mut writer = unsafe { VectorWriter::new(result) }; for i in 0..count as usize { match unsafe { FfiState::<WordCountState>::with_state(*source.add(i)) } { Some(st) => unsafe { writer.write_i64(offset as usize + i, st.count) }, None => unsafe { writer.set_null(offset as usize + i) }, } } } unsafe extern "C" fn first_word_scalar(_info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector) { let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..reader.row_count() { if !unsafe { reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } unsafe { writer.write_varchar(row, first_word(reader.read_str(row))) }; } } fn live_connection() -> libduckdb_sys::duckdb_connection { std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess); } con } /// First column of the first row, as BIGINT; `None` for NULL. fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); let reader = unsafe { chunk.reader(0) }; unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) } } unsafe fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> { unsafe { AggregateFunctionBuilder::new("word_count") .param(TypeId::Varchar) .returns(TypeId::BigInt) .ffi_state::<WordCountState>() // state_size + init + destructor .update(wc_update) .combine(wc_combine) .finalize(wc_finalize) .register(con)?; ScalarFunctionBuilder::new("first_word") .param(TypeId::Varchar) .returns(TypeId::Varchar) .function(first_word_scalar) .register(con)?; } Ok(()) } let con = live_connection(); unsafe { register(con) }.unwrap(); assert_eq!(query_i64(con, "SELECT word_count(s) FROM (VALUES ('hello world'), (NULL), ('one two three')) t(s)"), Some(5)); assert_eq!(query_i64(con, "SELECT length(first_word(' quack rs'))::BIGINT"), Some(5)); assert_eq!(query_i64(con, "SELECT count(*) FROM (SELECT first_word(NULL) AS w) WHERE w IS NULL"), Some(1)); // Many groups, so DuckDB runs combine on partial states. assert_eq!(query_i64(con, "SELECT sum(c)::BIGINT FROM (SELECT word_count('a b') AS c FROM range(100000) GROUP BY range % 997)"), Some(200000)); }
Both builders call the DuckDB C API internally. register returns Err if DuckDB reports
a failure — this propagates to the entry point and is surfaced to the user.
Part 4 — Entry point
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe fn register(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } quack_rs::entry_point!(hello_ext_init_c_api, |con| unsafe { register(con) }); }
This one line expands to the equivalent of:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_connection, duckdb_extension_access, duckdb_extension_info}; use quack_rs::error::ExtensionError; unsafe fn register(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } #[no_mangle] pub unsafe extern "C" fn hello_ext_init_c_api( info: duckdb_extension_info, access: *const duckdb_extension_access, ) -> bool { unsafe { quack_rs::entry_point::init_extension_with_policy( info, access, quack_rs::DUCKDB_API_VERSION, quack_rs::abi::AbiPolicy::Strict, |con| unsafe { register(con) }, ) } } }
Pass the full symbol name — hello_ext_init_c_api here. DuckDB looks up this exact
symbol when loading the extension. See The Entry Point for
the full initialization sequence.
Unit tests (no DuckDB process needed)
Test pure logic directly:
#![allow(unused)] fn main() { fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 } fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } #[test] fn count_words_whitespace_variants() { assert_eq!(count_words(" hello world "), 2); assert_eq!(count_words("\t\nhello\tworld\n"), 2); assert_eq!(count_words(" "), 0); // all whitespace → 0 } #[test] fn first_word_empty_and_whitespace() { assert_eq!(first_word(""), ""); assert_eq!(first_word(" "), ""); } }
Test aggregate state with AggregateTestHarness:
#![allow(unused)] fn main() { use quack_rs::prelude::*; use quack_rs::testing::AggregateTestHarness; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 } #[test] fn word_count_null_rows_are_skipped() { // DuckDB passes NULL rows to `update`; the callback's `is_valid` check // skips them, so they never reach the state. let mut h = AggregateTestHarness::<WordCountState>::new(); h.update(|s| s.count += count_words("hello")); // NULL row omitted — models the callback's `is_valid` skip h.update(|s| s.count += count_words("world")); assert_eq!(h.finalize().count, 2); } #[test] fn word_count_combine() { let mut h1 = AggregateTestHarness::<WordCountState>::new(); h1.update(|s| s.count += count_words("hello world")); // 2 let mut h2 = AggregateTestHarness::<WordCountState>::new(); h2.update(|s| s.count += count_words("one two three four")); // 4 h2.combine(&h1, |src, tgt| tgt.count += src.count); assert_eq!(h2.finalize().count, 6); } }
Run all tests with:
cargo test --manifest-path examples/hello-ext/Cargo.toml
See the Testing Guide for the full test strategy.
Project Scaffold
quack_rs::scaffold::generate_scaffold generates every file of a new DuckDB community
extension project in Rust — Cargo.toml, Makefile, CI workflow, description.yml, an
example function and its tests — from a single function call.
What it generates
my_extension/
├── Cargo.toml # cdylib crate, dependencies, release profile
├── Makefile # delegates to cargo + extension-ci-tools
├── extension_config.cmake # required by extension-ci-tools
├── src/
│ ├── lib.rs # entry point, example function, unit test
│ └── wasm_lib.rs # WASM staticlib shim
├── description.yml # community extension metadata
├── test/
│ └── sql/
│ └── my_extension.test # SQLLogicTest for the example function
├── .github/
│ └── workflows/
│ └── extension-ci.yml # cross-platform CI workflow
├── .gitmodules # extension-ci-tools submodule
├── .gitignore
└── .cargo/
└── config.toml # Windows CRT static linking
Usage
generate_scaffold returns paths relative to the project root (Cargo.toml,
src/lib.rs, ...) and writes nothing itself. Join them under the directory the
project should live in — writing them relative to the current directory would
overwrite whatever Cargo.toml and src/lib.rs are already there.
use quack_rs::scaffold::{ScaffoldConfig, generate_scaffold}; use std::path::Path; fn main() { let config = ScaffoldConfig { name: "my_extension".to_string(), description: "My DuckDB extension".to_string(), version: "0.1.0".to_string(), license: "MIT".to_string(), maintainer: "Your Name".to_string(), github_repo: "yourorg/duckdb-my-extension".to_string(), excluded_platforms: vec![], // `target_duckdb_version`, `use_unstable_c_api` and `git_ref` default // to the stable-ABI settings; see `concepts/abi.md`. ..ScaffoldConfig::default() }; let files = generate_scaffold(&config).expect("scaffold generation failed"); // Everything goes under ./my_extension/, never into the current directory. let root = Path::new(&config.name); for file in &files { let path = root.join(&file.path); if let Some(parent) = path.parent() { std::fs::create_dir_all(parent).unwrap(); } std::fs::write(&path, &file.content).unwrap(); println!("created {}", path.display()); } }
ScaffoldConfig fields
| Field | Type | Description |
|---|---|---|
name | String | Extension name — must match [lib] name in Cargo.toml and description.yml |
description | String | Description for description.yml and the //! docs of src/lib.rs. Quoted and escaped, so :, # and quotes are fine; must not be empty, padded with whitespace, or contain control characters other than newline and tab |
version | String | The extension's version — validated by validate_extension_version |
license | String | SPDX license identifier (e.g., "MIT", "Apache-2.0") |
maintainer | String | Your name or org, listed in description.yml — one non-empty line |
github_repo | String | "owner/repo", in the characters GitHub allows |
excluded_platforms | Vec<String> | Platforms to skip (e.g., ["wasm_mvp", "wasm_eh"]) |
git_ref | String | repo.ref — a commit hash (or tag), not a branch. Defaults to REF_PLACEHOLDER so it cannot be submitted unset |
target_duckdb_version | String | Written as TARGET_DUCKDB_VERSION in the Makefile. Defaults to DUCKDB_API_VERSION (v1.2.0); with use_unstable_c_api, an exact DuckDB release such as v1.5.5 |
use_unstable_c_api | bool | Set when the extension enables duckdb-1-5 / duckdb-1-5-3 / duckdb-1-5-4. Defaults to false |
Name validation
Extension names must satisfy all of:
- Match
^[a-z][a-z0-9_]*$(no hyphens: the entry point is<name>_init_c_api) - Not exceed 64 characters
- Be globally unique on community-extensions.duckdb.org
Use vendor-prefixed names to avoid collisions: myorg_analytics, not analytics.
The scaffold generator checks the first two rules before generating any files and returns an error if the name breaks one. Uniqueness is yours to check.
After scaffolding
cd my_extension
git init
git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools
make configure
make release
make test
git submodule add is not optional in a new project: the scaffold writes
.gitmodules, but git submodule update --init does nothing until the
submodule has been added once (Pitfall P4). The generated Makefile stops with
this command if the checkout is missing.
cargo test runs the generated unit test and make test runs
test/sql/my_extension.test against a real DuckDB. Replace the example
function in src/lib.rs with your own, give each new function a test in both
places, and push to GitHub — CI runs automatically.
Excluded platforms
Some extensions cannot be built for all platforms (e.g., extensions that depend on platform-specific system libraries, or WASM environments that lack threading).
#![allow(unused)] fn main() { use quack_rs::scaffold::ScaffoldConfig; let config = ScaffoldConfig { excluded_platforms: vec![ "wasm_mvp".to_string(), "wasm_eh".to_string(), "wasm_threads".to_string(), ], ..ScaffoldConfig::default() }; }
Validate individual platform names with quack_rs::validate::validate_platform, or a
semicolon-delimited string (as used in description.yml) with
quack_rs::validate::validate_excluded_platforms_str.
Extension Anatomy
A DuckDB loadable extension is a shared library (.so / .dylib / .dll) that DuckDB loads
at runtime. This page covers what DuckDB expects of a Rust extension built on the C Extension
API: the entry-point symbol, the initialization sequence, the loadable-extension dispatch
table, and version and binary compatibility.
The initialization sequence
When DuckDB loads your extension, it:
- Opens the shared library and looks up the symbol
{name}_init_c_api, where{name}is the extension name - Calls that function with an
infohandle and a pointer to aduckdb_extension_accessstruct (theset_error,get_databaseandget_apicallbacks) - Your function must:
- Call
duckdb_rs_extension_api_init(info, access, api_version)to initialize the dispatch table - Get the
duckdb_databasehandle viaaccess.get_database(info) - Open a
duckdb_connectionviaduckdb_connect - Register functions on that connection
- Disconnect
- Return
true(success) orfalse(failure), reporting any error throughaccess.set_error
- Call
quack_rs::entry_point::init_extension performs this sequence. It also checks the C API
layout before registration (see ABI Compatibility) and converts a panic in your
registration code into a load error. The entry_point! macro generates the required
#[no_mangle] extern "C" symbol:
#![allow(unused)] fn main() { use quack_rs::entry_point; use quack_rs::error::ExtensionError; fn register(_con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } entry_point!(my_extension_init_c_api, |con| register(con)); // emits: #[no_mangle] pub unsafe extern "C" fn my_extension_init_c_api(...) }
Symbol naming
The symbol name must be {extension_name}_init_c_api. Extension names use lowercase ASCII
letters, digits and underscores only (no hyphens, which cannot appear in a C symbol).
If the symbol is missing or misnamed, DuckDB fails to load the extension.
Extension name: "word_count_ext"
Required symbol: word_count_ext_init_c_api
Pass the full symbol name to entry_point!. The exported name then appears verbatim at the
call site; the macro does not build identifiers at compile time.
The loadable-extension feature
With features = ["loadable-extension"], libduckdb-sys does not link DuckDB; every C API
function dispatches through a table of function pointers instead:
Without feature: duckdb_query(...) → calls linked libduckdb directly
With feature: duckdb_query(...) → dispatches through an AtomicPtr table
The AtomicPtr table starts as null. The extension's entry point fills it by calling
duckdb_rs_extension_api_init, which copies the pointers out of the API struct DuckDB
provides. This means:
- Any call before
duckdb_rs_extension_api_initpanics with"DuckDB API not initialized or DuckDB feature omitted" - In a plain
cargo test, you cannot call anyduckdb_*function — no DuckDB host process ever initializes the table
This is why quack-rs offers AggregateTestHarness for testing: it simulates the aggregate
lifecycle in pure Rust, without calling the DuckDB API. For SQL-level tests, the
bundled-test / bundled-test-prebuilt features link a real DuckDB and fill the table
when InMemoryDb::open() runs (see Testing Guide).
Dependency model
graph TD
EXT["your-extension"]
QR["quack-rs"]
LDS["libduckdb-sys >=1.4.4, <2<br/>{loadable-extension}<br/>(headers only — no linked library)"]
EXT --> QR
EXT --> LDS
QR --> LDS
The loadable-extension feature produces a shared library that does not statically link
DuckDB. Instead, it receives DuckDB's function pointers at load time, so the extension runs
inside the host DuckDB process and uses that process's DuckDB instance, memory and threads.
Version support
libduckdb-sys = ">=1.4.4, <2" — the bounded range is intentional.
Every DuckDB 1.4.x and 1.5.x release loads extensions built for C API version v1.2.0
(the version string passed to duckdb_rs_extension_api_init, quack_rs::DUCKDB_API_VERSION).
DuckDB 1.5.6 declares C API version v1.5.6 and still accepts v1.2.0 extensions. CI loads
the example extension into DuckDB 1.4.4, 1.5.0, 1.5.5 and the latest release.
Using a range rather than an exact pin means:
- Extension authors can choose which
libduckdb-sysrelease to build against (for example=1.4.4for DuckDB 1.4.4, or~1.10505.0for DuckDB 1.5.5, whichlibduckdb-sysnumbers1.10505.x) and still resolve againstquack-rs quack-rsitself doesn't force a DuckDB downgrade on users
The <2 upper bound is equally intentional: it prevents silent adoption of a future major
release that may introduce breaking C API changes. Upgrading beyond the 1.x band requires
an explicit quack-rs release that audits the new C API surface.
For your own extension's
Cargo.toml: an extension that uses only the stable C API (the default quack-rs features) can keep the same">=1.4.4, <2"range; this is whatscaffold::generate_scaffoldwrites. An extension that enables theduckdb-1-5*features must pinlibduckdb-systo the bindings of the one DuckDB release it is stamped for (the scaffold writes~1.10505.0forv1.5.5). See ABI Compatibility.
Binary compatibility
Which DuckDB releases accept an extension binary depends on the ABI type stamped into its metadata footer:
- A
C_STRUCTbinary targeting C APIv1.2.0(the default, stable API only) loads into every DuckDB release whose C API version is at leastv1.2.0— all of 1.4.x and 1.5.x — on the platform it was built for - A
C_STRUCT_UNSTABLEbinary loads only into the exact DuckDB release it names. Stamp builds that use theduckdb-1-5*features this way, so DuckDB itself refuses a mismatched release; quack-rs's runtime layout check is the backstop when the stamp is missing - DuckDB checks the footer's platform and version fields at load time and refuses a mismatch
- Core and community extensions are signed; a binary you build locally is not
- To load an unsigned extension during development, start DuckDB with
allow_unsigned_extensionsenabled (duckdb -unsignedin the CLI); the setting cannot be changed on a running database - The community extension CI builds and signs each extension for every supported platform
The Entry Point
Every DuckDB loadable extension exports one C-callable entry-point function, which DuckDB
calls when it loads the extension. quack-rs generates it with the entry_point_v2! or
entry_point! macro, or you can write it by hand around init_extension.
Option A: entry_point_v2! with Connection (recommended)
Added in v0.4.0.
The entry_point_v2! macro gives your closure a &Connection instead of a raw
duckdb_connection. The Connection type implements the Registrar trait, which
registers every kind of function, macro, cast, copy function and config option;
replacement scans, which belong to the database rather than a connection, are
Connection's own methods:
use quack_rs::entry_point_v2;
use quack_rs::connection::{Connection, Registrar};
use quack_rs::error::ExtensionError;
unsafe fn register(con: &Connection) -> Result<(), ExtensionError> {
unsafe {
con.register_scalar(/* ScalarFunctionBuilder */)?;
con.register_aggregate(/* AggregateFunctionBuilder */)?;
con.register_table(/* TableFunctionBuilder */)?;
con.register_cast(/* CastFunctionBuilder */)?;
con.register_scalar_set(/* ScalarFunctionSetBuilder */)?;
con.register_aggregate_set(/* AggregateFunctionSetBuilder */)?;
con.register_sql_macro(/* SqlMacro */)?;
con.register_replacement_scan(/* callback, data, destructor */);
// con.register_copy_function(/* CopyFunctionBuilder */)?; // requires duckdb-1-5
}
Ok(())
}
entry_point_v2!(my_extension_init_c_api, |con| unsafe { register(con) });
(The register_* arguments above are placeholders, so that block is not
compilable as written.) The macro emits the equivalent of the following; the real
expansion also evaluates the policy and closure arguments inside a panic guard:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_extension_access, duckdb_extension_info}; use quack_rs::connection::Connection; use quack_rs::error::ExtensionError; unsafe fn register(_con: &Connection) -> Result<(), ExtensionError> { Ok(()) } #[no_mangle] pub unsafe extern "C" fn my_extension_init_c_api( info: duckdb_extension_info, access: *const duckdb_extension_access, ) -> bool { unsafe { quack_rs::entry_point::init_extension_v2_with_policy( info, access, quack_rs::DUCKDB_API_VERSION, quack_rs::abi::AbiPolicy::Strict, |con| unsafe { register(con) }, ) } } }
entry_point_v2!(name, policy, |con| ...) passes a different AbiPolicy.
Pass the full symbol name to the macro. The symbol {name}_init_c_api must match the
name field in description.yml and the [lib] name in Cargo.toml.
Why Connection over raw duckdb_connection?
entry_point! (raw) | entry_point_v2! (Connection) | |
|---|---|---|
| Receives | duckdb_connection | &Connection |
| Registration | Call builders' .register(con) | Call con.register_*() |
| Type safety | Raw pointer | Typed wrapper around the connection and database handles |
| Replacement scans | Need the database handle, which the closure does not receive | con.register_replacement_scan*() |
Option B: The entry_point! macro
The original macro passes a raw duckdb_connection to your closure. It performs the
same initialization, but you pass the connection to each builder's .register():
#![allow(unused)] fn main() { use quack_rs::entry_point; use quack_rs::error::ExtensionError; fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> { // Register each function on `con`, e.g. `unsafe { builder.register(con)? };` let _ = con; Ok(()) } entry_point!(my_extension_init_c_api, |con| register(con)); }
Option C: Manual entry point
If you need full control (e.g., multiple registration functions, conditional logic):
#![allow(unused)] fn main() { use quack_rs::entry_point::init_extension; use libduckdb_sys::{duckdb_extension_info, duckdb_extension_access}; use libduckdb_sys::duckdb_connection; use quack_rs::error::ExtensionError; fn register_scalar_functions(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } fn register_aggregate_functions(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } fn register_sql_macros(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } #[no_mangle] pub unsafe extern "C" fn my_extension_init_c_api( info: duckdb_extension_info, access: *const duckdb_extension_access, ) -> bool { unsafe { init_extension(info, access, quack_rs::DUCKDB_API_VERSION, |con| { register_scalar_functions(con)?; register_aggregate_functions(con)?; register_sql_macros(con)?; Ok(()) }) } } }
What init_extension does
flowchart TD
A["<b>1. duckdb_rs_extension_api_init</b>(info, access, version)<br/>Fills the global AtomicPtr dispatch table"]
L["<b>2. ABI layout check</b><br/>Applies the AbiPolicy (Strict by default)"]
B["<b>3. access.get_database</b>(info)<br/>Returns the duckdb_database handle"]
C["<b>4. duckdb_connect</b>(db, &mut con)<br/>Opens a connection for function registration"]
D["<b>5. register</b>(con) ← your closure<br/>A panic becomes an error"]
E["<b>6. duckdb_disconnect</b>(&mut con)<br/>Always runs, even if registration failed"]
F{Error?}
G["return <b>true</b>"]
H["return <b>false</b><br/>error reported via access.set_error"]
A --> L --> B --> C --> D --> E --> F
L -->|refused| H
F -->|no| G
F -->|yes| H
style G fill:#1c3b1c,stroke:#4a9e4a,color:#c8ecc8
style H fill:#3b1c1c,stroke:#9e4a4a,color:#ecc8c8
An error from any step, including an Err or a panic from your closure in step 5, is
reported to DuckDB via access.set_error, and the function returns false. DuckDB then
fails the LOAD with that message. (When DuckDB itself detects the failure, as when
get_database returns null, quack-rs returns false without overwriting DuckDB's own
message.) The layout check in step 2 only runs when a duckdb-1-5*
feature is enabled; see ABI Compatibility.
Registration is not transactional: functions registered before a failure stay registered
for the life of the database. Do fallible setup work (reading configuration, building
lookup tables) before the first register call.
The C API version constant
#![allow(unused)] fn main() { pub const DUCKDB_API_VERSION: &str = "v1.2.0"; }
Pitfall P2: This is the C API version, not the DuckDB release version. DuckDB 1.4.x and 1.5.0–1.5.5 declare C API version
v1.2.0; 1.5.6 declaresv1.5.6and still loads extensions that targetv1.2.0(DuckDB accepts any C API version up to its own). A binary stamped with a DuckDB release version instead (-dv v1.5.5) is refused atLOADby every DuckDB whose C API version is lower; quack-rs'sappend_metadatarejects such a value forC_STRUCTup front. See Pitfall P2.
No panics in the entry point
init_extension never panics. All error paths use Result and ?. If your registration
closure returns Err or panics, the message is reported to DuckDB via access.set_error
and the LOAD fails with an error instead of aborting the process. Catching a panic
requires panic = "unwind" in your release profile (the scaffold's default).
Never use unwrap() or expect() in FFI callbacks.
See Pitfall L3.
Error Handling
quack-rs reports every error through one type, ExtensionError, and the ExtResult<T>
alias. This page covers creating and propagating errors, how they reach DuckDB, and why
extension code must not panic.
ExtensionError
#![allow(unused)] fn main() { use quack_rs::error::{ExtensionError, ExtResult}; let (name, code) = ("my_fn", 1); let some_std_error = std::fmt::Error; // From a string literal let e = ExtensionError::from("something went wrong"); // From a format string let e = ExtensionError::new(format!("failed to register '{}': code {}", name, code)); // Wrapping another error let e = ExtensionError::from_error(some_std_error); }
ExtensionError implements:
std::error::ErrorDisplay,Debug,Clone,PartialEq,EqFrom<&str>,From<String>,From<Box<dyn Error>>,From<Box<dyn Error + Send + Sync>>From<std::io::Error>,From<std::ffi::NulError>,From<std::fmt::Error>
The From<std::io::Error> impl lets extensions that allocate runtime resources
during initialization (a tokio runtime, for example) use ? without .map_err():
#![allow(unused)] fn main() { use quack_rs::connection::Connection; use quack_rs::error::ExtensionError; // A stand-in with tokio's signature: `Runtime::new() -> std::io::Result<Runtime>`. mod tokio { pub mod runtime { pub struct Runtime; impl Runtime { pub fn new() -> std::io::Result<Self> { Ok(Runtime) } } } } fn register_all(con: &Connection) -> Result<(), ExtensionError> { let _rt = tokio::runtime::Runtime::new()?; // ← io::Error → ExtensionError // ... register functions ... Ok(()) } }
ExtResult<T>
A type alias for Result<T, ExtensionError>, used throughout the SDK:
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; pub type ExtResult<T> = Result<T, ExtensionError>; }
Propagating errors with ?
In your registration function:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::prelude::*; unsafe extern "C" fn my_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { ScalarFunctionBuilder::try_new("my_fn")? .param(TypeId::BigInt) .returns(TypeId::BigInt) .function(my_fn) .register(con)?; // ← ? propagates registration errors SqlMacro::scalar("my_macro", &["x"], "x + 1")? .register(con)?; Ok(()) } } }
If any registration call fails, ? returns the error from register, which
init_extension then reports to DuckDB via access.set_error.
Error reporting to DuckDB
init_extension converts the ExtensionError to a C string for DuckDB's set_error
callback, following the same rule as ExtensionError::to_c_string: a C string cannot
hold a NUL byte, so each one is replaced with ? and the rest of the message is kept:
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; let err = ExtensionError::new("bad\0input"); assert_eq!(err.to_c_string().to_str(), Ok("bad?input")); }
DuckDB surfaces this string to the user as the extension load error.
No panics
The central rule of DuckDB extension development:
Never
unwrap(),expect(), orpanic!()in any code path that DuckDB may call.
A panic cannot unwind out of an extern "C" function: since Rust 1.81 the runtime aborts
the process instead, taking the user's DuckDB session with it. quack-rs's callback macros
and typed builders catch panics and turn them into SQL errors (which requires
panic = "unwind" in the release profile), but that is a safety net for bugs, not an
error-handling strategy.
Safe patterns
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_aggregate_state; use quack_rs::aggregate::{AggregateState, FfiState}; use quack_rs::error::ExtensionError; #[derive(Default)] struct MyState { count: u64 } impl AggregateState for MyState {} fn some_fallible_call() -> Result<u64, ExtensionError> { Ok(1) } unsafe fn demo(state_ptr: duckdb_aggregate_state, maybe_count: Option<u64>) -> Result<(), ExtensionError> { // ✅ Use Option methods if let Some(s) = FfiState::<MyState>::with_state_mut(state_ptr) { s.count += 1; } // ✅ Use Result and ? let value = some_fallible_call()?; // ✅ Use unwrap_or / unwrap_or_else / map let count = maybe_count.unwrap_or(0); // ❌ Never in FFI callbacks let s = FfiState::<MyState>::with_state_mut(state_ptr).unwrap(); // panics if None Ok(()) } }
In init_extension
init_extension reports every error via set_error and does not panic itself. It also
runs your registration closure under catch_unwind, so a panic there becomes a load error
rather than a process abort (with panic = "unwind").
Type System
quack-rs describes DuckDB column types with two types: TypeId, a plain enum of DuckDB's
type ids, and LogicalType, an owned handle for full type descriptions such as
DECIMAL(18, 3), LIST(VARCHAR) or a STRUCT. This page also maps each DuckDB type to its
Rust type and vector read/write methods.
TypeId
TypeId is an enum covering DuckDB's column types (the GEOMETRY and
VARIANT types added in DuckDB 1.5.x are exposed behind the duckdb-1-5-3
feature — see Known Limitations). The list below
names the variants; it is a listing, not compilable code:
use quack_rs::types::TypeId;
TypeId::Boolean
TypeId::TinyInt // i8
TypeId::SmallInt // i16
TypeId::Integer // i32
TypeId::BigInt // i64
TypeId::UTinyInt // u8
TypeId::USmallInt // u16
TypeId::UInteger // u32
TypeId::UBigInt // u64
TypeId::HugeInt // i128
TypeId::UHugeInt // u128
TypeId::Float // f32
TypeId::Double // f64
TypeId::Timestamp
TypeId::TimestampTz
TypeId::TimestampS
TypeId::TimestampMs
TypeId::TimestampNs
TypeId::Date
TypeId::Time
TypeId::TimeTz
TypeId::Interval
TypeId::Varchar
TypeId::Blob
TypeId::Decimal
TypeId::Enum
TypeId::List
TypeId::Struct
TypeId::Map
TypeId::Uuid
TypeId::Union
TypeId::Bit
TypeId::Array
TypeId::TimeNs
TypeId::Any
TypeId::Varint // SQL name BIGNUM (VARINT before DuckDB 1.4)
TypeId::SqlNull
TypeId::IntegerLiteral
TypeId::StringLiteral
TypeId::Geometry // duckdb-1-5-3
TypeId::Variant // duckdb-1-5-3
TypeId implements Copy, Clone, Debug, PartialEq, Eq, Hash and Display.
SQL name
#![allow(unused)] fn main() { use quack_rs::types::TypeId; assert_eq!(TypeId::BigInt.sql_name(), "BIGINT"); assert_eq!(TypeId::Varchar.sql_name(), "VARCHAR"); assert_eq!(format!("{}", TypeId::Timestamp), "TIMESTAMP"); }
DuckDB constant
TypeId::to_duckdb_type() returns the DUCKDB_TYPE_* integer constant from libduckdb-sys.
You rarely need this directly — it's called internally by LogicalType::new.
Reverse conversion
TypeId::from_duckdb_type(raw) converts a raw DUCKDB_TYPE constant back into a TypeId.
It panics if the value does not match any known constant; TypeId::try_from_duckdb_type(raw)
returns None instead.
#![allow(unused)] fn main() { use quack_rs::types::TypeId; let type_id = TypeId::from_duckdb_type(libduckdb_sys::DUCKDB_TYPE_DUCKDB_TYPE_BIGINT); assert_eq!(type_id, TypeId::BigInt); }
LogicalType
LogicalType is an RAII wrapper around DuckDB's duckdb_logical_type. Creating one calls
into DuckDB, so this block is compiled but not run:
#![allow(unused)] fn main() { use quack_rs::types::{LogicalType, TypeId}; let lt = LogicalType::new(TypeId::Varchar); // lt.as_raw() returns the duckdb_logical_type pointer // Drop calls duckdb_destroy_logical_type automatically }
Pitfall L7:
duckdb_create_logical_typeallocates memory that must be freed withduckdb_destroy_logical_type.LogicalType'sDropimplementation does this automatically, preventing the memory leak that occurs when calling the DuckDB C API directly. See Pitfall L7.
For a type a TypeId fully describes, pass the TypeId to a builder's param or returns
method; the builder creates and destroys the LogicalType internally. For a parameterized
type (DECIMAL, LIST, MAP, STRUCT, UNION, ENUM, ARRAY), build a LogicalType and
pass it to param_logical or returns_logical.
Constructors
| Constructor | Creates |
|---|---|
LogicalType::new(type_id) | Simple type from a TypeId |
LogicalType::from_raw(ptr) | Takes ownership of a raw duckdb_logical_type handle (unsafe) |
LogicalType::decimal(width, scale) | DECIMAL(width, scale) |
LogicalType::list(element_type) | LIST<element_type> from a TypeId |
LogicalType::list_from_logical(element) | LIST<element> from an existing LogicalType |
LogicalType::map(key, value) | MAP<key, value> from TypeIds |
LogicalType::map_from_logical(key, value) | MAP<key, value> from existing LogicalTypes |
LogicalType::struct_type(fields) | STRUCT from &[(&str, TypeId)] |
LogicalType::struct_type_from_logical(fields) | STRUCT from &[(&str, LogicalType)] |
LogicalType::union_type(members) | UNION from &[(&str, TypeId)] |
LogicalType::union_type_from_logical(members) | UNION from &[(&str, LogicalType)] |
LogicalType::enum_type(members) | ENUM from &[&str] |
LogicalType::array(element_type, size) | ARRAY<element_type>[size] from a TypeId |
LogicalType::array_from_logical(element, size) | ARRAY<element>[size] from an existing LogicalType |
Every constructor except from_raw panics on invalid input (for example a composite
TypeId passed to new, or a DECIMAL width outside 1–38) and has a try_* counterpart
(try_new, try_decimal, try_struct_type, …) that returns
Result<LogicalType, LogicalTypeError> instead. UNION types accept at most
MAX_UNION_MEMBERS (255) members.
Introspection methods
All introspection methods are unsafe: they call into DuckDB, so the handle must be valid
and the C API initialized.
| Method | Returns | Applicable to |
|---|---|---|
get_type_id() | TypeId | Any |
get_alias() | Option<String> | Any |
set_alias(alias) | () | Any |
decimal_width() | u8 | DECIMAL |
decimal_scale() | u8 | DECIMAL |
decimal_internal_type() | TypeId | DECIMAL |
enum_internal_type() | TypeId | ENUM |
enum_dictionary_size() | u32 | ENUM |
enum_dictionary_value(index) | String | ENUM |
list_child_type() | LogicalType | LIST |
map_key_type() | LogicalType | MAP |
map_value_type() | LogicalType | MAP |
struct_child_count() | u64 | STRUCT |
struct_child_name(index) | String | STRUCT |
struct_child_type(index) | LogicalType | STRUCT |
union_member_count() | u64 | UNION |
union_member_name(index) | String | UNION |
union_member_type(index) | LogicalType | UNION |
array_size() | u64 | ARRAY |
array_child_type() | LogicalType | ARRAY |
Rust type ↔ DuckDB type mapping
When reading from or writing to vectors, use the corresponding VectorReader/VectorWriter
method. The most common types:
| DuckDB type | TypeId | Reader method | Writer method |
|---|---|---|---|
BOOLEAN | Boolean | read_bool | write_bool |
TINYINT | TinyInt | read_i8 | write_i8 |
SMALLINT | SmallInt | read_i16 | write_i16 |
INTEGER | Integer | read_i32 | write_i32 |
BIGINT | BigInt | read_i64 | write_i64 |
UTINYINT | UTinyInt | read_u8 | write_u8 |
USMALLINT | USmallInt | read_u16 | write_u16 |
UINTEGER | UInteger | read_u32 | write_u32 |
UBIGINT | UBigInt | read_u64 | write_u64 |
FLOAT | Float | read_f32 | write_f32 |
DOUBLE | Double | read_f64 | write_f64 |
HUGEINT | HugeInt | read_i128 | write_i128 |
UHUGEINT | UHugeInt | read_u128 | write_u128 |
VARCHAR | Varchar | read_str | write_varchar |
BLOB | Blob | read_blob | write_blob |
UUID | Uuid | read_uuid | write_uuid |
DATE | Date | read_date | write_date |
TIME | Time | read_time | write_time |
TIMESTAMP | Timestamp | read_timestamp | write_timestamp |
INTERVAL | Interval | read_interval | write_interval |
The timestamp variants (TIMESTAMP_S, _MS, _NS, TIMESTAMPTZ), TIMETZ and DECIMAL
have matching read_*/write_* methods as well.
NULLs are handled separately — see NULL Handling & Strings.
ABI Compatibility
This page explains DuckDB C Extension API ABI compatibility: which DuckDB releases a
quack-rs extension binary can load into, and how quack-rs guards against a layout
mismatch. A loadable extension does not link against DuckDB's symbols. DuckDB hands it a
pointer to a duckdb_ext_api_v1 struct — an array of function pointers — and the
extension calls through it.
Two things have to agree about that struct's layout: the header your extension
was compiled against (whichever libduckdb-sys version Cargo resolved), and
the DuckDB binary that is loading it. When they disagree, every call lands on
the wrong slot.
The struct has two halves
| Region | Slots | Guarantee |
|---|---|---|
| Stable | 0 .. 357 | Frozen since DuckDB v1.2.0 — same slots, same order, same signatures in every release through v1.5.6 (two slots, 114 and 138, were renamed varint → bignum in v1.4.0 with an identical struct layout) |
| Unstable | 357 .. | DuckDB inserts new entries in the middle, shifting every later slot |
The stable prefix is what makes "build once, load anywhere" possible. The unstable tail is not append-only:
| DuckDB | Total slots | What moved |
|---|---|---|
| v1.2.0 – v1.2.2 | 408 | baseline |
| v1.3.0 – v1.3.2 | 428 | appended |
| v1.4.0 – v1.4.5 | 459 | duckdb_create_varint → duckdb_create_bignum; appended |
| v1.5.0 – v1.5.1 | 545 | duckdb_appender_clear inserted at slot 410 |
| v1.5.2 – v1.5.6 | 546 | duckdb_geometry_type_get_crs inserted at slot 493 |
DuckDB v1.5.6 declares all 546 slots stable for extensions that target C API v1.5.6 (earlier releases declared only the first 357 stable). The layout itself is unchanged from v1.5.2, and quack-rs targets C API v1.2.0, so the guard below still applies.
Every family since v1.2 changed the unstable tail, and twice (v1.5.0 and v1.5.2) an insertion in the middle shifted every later slot.
Which half are you using?
Everything quack-rs exposes by default lives in the stable prefix: scalar,
aggregate, table and cast functions, vectors, data chunks, values, SQL macros,
replacement scans, the query API and the
datetime conversions.
The duckdb-1-5, duckdb-1-5-3 and duckdb-1-5-4 features wrap the unstable
half — 130 of its 189 functions, covering scalar bind/init, copy functions in both
directions, the Arrow C Data Interface bridge, catalog access, ErrorData,
FileSystem, Expression, SelectionVector, config options, table descriptions,
TIME_NS values and the client context.
abi::uses_unstable_api() reports which half a build uses. The book's own test
build enables duckdb-1-5-4 (and therefore duckdb-1-5), so this block is
compiled there but not run:
#![allow(unused)] fn main() { use quack_rs::abi; // false unless `duckdb-1-5` is enabled. assert!(!abi::uses_unstable_api()); }
Why DuckDB does not catch this for you
DuckDB validates the ABI metadata in your extension's footer:
| ABI type | Version field (-dv) means | Accepted by |
|---|---|---|
C_STRUCT | the C API version (v1.2.0) | any DuckDB whose C API version is at least that — then handed the whole struct, unstable region included |
C_STRUCT_UNSTABLE | an exact DuckDB release (v1.5.5) | that release only |
So a C_STRUCT binary that touches the unstable region loads happily into the
wrong DuckDB and then mis-dispatches. DuckDB's own extension-template-c says:
WARNING: When set to 1, the
duckdb_extension.hfrom theTARGET_DUCKDB_VERSIONmust be used, using any other version of the header is unsafe.
Built against v1.5.0's headers and loaded into v1.5.5, an extension calling
ClientContext::from_connection invokes duckdb_destroy_client_context on a
duckdb_connection. In practice:
double free or corruption (out)
Aborted (core dumped)
What quack-rs does
Two layers.
Build metadata. If you enable duckdb-1-5, stamp the binary
C_STRUCT_UNSTABLE with the DuckDB release you built against, so DuckDB refuses
the wrong engine at install time. With extension-ci-tools, set in the Makefile:
USE_UNSTABLE_C_API=1
TARGET_DUCKDB_VERSION=v1.5.5
generate_scaffold writes this pairing from a ScaffoldConfig, and rejects a
target_duckdb_version that does not match use_unstable_c_api:
#![allow(unused)] fn main() { use quack_rs::scaffold::ScaffoldConfig; let config = ScaffoldConfig { name: "my_ext".to_string(), use_unstable_c_api: true, target_duckdb_version: "v1.5.5".to_string(), ..ScaffoldConfig::default() }; }
Runtime guard. When a duckdb-1-5* feature is enabled, abi::check compares
the compiled-in slot count against the layout the running engine uses, resolved from
duckdb_library_version() — which lives at stable slot 7 and is therefore always
dispatched correctly. (Without those features the check reports StableOnly and
always passes.) The entry point applies it according to an AbiPolicy:
| Policy | Behaviour |
|---|---|
Strict (default) | Refuse to load, with a message naming both layouts and the fix |
AllowUnknownEngine | Refuse a layout the table knows is different; allow a release the table has no entry for |
Warn | Print the diagnostic to stderr, then load anyway (set_error would fail the load) |
Trust | Skip the check |
AllowUnknownEngine and Trust are only as safe as your knowledge of the
engine: calling into the unstable region of a layout the extension was not built
for is undefined behaviour. Before DuckDB v1.4.5 was in quack-rs's table, a
v1.5.5 build loaded into v1.4.5 under AllowUnknownEngine and segfaulted.
The two lines below are alternatives: both define the same exported symbol, so they do not compile together.
use quack_rs::abi::AbiPolicy;
// Default: Strict.
quack_rs::entry_point!(my_ext_init_c_api, register);
// Explicit — appropriate when the binary is stamped C_STRUCT_UNSTABLE, because
// DuckDB already refuses to load it into the wrong release.
quack_rs::entry_point!(my_ext_init_c_api, AbiPolicy::Trust, register);
A refused load looks like this (a v1.5.0 build loaded into DuckDB v1.5.5):
DuckDB C extension API layout mismatch: this extension was built against a
duckdb_ext_api_v1 with 545 slots, but DuckDB v1.5.5 provides 546. The extension
uses the unstable region of the C API (quack-rs feature `duckdb-1-5`), whose slot
indices differ between these releases, so loading it would dispatch to the wrong
functions. Rebuild the extension against DuckDB v1.5.5. Stamp every build that
uses the unstable region with `--abi-type C_STRUCT_UNSTABLE --duckdb-version <the
DuckDB release it was built against>` (or `USE_UNSTABLE_C_API=1` with
extension-ci-tools) so DuckDB refuses a mismatched binary at install time.
For an engine older than v1.5.0 the advice changes: such an engine cannot run a
duckdb-1-5 build at all, so the message says to build a variant without those
features or to upgrade DuckDB.
Unknown DuckDB versions
Strict also refuses a DuckDB release quack-rs has no verified layout for — a
newer release, or a -dev build. That is deliberate: DuckDB changed the unstable
region in every minor release from 1.3.0 on, and in the 1.5.2 patch release, so
"unknown" is not evidence of "compatible".
The refusal lists the fixes, best first:
- rebuild against the release you are targeting and set
QUACK_RS_TARGET_DUCKDB_VERSIONto it, which turns the check into a positive match without waiting for a quack-rs release; - upgrade quack-rs to a version whose layout table lists that release;
- if the extension does not need the unstable region, build it without the
duckdb-1-5features: the stable prefix is laid out identically in every DuckDB since v1.2.0.
It never suggests AllowUnknownEngine or Trust, for the reason above.
scripts/check-abi-table.py re-derives quack-rs's layout table from every
upstream release header and runs in CI, so the table tracks DuckDB.
Scalar Functions
A DuckDB scalar function returns one output value per input row, like the built-in
length(), upper() or sin(). This page shows how to write one in Rust with quack-rs:
as a raw callback registered through ScalarFunctionBuilder, as a safe closure with
map1 / map2, or as a set of overloads with ScalarFunctionSetBuilder.
Function signature
DuckDB calls your scalar function once per data chunk (not once per row). The signature is:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn my_fn( info: duckdb_function_info, // function metadata (rarely needed) input: duckdb_data_chunk, // input data — one or more columns output: duckdb_vector, // output vector — one value per input row ) {} }
Inside the function, you:
- Create a
VectorReaderfor each input column - Create a
VectorWriterfor the output - Loop over rows, checking for NULLs and transforming values
A panic that unwinds out of an extern "C" function aborts the process. Generate the
function with scalar_callback!, which reports a panic as a SQL error, or use the
closure constructors below (Error Handling).
Registration
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn my_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} use quack_rs::scalar::ScalarFunctionBuilder; use quack_rs::types::TypeId; unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { ScalarFunctionBuilder::new("my_fn") .param(TypeId::BigInt) // first parameter type .param(TypeId::BigInt) // second parameter type (if any) .returns(TypeId::BigInt) // return type .function(my_fn) // callback .register(con)?; } Ok(()) } }
Before calling duckdb_register_scalar_function, register checks that returns and
function are set and that no parameter or return type is a bare composite TypeId
(List, Decimal, …; use the *_logical methods for those). It also refuses a
signature that duckdb_functions() already lists under the same name, built-ins
included: DuckDB would otherwise silently replace that overload for every connection,
or make every call ambiguous. If DuckDB reports failure, register returns Err.
Validated registration
For user-configurable function names (e.g., from a config file), use try_new:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn my_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe fn demo(con: duckdb_connection, name: &str) -> Result<(), ExtensionError> { ScalarFunctionBuilder::try_new(name)? // validates name before building .param(TypeId::Varchar) .returns(TypeId::Varchar) .function(my_fn) .register(con)?; Ok(()) } }
try_new validates the name as an unquoted SQL identifier: [A-Za-z_][A-Za-z0-9_]*,
at most 256 characters, and not a DuckDB keyword that cannot be called as a function
unquoted (order, coalesce). Mixed case is allowed (DuckDB itself ships
formatReadableSize). new does not validate: it only panics if the name
contains an interior NUL byte, so use it for names known at compile time.
Closures
ScalarFunctionBuilder::map1 / map2 / map1_str / map2_str / map1_opt /
map2_opt build the whole function from a Rust closure and return a
TypedScalarFunctionBuilder. Its signature is fixed by the closure's types, so it
deliberately offers only name(), volatile() and register(con) — no returns,
param, function, extra_info, bind or init, any of which could make DuckDB
hand the closure vectors of a different type than it reads and writes. Register it
with register(con), or through a Registrar with register_typed_scalar.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { ScalarFunctionBuilder::map1("double_it", |x: i64| x * 2)?.register(con)?; Ok(()) } }
Complete example: double_it(BIGINT) → BIGINT
#![allow(unused)] fn main() { use quack_rs::prelude::*; fn live_connection() -> libduckdb_sys::duckdb_connection { std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess); } con } /// First column of the first row, as BIGINT; `None` for NULL. fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); let reader = unsafe { chunk.reader(0) }; unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) } } use quack_rs::vector::{VectorReader, VectorWriter}; use libduckdb_sys::{duckdb_function_info, duckdb_data_chunk, duckdb_vector}; unsafe extern "C" fn double_it( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { // SAFETY: DuckDB provides valid chunk and vector pointers. let reader = unsafe { VectorReader::new(input, 0) }; // column 0 let mut writer = unsafe { VectorWriter::new(output) }; let row_count = reader.row_count(); for row in 0..row_count { if unsafe { !reader.is_valid(row) } { // NULL input → NULL output // SAFETY: row < row_count, writer is valid. unsafe { writer.set_null(row) }; continue; } let value = unsafe { reader.read_i64(row) }; unsafe { writer.write_i64(row, value * 2) }; } } let con = live_connection(); unsafe { ScalarFunctionBuilder::new("double_it").param(TypeId::BigInt).returns(TypeId::BigInt) .function(double_it).register(con).unwrap(); } assert_eq!(query_i64(con, "SELECT double_it(21)"), Some(42)); // From a column, not a literal: `double_it(NULL::BIGINT)` is constant-folded. assert_eq!(query_i64(con, "SELECT double_it(i) FROM (VALUES (NULL::BIGINT)) t(i)"), None); }
Multi-parameter example: add(BIGINT, BIGINT) → BIGINT
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn add( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let col0 = unsafe { VectorReader::new(input, 0) }; // first param let col1 = unsafe { VectorReader::new(input, 1) }; // second param let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..col0.row_count() { if unsafe { !col0.is_valid(row) || !col1.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } let a = unsafe { col0.read_i64(row) }; let b = unsafe { col1.read_i64(row) }; unsafe { writer.write_i64(row, a + b) }; } } }
VARCHAR example: shout(VARCHAR) → VARCHAR
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn shout( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..reader.row_count() { if unsafe { !reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } let s = unsafe { reader.read_str(row) }; let upper = s.to_uppercase(); unsafe { writer.write_varchar(row, &upper) }; } } }
Overloading with function sets
If your function accepts different parameter types or arities, use ScalarFunctionSetBuilder
to register multiple overloads under a single name:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn add_ints(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe extern "C" fn add_doubles(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} use quack_rs::scalar::{ScalarFunctionSetBuilder, ScalarOverloadBuilder}; use quack_rs::types::TypeId; unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { ScalarFunctionSetBuilder::new("my_add") .overload( ScalarOverloadBuilder::new() .param(TypeId::Integer).param(TypeId::Integer) .returns(TypeId::Integer) .function(add_ints) ) .overload( ScalarOverloadBuilder::new() .param(TypeId::Double).param(TypeId::Double) .returns(TypeId::Double) .function(add_doubles) ) .register(con)?; } Ok(()) } }
Like AggregateFunctionSetBuilder, this builder calls duckdb_scalar_function_set_name
on every individual function before adding it to the set
(Pitfall L6).
ScalarOverloadBuilder has the same per-function settings as
ScalarFunctionBuilder, applied to that overload only: null_handling,
extra_info, varargs / varargs_logical, volatile, and (DuckDB 1.5+)
bind / init. register checks every overload for a return type and a
callback before it creates any DuckDB handle, and the error names the overload's
index. It also refuses two overloads that accept the same call: the same argument
types, or, with varargs, the same types at some argument count. f(BIGINT) and
f(BIGINT, BIGINT...) both accept f(1), so they cannot share a set; DuckDB would
accept the set and then fail every such call as ambiguous. An overload that matches
an existing function's signature is refused as for ScalarFunctionBuilder.
NULL handling
Your callback receives NULL rows whatever the setting: under the default,
DefaultNullHandling, it promises NULL-in-NULL-out and must write the NULLs
itself — call chunk.propagate_nulls(&mut writer) at the end, or use the closure
constructors (map1, map2 and their _str forms), which do it for you
(Pitfall L8,
NULL handling). A function that means to return non-NULL for
NULL input (e.g., a COALESCE-like function) sets SpecialNullHandling:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn my_coalesce_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { use quack_rs::types::NullHandling; ScalarFunctionBuilder::new("coalesce_custom") .param(TypeId::BigInt) .returns(TypeId::BigInt) .null_handling(NullHandling::SpecialNullHandling) .function(my_coalesce_fn) .register(con)?; Ok(()) } }
With SpecialNullHandling, the callback must check VectorReader::is_valid(row)
itself and decide what each NULL input produces. The closure constructors
map1_opt / map2_opt register SpecialNullHandling for you and pass the
closure an Option (None for NULL); returning None writes NULL.
Complex parameter and return types
For scalar functions that accept or return parameterized types like LIST(BIGINT),
use param_logical and returns_logical:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn flatten_list_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { use quack_rs::scalar::ScalarFunctionBuilder; use quack_rs::types::{LogicalType, TypeId}; ScalarFunctionBuilder::new("flatten_list") .param_logical(LogicalType::list(TypeId::BigInt)) // LIST(BIGINT) input .returns(TypeId::BigInt) .function(flatten_list_fn) .register(con)?; Ok(()) } }
These methods are also available on ScalarOverloadBuilder for function sets:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn my_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} fn demo() { let _ = ScalarOverloadBuilder::new() .param(TypeId::Varchar) .returns_logical(LogicalType::list(TypeId::Timestamp)) // LIST(TIMESTAMP) output .function(my_fn) ; } }
Key points
VectorReader::new(input, column_index)— the column index is zero-based- Always check
is_valid(row)before reading — skipping this reads garbage for NULL rows, and forVARCHAR/BLOBcan follow a stale pointer set_nullmust be called for NULL outputs — it callsensure_validity_writableautomatically (Pitfall L4)read_boolreturnsbool— handles DuckDB's non-0/1 boolean bytes correctly (Pitfall L5)read_strhandles both inline and pointer string formats automatically (Pitfall P7)
Varargs and volatility
These ScalarFunctionBuilder methods map to functions in DuckDB's stable C
API (v1.2.0), so they need no feature flag and work on DuckDB 1.4.x and 1.5.x
alike:
varargs(type_id: TypeId)
Declares that the function accepts a variable number of trailing arguments, all
of the given TypeId. Maps to duckdb_scalar_function_set_varargs. A composite
TypeId such as List or Decimal makes register return an error naming the
varargs slot; use varargs_logical for those.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn concat_all_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { ScalarFunctionBuilder::new("concat_all") .varargs(TypeId::Varchar) .returns(TypeId::Varchar) .function(concat_all_fn) .register(con)?; Ok(()) } }
varargs_logical(logical_type: LogicalType)
Like varargs, but accepts a LogicalType for parameterized variadic arguments.
Maps to duckdb_scalar_function_set_varargs.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn merge_lists_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { ScalarFunctionBuilder::new("merge_lists") .varargs_logical(LogicalType::list(TypeId::BigInt)) .returns_logical(LogicalType::list(TypeId::BigInt)) .function(merge_lists_fn) .register(con)?; Ok(()) } }
volatile()
Marks the function as volatile: DuckDB re-evaluates it for every row, even when
its arguments are constant (as for random()), and never merges two identical
calls in one query. Without it, DuckDB may evaluate a call with constant
arguments only once. Maps to
duckdb_scalar_function_set_volatile. Also available on the closure-built
TypedScalarFunctionBuilder.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn random_int_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { ScalarFunctionBuilder::new("random_int") .returns(TypeId::Integer) .volatile() .function(random_int_fn) .register(con)?; Ok(()) } }
DuckDB 1.5.0 additions (duckdb-1-5)
The following ScalarFunctionBuilder methods are available when the duckdb-1-5
feature is enabled:
bind(bind_fn)
Sets a custom bind callback that runs at plan time. Use this to inspect argument
types and set the return type dynamically. Maps to
duckdb_scalar_function_set_bind.
The callback takes a RawScalarBindInfo (wrap it with ScalarBindInfo::new),
not a bare duckdb_bind_info: DuckDB passes a table function's bind callback
a different, larger struct, and a callback written for one kind of function
corrupts memory on the other, so the types keep them apart. init takes a
RawScalarInitInfo for the same reason. Generate panic-safe callbacks with
scalar_bind_callback! and scalar_init_callback!; the table-function macros
do not type-check here.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn dynamic_return_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe extern "C" fn my_bind_fn(_: quack_rs::scalar::RawScalarBindInfo) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { ScalarFunctionBuilder::new("dynamic_return") .varargs(TypeId::Varchar) .returns(TypeId::Varchar) // default; overridden in bind .bind(my_bind_fn) .function(dynamic_return_fn) .register(con)?; Ok(()) } }
init(init_fn)
Sets a local-init callback invoked once per thread before execution begins. Use
this to allocate per-thread state. Maps to
duckdb_scalar_function_set_init.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn stateful_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe extern "C" fn my_init_fn(_: quack_rs::scalar::RawScalarInitInfo) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { ScalarFunctionBuilder::new("stateful_fn") .param(TypeId::BigInt) .returns(TypeId::BigInt) .init(my_init_fn) .function(stateful_fn) .register(con)?; Ok(()) } }
Typed bind data and local state
ScalarBindData<T> and ScalarLocalState<T> store a Rust value from the bind
and init callbacks with a generated, panic-safe destructor, so no hand-written
Box::from_raw is needed. A function that multiplies its argument by a factor
fixed at bind time, and counts the rows each thread processes:
#![allow(unused)] fn main() { use quack_rs::prelude::*; use quack_rs::scalar::{ ScalarBindData, ScalarBindInfo, ScalarFunctionInfo, ScalarInitInfo, ScalarLocalState, }; #[derive(Clone)] struct Factor(i64); quack_rs::scalar_bind_callback!(scaled_bind, |info| { // SAFETY: `info` is the argument of the running bind callback. let bind = unsafe { ScalarBindInfo::new(info) }; ScalarBindData::set(&bind, Factor(10)); }); quack_rs::scalar_init_callback!(scaled_init, |info| { // SAFETY: `info` is the argument of the running init callback. let init = unsafe { ScalarInitInfo::new(info) }; ScalarLocalState::set(&init, 0_u64); // rows this thread has processed }); quack_rs::scalar_callback!(scaled, |info, input, output| { // SAFETY: DuckDB passes valid handles to the running callback. let fn_info = unsafe { ScalarFunctionInfo::new(info) }; let chunk = unsafe { DataChunk::from_raw(input) }; // SAFETY: `scaled_bind` stored a `Factor`, and nothing else did. let factor = unsafe { ScalarBindData::<Factor>::get(&fn_info) }.map_or(1, |f| f.0); // SAFETY: `scaled_init` stored a `u64` on this thread; no other borrow is live. if let Some(rows) = unsafe { ScalarLocalState::<u64>::get_mut(&fn_info) } { *rows += chunk.size() as u64; } let reader = unsafe { chunk.reader(0) }; let mut writer = unsafe { VectorWriter::from_vector(output) }; for row in 0..chunk.size() { // A NULL row holds an arbitrary value; `wrapping_mul` keeps it from // overflowing, and `propagate_nulls` below overwrites it with NULL. unsafe { writer.write_i64(row, reader.read_i64(row).wrapping_mul(factor)) }; } unsafe { chunk.propagate_nulls(&mut writer) }; }); unsafe fn demo(con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> { ScalarFunctionBuilder::new("scaled") .param(TypeId::BigInt) .returns(TypeId::BigInt) .bind(scaled_bind) .init(scaled_init) .function(scaled) .register(con)?; Ok(()) } }
- Bind data must be
Clone + Send + Sync. Every executing thread reads the same value concurrently, and DuckDB copies the bound expression whenever the optimizer duplicates it (filter pushdown through a projection does). Without a copy callback the copy has no bind data at all, sosetregisters one that clonesT. Wrap data that is expensive or impossible to clone in anArc<T>. - Local state must be
Send: it is per thread, but may be freed on another. - Call
setat most once per callback. DuckDB overwrites the stored pointer on a second call without freeing the first value, so that value is leaked (never dropped). - Bind data must depend only on the call's arguments (their values when
constant, their types) and
extra_info. DuckDB'sCScalarFunctionBindData::Equalsignores bind data, so two calls with the same arguments —SELECT f(i), f(i)— are merged and share the first call's bind data. A bind callback that reads a counter, a clock or a random source needs the function markedvolatile(), which stops the merge.
Extra info
Attach arbitrary data to a scalar function using extra_info, for example a locale
or a configuration struct that parameterises its behaviour. The method is available
on both ScalarFunctionBuilder and ScalarOverloadBuilder. The pointee must be
Send + Sync: DuckDB passes the same pointer to the callback on every thread that
runs the function, and the destructor runs on whichever thread releases it.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe extern "C" fn locale_upper_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe extern "C" fn my_destroy(p: *mut std::os::raw::c_void) { drop(unsafe { Box::from_raw(p.cast::<String>()) }); } unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { use std::os::raw::c_void; let config = Box::into_raw(Box::new("en_US".to_string())).cast::<c_void>(); unsafe { ScalarFunctionBuilder::new("locale_upper") .param(TypeId::Varchar) .returns(TypeId::Varchar) .extra_info(config, Some(my_destroy)) .function(locale_upper_fn) .register(con)?; } Ok(()) } }
Inside the callback, retrieve the extra info with ScalarFunctionInfo::get_extra_info().
ScalarFunctionInfo
ScalarFunctionInfo wraps the duckdb_function_info handle provided to a scalar
function callback. It exposes:
get_extra_info() -> *mut c_void— retrieves the extra-info pointer set during registrationset_error(message)— reports an error, which fails the query
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; use quack_rs::scalar::ScalarFunctionInfo; unsafe extern "C" fn my_fn( info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let info = unsafe { ScalarFunctionInfo::new(info) }; let extra = unsafe { info.get_extra_info() }; // ... use extra info, or report errors via info.set_error("...") ... } }
With the duckdb-1-5 feature, ScalarFunctionInfo also provides:
get_bind_data() -> *mut c_void— retrieves bind data set during the bind callbackget_state() -> *mut c_void— retrieves per-thread state set during the init callback
ScalarBindInfo (duckdb-1-5)
ScalarBindInfo wraps the duckdb_bind_info handle provided to a scalar function
bind callback. It exposes:
argument_count() -> u64— number of argumentsargument(index) -> Option<Expression>— the argument atindexas an RAIIExpression, before DuckDB casts it to the parameter type. On DuckDB before 1.5.5 it fails the bind and returnsNone, because those releases abort the process when the argument is a scalar subqueryget_argument(index) -> duckdb_expression— the raw handle, with that hazard left to the callerget_extra_info() -> *mut c_void— the extra-info pointer from registrationset_bind_data(data, destroy)— stores per-query data retrievable during executionset_bind_data_copy(copy)— the callback DuckDB uses to duplicate that data when it copies the bound expression; without one the copy's bind data is NULLset_error(message)— reports an errorget_client_context() -> ClientContext— access to the connection's catalog and config
ScalarInitInfo (duckdb-1-5)
ScalarInitInfo wraps the duckdb_init_info handle provided to a scalar function
init callback. It exposes:
get_extra_info() -> *mut c_void— the extra-info pointer from registrationget_bind_data() -> *mut c_void— the bind data from the bind callbackset_state(state, destroy)— stores per-thread state retrievable during executionset_error(message)— reports an errorget_client_context() -> ClientContext— access to the connection's catalog and config
Aggregate Functions
This page shows how to write a DuckDB aggregate function in Rust with quack-rs:
the callbacks DuckDB calls, their signatures, and how AggregateFunctionBuilder
registers them. An aggregate function reduces many rows to one value per group,
like SUM(), COUNT() or AVG(). Because DuckDB aggregates in parallel, it
also has a combine step that merges partial results from parallel workers.
Known DuckDB limitation
Out-of-bounds state reads. Two query shapes make DuckDB call every C-API aggregate's
updatewith a state array holding one state while passingcount > 1rows, so the callback readsstates[1..count]past the end of the array (undefined behaviour, in any C-API aggregate, whether built with quack-rs or by hand):
- Window aggregates whose frame is the whole partition, e.g.
agg(x) OVER ()—WindowConstantAggregator(src/function/window/window_constant_aggregator.cpp, ~lines 106 and 296–299 in DuckDB 1.5.5).- Ordered aggregates, e.g.
agg(x ORDER BY y)—src/function/aggregate/sorted_aggregate_function.cpp, ~lines 630–633.Both paths pass a
CONSTANT_VECTORof states because the function has nosimple_update(which the C API cannot set), andCAPIAggregateUpdate(src/main/capi/aggregate_function-c.cpp, ~lines 92–110) passes the vector's data pointer to the extension without flattening it. This is a defect in DuckDB's C API, not in quack-rs, and it cannot be detected from inside the callback: readingstates[1]to check is itself the out-of-bounds read. Until DuckDB fixes it, do not use C-API aggregates in those two query shapes.Reported upstream as duckdb/duckdb#26109.
The aggregate lifecycle
flowchart TD
REG["<b>Registration</b><br/>AggregateFunctionBuilder<br/>→ duckdb_register_aggregate_function"]
REG --> SIZE
SIZE --> INIT
INIT --> UPDATE
UPDATE --> COMBINE
COMBINE --> FINAL
FINAL --> DESTROY
SIZE["<b>state_size</b>()<br/>How many bytes to allocate per group?"]
INIT["<b>state_init</b>(state)<br/>Initialise a fresh state"]
UPDATE["<b>update</b>(chunk, states[])<br/>Process one input batch<br/>(NULL rows included — check is_valid)"]
COMBINE["<b>combine</b>(src[], tgt[], count)<br/>Merge partial results from parallel workers<br/>⚠️ Pitfall L1: target starts fresh — copy ALL config fields"]
FINAL["<b>finalize</b>(states[], out, count, offset)<br/>Write count results at out[offset..], once per result batch"]
DESTROY["<b>state_destroy</b>(states[], count)<br/>Free memory — after finalize and for combine<br/>sources after the merge (not every state: see Known Limitations)"]
style COMBINE fill:#fff3cd,stroke:#e6ac00,color:#333
DuckDB may call combine many times as it merges partial results. A combine
target holds whatever state_init set up, not a copy of the source, so combine
must carry every field across (Pitfall L1). state_size is called whenever an
operator sizes its state buffers, not once at registration, so it must always
return the same value. destroy runs after finalize, and on combine's source
states once they have been merged.
combine must leave its source states unchanged
(Pitfall L15).
A window's segment tree combines the same state into every frame that covers it,
from several threads at once. A combine that moves data out of its source
(mem::take, or zeroing a counter) is right for the first frame and wrong for the
rest: in quack-rs's regression test, a sliding-window sum written that way was
wrong on 4985 of 5000 rows. Read the source and copy or clone what the target
needs.
Registration
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { count: i64 } impl AggregateState for MyState {} unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} use quack_rs::aggregate::AggregateFunctionBuilder; use quack_rs::types::TypeId; unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { AggregateFunctionBuilder::new("my_agg") .param(TypeId::BigInt) // input type(s) .returns(TypeId::BigInt) // output type .ffi_state::<MyState>() // state_size + init + destructor .update(update) .combine(combine) .finalize(finalize) .register(con)?; } Ok(()) } }
register returns an error if the return type or any of the five required
callbacks (state_size, init, update, combine, finalize) is missing. The
builder treats the destructor as optional (without one, register installs a
no-op destructor; see
Pitfall L13),
but FfiState<T> needs its destroy_callback to drop each T.
.ffi_state::<MyState>() sets state_size, init and destructor together from
FfiState<MyState>, so the three cannot describe different states; see
State Management. update,
combine and finalize read the state through FfiState::<MyState>::with_state /
with_state_mut with the same type.
Callback signatures
With FfiState<T> you do not write state_size, init or destroy yourself:
ffi_state::<T>() installs FfiState::<T>::size_callback, init_callback and
destroy_callback. The wrappers below show the signature DuckDB calls each one
with, and what it does. update, combine and finalize are yours to write.
state_size
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { config_field: i64, accumulator: i64 } impl MyState { fn accumulate(&mut self, v: i64) { self.accumulator += v; } fn result(&self) -> i64 { self.accumulator } } impl AggregateState for MyState {} unsafe extern "C" fn state_size(info: duckdb_function_info) -> idx_t { unsafe { FfiState::<MyState>::size_callback(info) } } }
Returns the number of bytes DuckDB allocates per group, FfiState::<MyState>::size():
a tag word followed by MyState itself, padded to whole words (or, for a state
larger than 256 bytes or aligned more strictly than usize, a Box<MyState>
pointer).
state_init
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { config_field: i64, accumulator: i64 } impl MyState { fn accumulate(&mut self, v: i64) { self.accumulator += v; } fn result(&self) -> i64 { self.accumulator } } impl AggregateState for MyState {} unsafe extern "C" fn state_init(info: duckdb_function_info, state: duckdb_aggregate_state) { unsafe { FfiState::<MyState>::init_callback(info, state) }; } }
Writes MyState::default() into the DuckDB-allocated state slot (or, for a boxed
state, a Box holding it), then writes the tag that marks the slot initialised. A
panic in default() is caught and reported to DuckDB as a query error.
update
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { config_field: i64, accumulator: i64 } impl MyState { fn accumulate(&mut self, v: i64) { self.accumulator += v; } fn result(&self) -> i64 { self.accumulator } } impl AggregateState for MyState {} unsafe extern "C" fn update( _info: duckdb_function_info, input: duckdb_data_chunk, states: *mut duckdb_aggregate_state, ) { let reader = unsafe { VectorReader::new(input, 0) }; let row_count = reader.row_count(); for row in 0..row_count { if unsafe { !reader.is_valid(row) } { continue; } let value = unsafe { reader.read_i64(row) }; let state_ptr = unsafe { *states.add(row) }; if let Some(st) = unsafe { FfiState::<MyState>::with_state_mut(state_ptr) } { st.accumulate(value); } } } }
states[row] is the state of that row's group; rows in the same group share one
state. update receives NULL rows too, so skip rows where is_valid is false (see
NULL Handling).
combine
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { config_field: i64, accumulator: i64 } impl MyState { fn accumulate(&mut self, v: i64) { self.accumulator += v; } fn result(&self) -> i64 { self.accumulator } } impl AggregateState for MyState {} unsafe extern "C" fn combine( _info: duckdb_function_info, source: *mut duckdb_aggregate_state, target: *mut duckdb_aggregate_state, count: idx_t, ) { for i in 0..count as usize { let src = unsafe { FfiState::<MyState>::with_state(*source.add(i)) }; let tgt = unsafe { FfiState::<MyState>::with_state_mut(*target.add(i)) }; if let (Some(s), Some(t)) = (src, tgt) { // ⚠️ MUST copy ALL fields — see Pitfall L1 t.config_field = s.config_field; // configuration t.accumulator += s.accumulator; // data } } } }
Pitfall L1 — critical: Target states are fresh states, set up by
state_init(withFfiState<T>::init_callback, aT::default()), not copies of the source. You must copy every field, including configuration fields set duringupdate. Forgetting even one config field produces silently wrong results. See Pitfall L1.
finalize
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { config_field: i64, accumulator: i64 } impl MyState { fn accumulate(&mut self, v: i64) { self.accumulator += v; } fn result(&self) -> i64 { self.accumulator } } impl AggregateState for MyState {} unsafe extern "C" fn finalize( _info: duckdb_function_info, source: *mut duckdb_aggregate_state, result: duckdb_vector, count: idx_t, offset: idx_t, ) { let mut writer = unsafe { VectorWriter::new(result) }; for i in 0..count as usize { let state_ptr = unsafe { *source.add(i) }; match unsafe { FfiState::<MyState>::with_state(state_ptr) } { Some(st) => unsafe { writer.write_i64(offset as usize + i, st.result()) }, None => unsafe { writer.set_null(offset as usize + i) }, } } } }
offset is non-zero when DuckDB writes the results into part of a larger vector.
Always add it to the output index.
state_destroy
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { config_field: i64, accumulator: i64 } impl AggregateState for MyState {} unsafe extern "C" fn state_destroy(states: *mut duckdb_aggregate_state, count: idx_t) { unsafe { FfiState::<MyState>::destroy_callback(states, count) }; } }
destroy_callback drops the T in each state whose tag matches (freeing its box,
if T is boxed), clearing the tag first, so a second call on the same state is a
no-op. See Pitfall L2.
Complex parameter and return types
For functions that accept or return parameterized types like LIST(BIGINT),
MAP(VARCHAR, INTEGER), or STRUCT(...), use param_logical and
returns_logical instead of param and returns:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { count: i64 } impl AggregateState for MyState {} unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} use quack_rs::aggregate::AggregateFunctionBuilder; use quack_rs::types::{LogicalType, TypeId}; unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { AggregateFunctionBuilder::new("retention") .param(TypeId::Boolean) .param(TypeId::Boolean) .returns_logical(LogicalType::list(TypeId::Boolean)) // LIST(BOOLEAN) .ffi_state::<MyState>() // state_size + init + destructor .update(update) .combine(combine) .finalize(finalize) .register(con)?; } Ok(()) } }
param_logical and param can be interleaved — the parameter position is
determined by the total number of calls made so far:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; fn demo() { let _ = AggregateFunctionBuilder::new("my_func") .param(TypeId::Varchar) // position 0: VARCHAR .param_logical(LogicalType::list(TypeId::BigInt)) // position 1: LIST(BIGINT) .param(TypeId::Integer) // position 2: INTEGER .returns(TypeId::BigInt) // ... ; } }
If both returns and returns_logical are called, the logical type takes precedence.
Extra info
extra_info attaches arbitrary data to an aggregate function, for example
configuration that parameterises its behaviour. DuckDB calls the destroy callback
to free the data when the function is dropped:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { count: i64 } impl AggregateState for MyState {} unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} unsafe extern "C" fn my_destroy(p: *mut std::os::raw::c_void) { drop(unsafe { Box::from_raw(p.cast::<u64>()) }); } unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { use std::os::raw::c_void; let config = Box::into_raw(Box::new(42u64)).cast::<c_void>(); unsafe { AggregateFunctionBuilder::new("my_agg") .param(TypeId::BigInt) .returns(TypeId::BigInt) .extra_info(config, Some(my_destroy)) .ffi_state::<MyState>() // state_size + init + destructor .update(update) .combine(combine) .finalize(finalize) .register(con)?; } Ok(()) } }
Inside callbacks, retrieve the extra info with AggregateFunctionInfo::get_extra_info().
AggregateFunctionInfo
AggregateFunctionInfo wraps the duckdb_function_info handle that DuckDB passes
to every aggregate callback except the destructor. It exposes:
get_extra_info() -> *mut c_void: the extra-info pointer set at registration.set_error(message): fails the current query withmessage. Called fromfinalize, it also leaves some states undestroyed; see Known Limitations.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; use quack_rs::aggregate::AggregateFunctionInfo; unsafe extern "C" fn update( info: duckdb_function_info, input: duckdb_data_chunk, states: *mut duckdb_aggregate_state, ) { let info = unsafe { AggregateFunctionInfo::new(info) }; let extra = unsafe { info.get_extra_info() }; // ... use extra info, or report errors via info.set_error("...") ... } }
Next steps
- State Management —
FfiState<T>,AggregateState, and lifecycle details - Overloading with Function Sets — register multiple signatures under one name
Aggregate State
This page covers aggregate state in a DuckDB aggregate function written in Rust:
the AggregateState trait and FfiState<T>, which manages each state's lifecycle
(allocation, initialisation, access and destruction) so that you do not write
raw-pointer code for it.
Known DuckDB limitation. Every C-API aggregate, and so every aggregate that uses
FfiState<T>, reads out of bounds underagg(x) OVER ()(whole-partition window frames) andagg(x ORDER BY y). This is a DuckDB C API defect; see Aggregate Functions for the details and DuckDB source lines. Do not use C-API aggregates in those two query shapes.Reported upstream as duckdb/duckdb#26109.
AggregateState trait
Any type that is Default + Send + Sync + 'static can be used as aggregate state by
implementing the AggregateState marker trait. The Sync bound is new in 0.18.0: a
window's segment tree lets several threads read the same state as a combine source at
once, so a state containing a Cell or RefCell would race. Use atomics or a Mutex
instead, or keep that data outside the state.
#![allow(unused)] fn main() { use quack_rs::aggregate::AggregateState; #[derive(Default, Debug)] struct MyState { config: usize, // set in update, must be propagated in combine total: i64, // accumulated data } impl AggregateState for MyState {} }
AggregateState has no required methods. state_init uses Default to create each
fresh state.
FfiState<T>
FfiState<T> names the layout of the bytes DuckDB allocates for each group's
state, and the callbacks that manage them. The type itself is never constructed.
A small T, aligned no more strictly than usize and at most 256 bytes, is
stored in those bytes directly. A larger or more strictly aligned T is boxed,
and the slot holds the pointer. On wasm32, where usize is 4 bytes but u64,
i64 and f64 are 8-byte aligned, a state containing one of them is therefore
boxed. Either way the slot starts with a tag.
Memory layout
DuckDB-allocated slot (state_size bytes, a multiple of sizeof(usize)):
[ tag: usize ][ T, padded to whole words ] T stored inline
[ tag: usize ][ *mut T ] T boxed
│
└──→ Box<T> (on the Rust heap)
Storing T inline matters because DuckDB 1.4.4 to 1.5.5 does not destroy
every state. When a grouped aggregate's result scan stops early (a LIMIT
above it, an error, an interrupt), the states it never reached are never
destroyed; so is one state per row of a window frame with EXCLUDE. An inline
T's bytes belong to DuckDB, which frees them with the hash table; only what
T itself owns on the heap, or a boxed T's box, leaks. See
Known Limitations.
The tag marks the slot initialised. When one state_init call fails (a
panicking T::default(), say), DuckDB 1.4.4 to 1.5.5 still runs the
destructor over every state it created, including states whose state_init
never ran, so a slot can hold arbitrary bytes. destroy_callback drops only a
slot carrying the tag init_callback wrote, and clears the tag first. The tag
is derived from a hash of T's TypeId (and, for a boxed T, the box's
address), not from the slot's address, because DuckDB moves states by copying
their bytes. The check turns dropping garbage from a certainty into a matter
of chance: uninitialised bytes that happen to equal the tag. It mitigates the
DuckDB defect (described in the repository's docs/upstream-duckdb-reports.md);
it cannot guarantee against it.
Lifecycle callbacks
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::aggregate::{AggregateState, FfiState}; #[derive(Default, Debug)] struct MyState { config: usize, total: i64 } impl AggregateState for MyState {} unsafe fn demo(_info: duckdb_function_info, info: duckdb_function_info, state: duckdb_aggregate_state, states: *mut duckdb_aggregate_state, count: idx_t) { // state_size: DuckDB calls this whenever an operator sizes its state buffers FfiState::<MyState>::size_callback(_info); // Returns: FfiState::<MyState>::size() (on a 64-bit target, a tag word, then the // usize and i64 inline) // state_init: DuckDB calls this for every state slot it allocates, combine // targets included FfiState::<MyState>::init_callback(info, state); // Effect: writes MyState::default() into the slot (or a box holding it), then the tag // destructor: DuckDB calls this after finalize, on combine's source states once // merged, and (after a failed state_init) on states never initialised, which // the tag makes it skip; not on every state (see Known Limitations) FfiState::<MyState>::destroy_callback(states, count); // Effect: for each state whose tag matches: clear the tag, then drop the T } }
Wiring them up: ffi_state::<T>()
The recommended way to register those three callbacks is ffi_state::<T>()
(new in 0.18.0). It installs size_callback, init_callback and
destroy_callback for the same T in one call, so the size DuckDB allocates
and the state init writes cannot disagree. Wired one by one with the
state_size, init and destructor setters, a size callback for one type
paired with an init callback for a larger one writes past DuckDB's allocation.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct MyState { config: usize, total: i64 } impl AggregateState for MyState {} unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { AggregateFunctionBuilder::new("my_agg") .param(TypeId::BigInt) .returns(TypeId::BigInt) .ffi_state::<MyState>() // state_size + init + destructor .update(update) .combine(combine) .finalize(finalize) .register(con)?; } Ok(()) } }
AggregateOverloadBuilder has the same method, for each overload of an
AggregateFunctionSetBuilder; the set builder itself has
none. update, combine and finalize still read the state through
FfiState::<T>::with_state / with_state_mut with the same T.
Accessing state in callbacks
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::aggregate::{AggregateState, FfiState}; #[derive(Default, Debug)] struct MyState { config: usize, total: i64 } impl AggregateState for MyState {} unsafe fn demo(state_ptr: duckdb_aggregate_state, delta: i64) { // Immutable access (in finalize, combine source): if let Some(st) = FfiState::<MyState>::with_state(state_ptr) { let value = st.total; } // Mutable access (in update, combine target): if let Some(st) = FfiState::<MyState>::with_state_mut(state_ptr) { st.total += delta; } } }
The methods return Option<&T> and Option<&mut T> respectively: None if the
slot's tag does not match, which happens after destroy_callback has run or when
T::default() panicked in state_init. Returning Option instead of panicking
keeps a panic from unwinding across the FFI boundary
(Pitfall L3).
The double-free problem — solved
Without quack-rs, a naive destructor looks like:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, idx_t}; struct MyState; // A hand-written boxed layout: this is the code *without* quack-rs. #[repr(C)] struct FfiState<T> { inner: *mut T } // ❌ Naive — causes double-free if DuckDB calls destroy twice unsafe extern "C" fn destroy(states: *mut duckdb_aggregate_state, count: idx_t) { for i in 0..count as usize { let ffi = &mut *(*states.add(i) as *mut FfiState<MyState>); drop(Box::from_raw(ffi.inner)); // inner is now dangling — crash on second call } } }
FfiState::destroy_callback clears the slot's tag before dropping the T,
and drops only a slot whose tag matches. If DuckDB calls destroy again, the tag
no longer matches, the slot is skipped, and with_state returns None.
Testing state logic without DuckDB
AggregateTestHarness<S> simulates the DuckDB aggregate lifecycle in pure Rust:
#![allow(unused)] fn main() { use quack_rs::aggregate::AggregateState; #[derive(Default, Debug)] struct MyState { config: usize, total: i64 } impl AggregateState for MyState {} use quack_rs::testing::AggregateTestHarness; #[test] fn combine_propagates_config() { let mut source = AggregateTestHarness::<MyState>::new(); source.update(|s| { s.config = 5; // config field set during update s.total += 100; }); let mut target = AggregateTestHarness::<MyState>::new(); target.combine(&source, |src, tgt| { tgt.config = src.config; // must propagate config — Pitfall L1 tgt.total += src.total; }); let result = target.finalize(); assert_eq!(result.config, 5, "config must be propagated in combine"); assert_eq!(result.total, 100); } }
See the Testing Guide for the full test strategy.
Overloading with Function Sets
A DuckDB aggregate function can have several signatures under one name, registered
together as a function set. This page shows how to overload an aggregate
function in Rust with AggregateFunctionSetBuilder, including variadic aggregates
such as retention(c1, c2, ..., c32) and overloads with different return types.
Known DuckDB limitation. C-API aggregates — including every overload in a set — read out of bounds under
agg(x) OVER ()(whole-partition window frames) andagg(x ORDER BY y). This is a DuckDB C API defect; see Aggregate Functions for the details and DuckDB source lines. Do not use C-API aggregates in those two query shapes.Reported upstream as duckdb/duckdb#26109.
Note: For scalar function overloads, see
ScalarFunctionSetBuilder.
When to use function sets
Use AggregateFunctionSetBuilder when you need:
- Multiple type signatures for the same function name (e.g.,
my_agg(INT)andmy_agg(BIGINT)) - Variadic arity under one name (e.g.,
retention(2 columns),retention(3 columns), ...) - Overloads that return different types (see Per-overload return types)
For a single signature, use AggregateFunctionBuilder directly.
Registration
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct RetentionState { hits: u32 } impl AggregateState for RetentionState {} unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} use quack_rs::aggregate::AggregateFunctionSetBuilder; use quack_rs::types::TypeId; unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { AggregateFunctionSetBuilder::new("retention") .returns(TypeId::Varchar) .overloads(2..=3, |n, builder| { // Each overload gets `n` BOOLEAN parameters let b = (0..n).fold(builder, |b, _| b.param(TypeId::Boolean)); b.ffi_state::<RetentionState>() // state_size + init + destructor .update(update) .combine(combine) .finalize(finalize) }) .register(con)?; } Ok(()) } }
overloads takes a RangeInclusive<usize> and a closure, called once per arity
n with a fresh AggregateOverloadBuilder, that returns the configured overload.
The set builder gives every member the set's name when it registers them.
Per-overload return types
DuckDB resolves an aggregate overload from its parameter types and arity
alone — the return type plays no part in resolution. Members of one set are
therefore free to return different types. This is how DuckDB's own arg_max
works:
arg_max(ANY, ANY) -> ANY
arg_max(ANY, ANY, ANY) -> ANY[]
Set the return type on the overload with AggregateOverloadBuilder::returns (or
returns_logical), and add each one with overload:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct IntState { sum: i64 } impl AggregateState for IntState {} unsafe extern "C" fn int_update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn int_combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn int_finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} #[derive(Default)] struct StrState { longest: String } impl AggregateState for StrState {} unsafe extern "C" fn str_update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn str_combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn str_finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { use quack_rs::aggregate::{AggregateFunctionSetBuilder, AggregateOverloadBuilder}; use quack_rs::types::TypeId; unsafe { AggregateFunctionSetBuilder::new("my_agg") .overload( AggregateOverloadBuilder::new() .param(TypeId::Integer) .returns(TypeId::Integer) // my_agg(INTEGER) -> INTEGER .ffi_state::<IntState>() .update(int_update) .combine(int_combine) .finalize(int_finalize), ) .overload( AggregateOverloadBuilder::new() .param(TypeId::Varchar) .returns(TypeId::Varchar) // my_agg(VARCHAR) -> VARCHAR .ffi_state::<StrState>() .update(str_update) .combine(str_combine) .finalize(str_finalize), ) .register(con)?; } Ok(()) } }
Each overload carries its own callbacks, so overloads with different parameter
types can use different state types: call ffi_state::<T>() on each overload
with that overload's state type (AggregateFunctionSetBuilder itself has no
ffi_state), and read it in that overload's update, combine and finalize
with FfiState::<T>::with_state / with_state_mut for the same T.
Which return type wins
For each overload, in order:
AggregateOverloadBuilder::returns_logicalAggregateOverloadBuilder::returnsAggregateFunctionSetBuilder::returns_logical(the set-level default)AggregateFunctionSetBuilder::returns(the set-level default)
Registration fails, naming the overload index, if an overload reaches the end of that list with nothing set or is missing a required callback. It also fails if two overloads take the same parameter types. All of this is checked for every overload before any DuckDB handle is created.
Other per-overload settings
Each overload can also carry its own extra_info
(AggregateOverloadBuilder::extra_info), read in that overload's callbacks with
AggregateFunctionInfo::get_extra_info. Ownership works as on
AggregateFunctionBuilder: DuckDB frees it once registration has handed it over,
even if registration then fails; the builder frees it if it never gets that far.
AggregateOverloadBuilder::null_handling sets an overload's
NULL handling.
overload and overloads may be mixed on one builder; overloads register in
the order they were added.
Note:
AggregateOverloadBuilderwas calledOverloadBuilderbefore v0.18.0. The old name is still exported as a deprecated alias atquack_rs::aggregate::builder::OverloadBuilder.
The silent name bug — solved
Pitfall L6: When using a function set, the name must be set on each individual
duckdb_aggregate_functionviaduckdb_aggregate_function_set_name, not just on the set. If any member lacks a name, it is silently not registered — no error is returned.DuckDB's C API documentation does not mention this. It was found by reading DuckDB's C++ test code at
test/api/capi/test_capi_aggregate_functions.cpp. Induckdb-behavioral, 6 of 7 functions failed to register silently due to this bug.
AggregateFunctionSetBuilder::register calls duckdb_aggregate_function_set_name
on every member, whether it was added with overload or overloads.
See Pitfall L6.
Complex return types
If all overloads share one complex return type, set it once on the set builder as a default, rather than repeating it on every overload:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct RetentionState { hits: u32 } impl AggregateState for RetentionState {} unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { use quack_rs::aggregate::AggregateFunctionSetBuilder; use quack_rs::types::{LogicalType, TypeId}; unsafe { AggregateFunctionSetBuilder::new("retention") .returns_logical(LogicalType::list(TypeId::Boolean)) // default for every overload .overloads(2..=32, |n, builder| { (0..n).fold(builder, |b, _| b.param(TypeId::Boolean)) .ffi_state::<RetentionState>() .update(update) .combine(combine) .finalize(finalize) }) .register(con)?; } Ok(()) } }
Individual overloads can also use param_logical for complex parameter types:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; fn demo() { let _ = AggregateFunctionSetBuilder::new("retention") .overloads(2..=8, |n, builder| { builder .param(TypeId::Interval) .param_logical(LogicalType::list(TypeId::Timestamp)) // LIST(TIMESTAMP) parameter // ... }) ; } }
Why not varargs?
DuckDB's C API has no duckdb_aggregate_function_set_varargs. A variadic aggregate
must therefore be registered as one overload per supported arity, which overloads
does in one call.
Note: Scalar functions do support varargs, through
ScalarFunctionBuilder::varargs()(stable C API, no feature flag needed).
Table Functions
A DuckDB table function returns a result set rather than a single value, and is
called in the FROM clause: SELECT * FROM my_function(args). This page shows how
to write one in Rust with quack-rs. DuckDB drives a table function through three
callbacks: bind, init and scan.
quack-rs provides two layers for registering table functions:
TypedTableFunctionBuilder<S>(recommended for new extensions) — closure-based API that hides bind/init/scan trampolines behind safe Rust closures and gives every execution a fresh, typed scan state built from whatbindproduced.TableFunctionBuilder— the underlying raw builder used byTypedTableFunctionBuilderinternally. Reach for it when you need fine-grained control: parallel scans (InitInfo::set_max_threadsabove 1, usually withlocal_initfor per-thread state), projection pushdown with column filtering, or callback shapes that don't fit the "produce state in bind, mutate it in scan" model.
Both builders are backed by the helper types BindInfo, InitInfo, FunctionInfo,
FfiBindData<T>, FfiInitData<T>, and FfiLocalInitData<T>.
Lifecycle
| Phase | Callback | Called when | Typical work |
|---|---|---|---|
| bind | bind_fn | Query is planned (once per plan) | Extract parameters; declare output columns; store configuration in bind data |
| init | init_fn | Each execution of the plan starts | Allocate per-scan state (cursor, row index, etc.) |
| scan | scan_fn | Each output batch | Fill the output chunk with rows; set its size with DataChunk::set_size |
DuckDB calls the scan callback repeatedly until it sets the chunk size to 0, which signals the end of the results.
Bind once, init many times. DuckDB keeps the bind data for as long as the bound plan lives and runs
initagainst it on every execution: eachEXECUTEof a prepared statement, each iteration of a recursive CTE that references the function. Treat bind data as immutable after bind and build anything a scan consumes (cursors, open files) ininit.
Closure-based typed state (with_state)
For the common "take parameters at bind, stream rows until exhausted" pattern,
TypedTableFunctionBuilder<S> replaces all three callback trampolines with two
closures. With with_state, the state returned by bind is a template: every
execution of the plan scans a fresh clone() of it, so S must be Clone + Send.
#![allow(unused)] fn main() { use quack_rs::prelude::*; #[derive(Clone)] struct State { remaining: u64, } fn register(reg: &impl Registrar) -> ExtResult<()> { let builder = TableFunctionBuilder::new("count_down") .param(TypeId::BigInt) // 1. bind closure: declare the output schema, read parameters, // return the template scan state (cloned for every execution). .with_state::<State, _>(|bind| { bind.add_result_column("n", TypeId::BigInt); let raw = unsafe { bind.get_parameter_value(0) }; Ok(State { remaining: raw.as_i64_or(0).max(0) as u64 }) }) // 2. scan closure: mutate state, write rows, set chunk size. .scan(|state, chunk| { if state.remaining == 0 { unsafe { chunk.set_size(0) }; return Ok(()); } let mut writer = unsafe { chunk.writer(0) }; unsafe { writer.write_i64(0, state.remaining as i64) }; state.remaining -= 1; unsafe { chunk.set_size(1) }; Ok(()) }) .build()?; unsafe { reg.register_table(builder) } } }
Separate bind data and scan state (with_bind_init)
When the scan state is expensive or impossible to clone (it owns a file handle,
a large buffer, a connection), or the parameters and the cursor are naturally
separate, use with_bind_init. bind returns immutable bind data B
(Send + Sync); init builds a fresh scan state S from &B for every
execution:
#![allow(unused)] fn main() { use quack_rs::prelude::*; struct Params { n: i64 } struct Cursor { next: i64, end: i64 } // no Clone needed fn register(reg: &impl Registrar) -> ExtResult<()> { let builder = TableFunctionBuilder::new("count_up") .param(TypeId::BigInt) .with_bind_init( |bind| { bind.add_result_column("n", TypeId::BigInt); let n = unsafe { bind.get_parameter_value(0) }.as_i64_or(0); Ok(Params { n }) }, |params: &Params| Ok(Cursor { next: 1, end: params.n }), ) .scan(|cursor, chunk| { if cursor.next > cursor.end { unsafe { chunk.set_size(0) }; return Ok(()); } unsafe { chunk.writer(0).write_i64(0, cursor.next); chunk.set_size(1); } cursor.next += 1; Ok(()) }) .build()?; unsafe { reg.register_table(builder) } } }
What you get for free
- No hand-written
unsafe extern "C" fntrampolines.TypedTableFunctionBuildergenerates them internally. - Typed scan state. The
scanclosure receives&mut S, freshly built for each execution (a clone of thewith_statetemplate, orinit(&B)forwith_bind_init) — no manualFfiBindData/FfiInitDatashuffling, and a prepared statement can be executed any number of times. - Panic safety. User closures run inside
catch_unwind. A panic is reported as a bind, init or scan error, and the scan sets the chunk size to zero, so the query fails cleanly instead of unwinding across the FFI boundary. - Error propagation. Return
Err(ExtensionError::new("..."))from any closure to report a SQL error to DuckDB.
Trade-offs and threading
Smust beSend + 'static(plusCloneforwith_state).Syncis not required, soTypedTableFunctionBuilderforces scans to run on a single worker by callingInitInfo::set_max_threads(1)internally.- The typed builder does not offer projection pushdown: with pushdown on, the
scan's chunk holds only the projected columns and a closure written against the
declared schema would write the wrong column.
build()returns an error ifprojection_pushdown(true)was set on the raw builder beforewith_state/with_bind_init. Use the raw builder for pushdown. - The bind closure must declare at least one column. With
duckdb-1-5, a bind that declares none is reported as an ordinary bind error; DuckDB itself raises anINTERNAL Errorwith a C++ stack trace for it (and does so for a raw bind callback, which quack-rs cannot check after it returns — callset_erroryourself). - Extensions that need multi-worker parallelism (
set_max_threadsabove 1, withlocal_init+ thread-local buffers) should use the rawTableFunctionBuilderdirectly. TypedTableFunctionBuilder::build()returns a fully configuredTableFunctionBuilder, so you can still pass it through anyRegistrar— includingMockRegistrarfor unit tests. Registration fails if you then enableprojection_pushdownon it or replace itsbind,init,local_init,scanorextra_info: the generated callbacks read one another's data.
Builder API
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info}; unsafe extern "C" fn my_bind_callback(_: duckdb_bind_info) {} unsafe extern "C" fn my_init_callback(_: duckdb_init_info) {} unsafe extern "C" fn my_scan_callback(_: duckdb_function_info, _: duckdb_data_chunk) {} unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { use quack_rs::table::{TableFunctionBuilder, BindInfo, FfiBindData, FfiInitData}; use quack_rs::types::TypeId; TableFunctionBuilder::new("my_function") .param(TypeId::BigInt) // positional parameter types .bind(my_bind_callback) // declare output columns inside bind .init(my_init_callback) .scan(my_scan_callback) .register(con)?; Ok(()) } }
Output columns are declared inside the bind callback using BindInfo::add_result_column,
not on the builder itself.
State management
Bind data
Bind data persists from the bind phase through all scan batches — and through every
later execution of the same plan (see Bind once, init many times above), possibly
read from several threads at once, so FfiBindData::set requires T: Send + Sync.
Use FfiBindData<T> to allocate it safely:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_bind_info; use quack_rs::table::{BindInfo, FfiBindData}; struct MyBindData { limit: i64, } unsafe extern "C" fn my_bind(info: duckdb_bind_info) { // `get_parameter_value` returns an RAII `Value`; a NULL argument reads as the default. let n = unsafe { BindInfo::new(info).get_parameter_value(0) }.as_i64_or(0); unsafe { FfiBindData::<MyBindData>::set(info, MyBindData { limit: n }) }; } }
FfiBindData::set stores the value and registers a destructor so DuckDB frees
it at the right time — no Box::into_raw / Box::from_raw needed.
Init (scan) state
Per-scan state (e.g., a current row index) uses FfiInitData<T> (T: Send + Sync,
since concurrent scan threads share it):
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_init_info; use quack_rs::table::FfiInitData; struct MyScanState { pos: i64, } unsafe extern "C" fn my_init(info: duckdb_init_info) { unsafe { FfiInitData::<MyScanState>::set(info, MyScanState { pos: 0 }) }; } }
Complete example: generate_series_ext
The hello-ext example registers generate_series_ext(n BIGINT) which emits
integers 0 .. n-1. See examples/hello-ext/src/lib.rs for the full source.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, DuckDBSuccess}; use quack_rs::data_chunk::DataChunk; use quack_rs::table::{BindInfo, FfiBindData, FfiInitData, TableFunctionBuilder}; use quack_rs::types::TypeId; struct GsBindData { total: i64 } struct GsScanState { pos: i64 } // Bind: extract `n`, register one output column unsafe extern "C" fn gs_bind(info: duckdb_bind_info) { let bind_info = unsafe { BindInfo::new(info) }; // Value is RAII — automatically destroyed when dropped. // A NULL argument reads as the default rather than aborting. let n = unsafe { bind_info.get_parameter_value(0) }.as_i64_or(0); bind_info.add_result_column("value", TypeId::BigInt); unsafe { FfiBindData::<GsBindData>::set(info, GsBindData { total: n }) }; } // Init: zero-initialise the scan cursor unsafe extern "C" fn gs_init(info: duckdb_init_info) { unsafe { FfiInitData::<GsScanState>::set(info, GsScanState { pos: 0 }) }; } // Scan: emit a batch of rows using DataChunk wrapper unsafe extern "C" fn gs_scan(info: duckdb_function_info, output: duckdb_data_chunk) { let chunk = unsafe { DataChunk::from_raw(output) }; // Never unwrap in a callback: a missing state ends the scan instead. let bind = unsafe { FfiBindData::<GsBindData>::get_from_function(info) }; let state = unsafe { FfiInitData::<GsScanState>::get_mut(info) }; let (Some(bind), Some(state)) = (bind, state) else { unsafe { chunk.set_size(0) }; return; }; let remaining = bind.total - state.pos; let batch = remaining.min(2048).max(0) as usize; let mut writer = unsafe { chunk.writer(0) }; for i in 0..batch { unsafe { writer.write_i64(i, state.pos + i as i64) }; } unsafe { chunk.set_size(batch) }; state.pos += batch as i64; } std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), DuckDBSuccess); TableFunctionBuilder::new("generate_series_ext") .param(TypeId::BigInt) .bind(gs_bind) .init(gs_init) .scan(gs_scan) .register(con) .unwrap(); } let sum = |sql: &str| -> i64 { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); unsafe { chunk.reader(0).read_i64(0) } }; assert_eq!(sum("SELECT sum(value)::BIGINT FROM generate_series_ext(5)"), 10); assert_eq!(sum("SELECT count(*) FROM generate_series_ext(5000)"), 5000); }
Registration
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info}; use quack_rs::table::TableFunctionBuilder; use quack_rs::types::TypeId; unsafe extern "C" fn gs_bind(_: duckdb_bind_info) {} unsafe extern "C" fn gs_init(_: duckdb_init_info) {} unsafe extern "C" fn gs_scan(_: duckdb_function_info, _: duckdb_data_chunk) {} unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { TableFunctionBuilder::new("generate_series_ext") .param(TypeId::BigInt) .bind(gs_bind) .init(gs_init) .scan(gs_scan) .register(con)?; Ok(()) } }
Advanced features
Named parameters
Named parameters let callers pass optional arguments by name (e.g., step := 10):
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info}; use quack_rs::table::TableFunctionBuilder; use quack_rs::types::TypeId; unsafe extern "C" fn gs_v2_bind(_: duckdb_bind_info) {} unsafe extern "C" fn gs_v2_init(_: duckdb_init_info) {} unsafe extern "C" fn gs_v2_scan(_: duckdb_function_info, _: duckdb_data_chunk) {} unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { TableFunctionBuilder::new("gen_series_v2") .param(TypeId::BigInt) // positional: n .named_param("step", TypeId::BigInt) // named: step := <value> .bind(gs_v2_bind) .init(gs_v2_init) .scan(gs_v2_scan) .register(con)?; Ok(()) } }
In the bind callback, read the named parameter with
BindInfo::get_named_parameter_value("step"). Named parameters are optional: if the
query omits step := …, the returned Value wraps a null handle (is_null() is
true), so use a defaulting accessor such as as_i64_or(1).
Registering a name twice
The C API has no table function sets: a second registration under a name that
already exists — your own earlier one, another extension's, or a built-in such as
range — is dropped by DuckDB while duckdb_register_table_function still
reports success, and the old function keeps answering. register therefore checks
duckdb_functions() first and returns an error if the name already belongs to a
table function or table macro (compared case-insensitively). Table
functions registered through the C API live in the in-memory system catalog and
are never persisted, so reloading an extension into a database file never trips
this check.
Local init (per-thread state)
local_init allocates per-thread state for a scan that runs on several threads.
It does not make the scan parallel by itself — that is
InitInfo::set_max_threads (see Thread control):
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info}; use quack_rs::table::TableFunctionBuilder; use quack_rs::types::TypeId; unsafe extern "C" fn gs_v2_bind(_: duckdb_bind_info) {} unsafe extern "C" fn gs_v2_init(_: duckdb_init_info) {} unsafe extern "C" fn gs_v2_local_init(_: duckdb_init_info) {} unsafe extern "C" fn gs_v2_scan(_: duckdb_function_info, _: duckdb_data_chunk) {} unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { TableFunctionBuilder::new("gen_series_v2") .param(TypeId::BigInt) .bind(gs_v2_bind) .init(gs_v2_init) .local_init(gs_v2_local_init) // per-thread state allocation .scan(gs_v2_scan) .register(con)?; Ok(()) } }
The local init callback receives duckdb_init_info and can use
FfiLocalInitData::<T>::set to store per-thread state (T: Send).
Thread control
Use InitInfo::set_max_threads in the global init callback to tell DuckDB how
many threads can scan concurrently. The default is 1. Above 1, DuckDB calls the
scan from that many threads at the same time whether or not local_init is
set — and all of them share the one global init data and bind data. Do not use
FfiInitData::get_mut then; keep shared mutable state behind a Mutex or
atomics and read it with FfiInitData::get:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_init_info; use quack_rs::table::{FfiInitData, InitInfo}; struct MyState { pos: i64 } unsafe extern "C" fn gs_v2_init(info: duckdb_init_info) { let init_info = unsafe { InitInfo::new(info) }; init_info.set_max_threads(1); unsafe { FfiInitData::<MyState>::set(info, MyState { pos: 0 }) }; } }
Projection pushdown
Enable projection pushdown to let DuckDB skip unrequested columns:
#![allow(unused)] fn main() { use quack_rs::table::TableFunctionBuilder; fn demo() { let _ = TableFunctionBuilder::new("my_func") .projection_pushdown(true) // ... ; } }
Caution: When projection pushdown is enabled, your scan callback must check which columns DuckDB actually needs using
InitInfo::projected_column_countandInitInfo::projected_column_index. Writing to non-projected columns causes crashes.projected_column_indexreturnsNonepast the end of the projection (the C API itself answers0there, which is indistinguishable from the first column).
See examples/hello-ext/src/lib.rs for a complete example using named_param,
local_init, and set_max_threads.
Complex parameter types
For parameterised types that TypeId cannot express (e.g. LIST(BIGINT),
MAP(VARCHAR, INTEGER), STRUCT(...)), use param_logical and
named_param_logical:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info}; use quack_rs::table::TableFunctionBuilder; use quack_rs::types::TypeId; unsafe extern "C" fn bind_fn(_: duckdb_bind_info) {} unsafe extern "C" fn init_fn(_: duckdb_init_info) {} unsafe extern "C" fn scan_fn(_: duckdb_function_info, _: duckdb_data_chunk) {} unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { use quack_rs::types::LogicalType; TableFunctionBuilder::new("read_data") .param_logical(LogicalType::list(TypeId::Varchar)) // positional LIST param .named_param_logical("options", LogicalType::map( // named MAP param TypeId::Varchar, TypeId::Varchar, )) .bind(bind_fn) .init(init_fn) .scan(scan_fn) .register(con)?; Ok(()) } }
BindInfo helpers
BindInfo wraps duckdb_bind_info and exposes these methods:
| Method | Description |
|---|---|
add_result_column(name, TypeId) | Declares an output column (a type DuckDB would silently drop, like ANY, is a bind error instead) |
add_result_column_with_type(name, &LogicalType) | Output column with complex type (same check, including nested ANY/INVALID) |
set_cardinality(rows, is_exact) | Cardinality hint for the optimizer. DuckDB 1.5.5 treats is_exact = false as an estimate and an upper bound, true as an estimate only — the reverse of duckdb.h |
set_error(message) | Report a bind-time error (an empty message is replaced by a placeholder) |
parameter_count() | Number of positional parameters |
get_parameter_value(index) | Positional parameter as an RAII Value |
get_named_parameter_value(name) | Named parameter as an RAII Value; a null handle if the query omitted it |
get_parameter(index) / get_named_parameter(name) | The same as a raw duckdb_value, which the caller must destroy |
get_extra_info() | Returns the extra-info pointer set on the function |
get_client_context() | Returns a ClientContext (duckdb-1-5) |
result_column_count() / result_column_name(i) / result_column_type(i) | The target table's columns in a COPY … FROM reader (duckdb-1-5); zero columns otherwise |
InitInfo helpers
InitInfo wraps duckdb_init_info:
| Method | Description |
|---|---|
projected_column_count() | Number of projected columns (with pushdown) |
projected_column_index(idx) | Declared column index at projection position; None when idx is out of range |
set_max_threads(n) | Maximum concurrent scan threads (default 1; shared global state above 1) |
set_error(message) | Report an init-time error (an empty message is replaced by a placeholder) |
get_extra_info() | Returns the extra-info pointer set on the function |
FunctionInfo helpers
FunctionInfo wraps duckdb_function_info (scan callbacks):
| Method | Description |
|---|---|
set_error(message) | Report a scan-time error (an empty message is replaced by a placeholder) |
get_extra_info() | Returns the extra-info pointer set on the function |
Extra info
Use TableFunctionBuilder::extra_info to attach function-level data that is
accessible from all callbacks (bind, init, and scan) via get_extra_info(). The
pointee must be Send + Sync: DuckDB passes the same pointer to callbacks running
on several threads at once, and frees it on whichever thread releases the function.
Example output
SELECT * FROM generate_series_ext(5);
-- 0
-- 1
-- 2
-- 3
-- 4
SELECT value * value AS sq FROM generate_series_ext(4);
-- 0
-- 1
-- 4
-- 9
See also
tablemodule documentationreplacement_scan— for file-path-triggered table scanshello-extREADME
Replacement Scans
A DuckDB replacement scan lets users query a file by its path alone:
SELECT * FROM 'myfile.myformat'
When DuckDB finds no table with that name, it calls each registered replacement
scan in registration order, and one of them can redirect the query to a table
function (in the example below, read_myformat('myfile.myformat')). DuckDB's own
CSV, Parquet and JSON readers use the same mechanism. This page shows how to
register one from a Rust extension with quack-rs.
quack-rs provides ReplacementScanBuilder (a static registration helper) and
ReplacementScanInfo (an ergonomic wrapper for callbacks).
Registration API
Unlike the other builders in quack-rs, ReplacementScanBuilder uses a single
static call because the DuckDB C API takes all arguments at once:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_database, duckdb_replacement_scan_info}; use std::os::raw::{c_char, c_void}; unsafe extern "C" fn my_scan_callback(_: duckdb_replacement_scan_info, _: *const c_char, _: *mut c_void) {} fn demo(db: duckdb_database, my_state: String) { use quack_rs::replacement_scan::ReplacementScanBuilder; // Low-level: pass raw extra_data and an optional delete callback. unsafe { ReplacementScanBuilder::register( db, // duckdb_database my_scan_callback, // ReplacementScanFn std::ptr::null_mut(), // extra_data (or a raw pointer) None, // delete_callback ); } // Ergonomic: pass owned Rust data; boxing and destructor are handled for you. unsafe { ReplacementScanBuilder::register_with_data(db, my_scan_callback, my_state); } } }
Note: Replacement scans are registered on a database handle (
duckdb_database), not a connection, and apply to every connection to that database. In an entry point,Connection::register_replacement_scanandregister_replacement_scan_with_datado the same through theConnectionyou are given.
A raw delete_callback must accept a null argument: unlike DuckDB's other
destructor slots, it is called with extra_data even when that is null.
Callback signature
The raw callback receives duckdb_replacement_scan_info, but you can wrap it
with ReplacementScanInfo for ergonomic, safe access:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_replacement_scan_info; use quack_rs::replacement_scan::ReplacementScanInfo; unsafe extern "C" fn my_scan_callback( info: duckdb_replacement_scan_info, table_name: *const ::std::os::raw::c_char, _data: *mut ::std::os::raw::c_void, ) { let path = unsafe { std::ffi::CStr::from_ptr(table_name) } .to_str() .unwrap_or(""); if !path.ends_with(".myformat") { return; // not ours: DuckDB tries the next replacement scan } // Use ReplacementScanInfo for ergonomic access unsafe { ReplacementScanInfo::new(info) .set_function("read_myformat") .add_varchar_parameter(path); } } }
table_name is the last part of the table reference: FROM myschema."f.myformat"
reaches the callback as f.myformat, so a callback cannot see or honour a schema.
ReplacementScanInfo methods
| Method | Description |
|---|---|
set_function(name) | Redirect to the named table function |
add_varchar_parameter(value) | Add a VARCHAR parameter to the redirected call |
add_i64_parameter(value) | Add a BIGINT (i64) parameter |
add_bool_parameter(value) | Add a BOOLEAN parameter |
add_parameter_raw(duckdb_value) | Add a parameter of any type; DuckDB copies it, so the caller still destroys the value |
set_error(message) | Report an error, which fails the query. DuckDB ignores an empty message, so an empty one is replaced with a placeholder |
Data passed to ReplacementScanBuilder::register_with_data must be Send + Sync:
it lives in the database-wide configuration, is read by the callback from any
connection's thread (concurrently), and is dropped by whichever thread closes the
database. The raw register has the same requirement, stated in its # Safety.
When to use replacement scans vs table functions
| Scenario | Use |
|---|---|
SELECT * FROM my_function('file.ext') | Table function |
SELECT * FROM 'file.ext' (bare path) | Replacement scan → delegates to a table function |
| File type auto-detection | Replacement scan |
Most extensions implement both: a table function that does the actual work, and a replacement scan that detects the file extension and transparently routes bare-path queries to the table function.
See also
replacement_scanmodule documentation- Table Functions
Cast Functions
A DuckDB cast function defines how values of one type are converted to another.
This page shows how to register one from a Rust extension with quack-rs'
CastFunctionBuilder. Once registered, CAST(x AS T) and TRY_CAST(x AS T) use
your callback, and so do implicit conversions if you give the cast an implicit cost.
When to use cast functions
- Your extension introduces a new logical type and needs
CASTto/from standard types. - You want to override DuckDB's built-in cast behaviour for a specific type pair.
- You need to control implicit cast priority relative to other registered casts.
Registering a cast
#![allow(unused)] fn main() { use quack_rs::cast::{CastFunctionBuilder, CastFunctionInfo, CastMode}; use quack_rs::types::TypeId; use quack_rs::vector::{VectorReader, VectorWriter}; use libduckdb_sys::{duckdb_function_info, duckdb_vector, idx_t}; unsafe extern "C" fn varchar_to_int( info: duckdb_function_info, count: idx_t, input: duckdb_vector, output: duckdb_vector, ) -> bool { let cast_info = unsafe { CastFunctionInfo::new(info) }; let reader = unsafe { VectorReader::from_vector(input, count as usize) }; let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..count as usize { if !unsafe { reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } let s = unsafe { reader.read_str(row) }; match s.parse::<i32>() { Ok(v) => unsafe { writer.write_i32(row, v) }, Err(e) => { let msg = format!("cannot cast {:?} to INTEGER: {e}", s); if cast_info.cast_mode() == CastMode::Try { // TRY_CAST: record a per-row error, which also sets the row to NULL unsafe { cast_info.set_row_error(&msg, row as idx_t, output) }; } else { // Regular CAST: fail the whole query cast_info.set_error(&msg); return false; } } } } true } fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { unsafe { CastFunctionBuilder::new(TypeId::Varchar, TypeId::Integer) .function(varchar_to_int) .register(con) } } }
Implicit casts
Provide an implicit_cost to allow DuckDB to use the cast automatically in
expressions where the types do not match:
#![allow(unused)] fn main() { use quack_rs::cast::CastFunctionBuilder; use quack_rs::types::TypeId; use libduckdb_sys::{duckdb_function_info, duckdb_vector, idx_t}; unsafe extern "C" fn my_cast(_: duckdb_function_info, _: idx_t, _: duckdb_vector, _: duckdb_vector) -> bool { true } fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { unsafe { CastFunctionBuilder::new(TypeId::Varchar, TypeId::Integer) .function(my_cast) .implicit_cost(100) // lower = higher priority .register(con) } } }
Extra info
Attach arbitrary data to a cast function using extra_info. This is useful for
parameterising the cast behaviour (e.g., a rounding mode):
#![allow(unused)] fn main() { use quack_rs::cast::CastFunctionBuilder; use quack_rs::types::TypeId; use libduckdb_sys::{duckdb_function_info, duckdb_vector, idx_t}; use std::os::raw::c_void; unsafe extern "C" fn my_cast(_: duckdb_function_info, _: idx_t, _: duckdb_vector, _: duckdb_vector) -> bool { true } unsafe extern "C" fn my_destroy(_: *mut c_void) {} fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { let mode = Box::into_raw(Box::new("round".to_string())).cast::<c_void>(); unsafe { CastFunctionBuilder::new(TypeId::Double, TypeId::BigInt) .function(my_cast) .implicit_cost(100) .extra_info(mode, Some(my_destroy)) .register(con) } } }
Inside the cast callback, retrieve the extra info with
CastFunctionInfo::get_extra_info(). The pointee must be Send + Sync: DuckDB
passes the same pointer to the callback on every thread that runs the cast, and the
destructor runs on whichever thread releases it. CastFunctionBuilder is Send, so
an unregistered builder may also run the destructor on the thread that drops it.
If register returns an error — a missing callback or type, a bare composite
TypeId (use new_logical), a null connection, or a source or target type that is
or contains ANY/INVALID, which DuckDB refuses — the builder still owns the extra
info and runs its destructor exactly once. These cases are checked in Rust before
DuckDB is called, because duckdb_register_cast_function rejects them before
taking ownership of the pointer.
TRY_CAST vs CAST
Inside your callback, check CastFunctionInfo::cast_mode() to distinguish between
the two modes:
| Mode | User wrote | Expected behaviour on error |
|---|---|---|
CastMode::Normal | CAST(x AS T) | Call set_error and return false |
CastMode::Try | TRY_CAST(x AS T) | Call set_row_error for each failed row, continue |
set_row_error records the message and sets that row of the output to NULL
(FlatVector::SetNull inside the C API); row must be less than the callback's
count — DuckDB does not check it in release builds.
The return value does nothing in
TRY_CASTmode. DuckDB discards it (execute_cast.cppcalls the cast and ignores the result), so returningfalsedoes not turn the chunk intoNULLs: every row you did not null keeps whatever the output vector held — possibly a previous chunk's value. Null each failed row yourself. The one exception is a panic inside acast_callback!body: the macro then sets every row of the chunk toNULL, since nothing the body wrote can be trusted.
In Normal mode, set a message before returning false. A cast_callback! body
that sets none fails the query with "cast function failed without reporting an
error message"; a hand-written extern "C" callback that sets none makes DuckDB
report Conversion Error: followed by nothing. An empty message passed to
set_error / set_row_error is replaced by a placeholder.
Working example
The examples/hello-ext extension registers two cast functions:
CAST(VARCHAR AS INTEGER)/TRY_CAST(VARCHAR AS INTEGER)— basic castCAST(DOUBLE AS BIGINT)— withimplicit_cost(100)andextra_infofor rounding mode
See examples/hello-ext/src/lib.rs for complete, copy-paste-ready references.
Complex source and target types
For casts involving complex types like DECIMAL(18, 3) or LIST(VARCHAR), use
the new_logical constructor instead of new:
#![allow(unused)] fn main() { use quack_rs::cast::CastFunctionBuilder; use quack_rs::types::{LogicalType, TypeId}; use libduckdb_sys::{duckdb_function_info, duckdb_vector, idx_t}; unsafe extern "C" fn my_cast(_: duckdb_function_info, _: idx_t, _: duckdb_vector, _: duckdb_vector) -> bool { true } fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { unsafe { CastFunctionBuilder::new_logical( LogicalType::list(TypeId::Varchar), // LIST(VARCHAR) source LogicalType::list(TypeId::Integer), // LIST(INTEGER) target ) .function(my_cast) .register(con) } } }
The source() and target() accessor methods return Option<TypeId> — they
return None when the type was set via new_logical (since a LogicalType
cannot always be expressed as a simple TypeId).
API reference
CastFunctionBuilder— the builderCastFunctionInfo— the info handle inside callbacksCastMode—NormalorTrycast_callback!— generates a panic-safe cast callback
NULL Handling in Functions
This page explains how NULL inputs reach the scalar and aggregate functions of a DuckDB extension written in Rust, and how to give them SQL NULL semantics.
The one thing to take away: for a scalar function,
DefaultNullHandlingdoes not make DuckDB return NULL for you. Your callback is invoked for NULL rows too, and if it writes a value there, that value is the answer. CallDataChunk::propagate_nulls, or useScalarFunctionBuilder::map1/map2, which do it for you. An aggregate'supdatelikewise receives NULL rows under either setting.
What DuckDB actually does
DuckDB's FunctionNullHandling has two settings, and quack-rs mirrors them as
NullHandling. The names suggest that the default makes the
engine handle NULL propagation. For functions registered through the C API it
does not: a scalar function's output for a NULL row is kept as written, and an
aggregate's update is handed NULL rows (see
Aggregate functions). For scalars the result is silent
wrong answers rather than an error.
For scalar functions, two pieces of DuckDB source settle it (quoted from v1.5.4):
src/main/capi/scalar_function-c.cpp — the C API bridge calls your callback for
the whole flattened chunk and never looks at the result's validity:
void CAPIScalarFunction(DataChunk &input, ExpressionState &state, Vector &result) {
...
input.Flatten();
...
c_bind_info.info.function(c_function_info, c_input, c_result);
if (!function_info.success) {
throw InvalidInputException(function_info.error);
}
...
}
src/execution/expression_executor/execute_function.cpp — the only NULL check is
a debug-only assertion that your function already did the right thing:
static void VerifyNullHandling(const BoundFunctionExpression &expr, DataChunk &args, Vector &result) {
#ifdef DEBUG
if (args.data.empty() || expr.function.GetNullHandling() != FunctionNullHandling::DEFAULT_NULL_HANDLING) {
return;
}
// ... D_ASSERT(!result_data.validity.RowIsValid(idx));
#endif
}
Every DuckDB a user installs is a release build, so that assertion is compiled out.
Why the obvious test passes anyway
SELECT my_func(NULL); -- NULL, even for a broken function
A literal NULL is constant-folded during binding: DuckDB evaluates the
expression once, sees a NULL argument to a DEFAULT_NULL_HANDLING function, and
substitutes NULL without calling anything. The wrong answers only appear when the
argument comes from a column:
CREATE TABLE t(i BIGINT);
INSERT INTO t VALUES (1), (NULL), (3);
SELECT i, my_func(i) FROM t;
Against DuckDB 1.5.4, a function that unconditionally writes 999 returns:
| i | my_func(i) |
|---|---|
| 1 | 999 |
| NULL | 999 ← not NULL |
| 3 | 999 |
quack-rs pins this behaviour in tests/ffi_roundtrip.rs, so a future DuckDB that
starts propagating will show up as a test failure rather than as a surprise.
Doing it right
The safe route: typed scalar functions
ScalarFunctionBuilder::map1 / map2 take an ordinary Rust closure and handle
validity for you — a NULL argument short-circuits to a NULL result without ever
calling your code:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; fn live_connection() -> libduckdb_sys::duckdb_connection { std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess); } con } /// First column of the first row, as BIGINT; `None` for NULL. fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); let reader = unsafe { chunk.reader(0) }; unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) } } let con = live_connection(); let run = || -> Result<(), ExtensionError> { unsafe { ScalarFunctionBuilder::map1("double_it", |x: i64| x * 2)? .register(con)?; } Ok(()) }; run().unwrap(); assert_eq!(query_i64(con, "SELECT double_it(21)"), Some(42)); // From a column, not a literal: `double_it(NULL::BIGINT)` is constant-folded // to NULL whatever the function does, so it cannot tell a broken one apart. assert_eq!(query_i64(con, "SELECT double_it(i) FROM (VALUES (NULL::BIGINT)) t(i)"), None); }
Use map1_opt / map2_opt when the function needs to see NULLs; those
register SpecialNullHandling for you and hand the closure Option<T>.
The raw route: propagate_nulls
When you write the extern "C" callback yourself, restore SQL semantics with one
call at the end:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; fn live_connection() -> libduckdb_sys::duckdb_connection { std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess); } con } /// First column of the first row, as BIGINT; `None` for NULL. fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); let reader = unsafe { chunk.reader(0) }; unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) } } quack_rs::scalar_callback!(double_it, |_info, input, output| { let chunk = unsafe { DataChunk::from_raw(input) }; let reader = unsafe { chunk.reader(0) }; let mut writer = unsafe { VectorWriter::from_vector(output) }; for row in 0..chunk.size() { unsafe { writer.write_i64(row, reader.read_i64(row) * 2) }; } // Without this, double_it(i) for a NULL `i` from a column is 0, not NULL. unsafe { chunk.propagate_nulls(&mut writer) }; }); let con = live_connection(); unsafe { ScalarFunctionBuilder::new("double_it").param(TypeId::BigInt).returns(TypeId::BigInt) .function(double_it).register(con).unwrap(); } assert_eq!(query_i64(con, "SELECT double_it(21)"), Some(42)); // From a column, not a literal: `double_it(NULL::BIGINT)` is constant-folded // to NULL whatever the function does, so it cannot tell a broken one apart. assert_eq!(query_i64(con, "SELECT double_it(i) FROM (VALUES (NULL::BIGINT)) t(i)"), None); }
propagate_nulls resolves each column's validity pointer once and marks the
output NULL wherever any input column is NULL. A column with no validity mask has
no NULLs and costs nothing. DataChunk::any_null(row) is the per-row form when
you need the decision inline.
NullHandling enum
#![allow(unused)] fn main() { use quack_rs::types::NullHandling; // Default: the function promises NULL in -> NULL out. // Scalar: you must keep that promise (see above). // Aggregate: `update` still receives NULL rows; skip them yourself. NullHandling::DefaultNullHandling; // The function means to see NULLs and may return non-NULL for them. NullHandling::SpecialNullHandling; }
Aggregate functions
Aggregates behave like scalar functions here: under either setting,
update receives every row of the chunk, NULL rows included. CAPIAggregateUpdate
in DuckDB's aggregate_function-c.cpp flattens the inputs and passes the whole
chunk through; nothing on the way filters by validity. An aggregate that ignores
NULLs skips them itself:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe fn demo(chunk: &DataChunk, reader: &VectorReader) { for row in 0..chunk.size() { if !unsafe { reader.is_valid(row) } { continue; // a NULL row: its data slot holds no meaningful value } // ... accumulate reader.read_i64(row) into *states.add(row) ... } } }
SpecialNullHandling declares that the aggregate may return non-NULL for NULL
input (a count_with_nulls, say). For an aggregate DuckDB reads the setting in
one place only — the correlated-subquery decorrelator, to pick an INNER or
LEFT join — and no query we tried (correlated scalar subqueries, with and
without arithmetic or coalesce around the aggregate, LATERAL, a correlated
subquery in WHERE) answered differently under the two settings on DuckDB 1.5.5.
Set it anyway when it is true; it is what DuckDB expects.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct CountState { count: i64 } impl AggregateState for CountState {} unsafe extern "C" fn my_update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {} unsafe extern "C" fn my_combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {} unsafe extern "C" fn my_finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {} unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { use quack_rs::aggregate::AggregateFunctionBuilder; use quack_rs::types::{TypeId, NullHandling}; unsafe { AggregateFunctionBuilder::new("count_with_nulls") .param(TypeId::BigInt) .returns(TypeId::BigInt) .null_handling(NullHandling::SpecialNullHandling) .ffi_state::<CountState>() .update(my_update) // counts rows whose value is NULL, too .combine(my_combine) .finalize(my_finalize) .register(con)?; } Ok(()) } }
Empty groups in a correlated subquery
One difference from an uncorrelated query holds under both settings. In
SELECT (SELECT my_count(x) FROM t2 WHERE t2.k = t1.k) FROM t1, an outer row
with no matching t2 rows gets NULL: the decorrelated plan joins the aggregate's
groups back to the outer rows, and an outer row with no group never has an empty
state finalized. DuckDB rewrites that NULL to 0 for its own count and
count(*) only. A count-like aggregate of yours that returns 0 for empty input
therefore returns NULL here; write coalesce((SELECT ...), 0) if the query needs
0. (SELECT my_count(x) FROM t2 WHERE false, uncorrelated, does finalize an
empty state and returns 0.)
When to use special NULL handling
| Use case | NULL handling | Who propagates |
|---|---|---|
| Scalar function, NULL in → NULL out | DefaultNullHandling | you (propagate_nulls, or map1/map2) |
Scalar function that inspects NULLs (COALESCE-like, IS_NULL-like) | SpecialNullHandling | you |
| Aggregate, ignore NULL rows | DefaultNullHandling (the default) | you (skip rows where is_valid is false) |
| Aggregate that counts NULLs | SpecialNullHandling | you |
If you don't call .null_handling(), DefaultNullHandling is used.
SQL Macros
A DuckDB SQL macro packages a SQL expression or query as a named function, with no
FFI callback behind it. This page shows how to create scalar and table macros from a
Rust extension with quack-rs' SqlMacro: you give the macro body as a string and call
.register(con), which runs CREATE OR REPLACE MACRO.
Two macro types
| Type | SQL generated | Returns |
|---|---|---|
| Scalar | CREATE OR REPLACE MACRO "name"("a", "b") AS (expression) | one value per row |
| Table | CREATE OR REPLACE MACRO "name"("a", "b") AS TABLE query | a result set |
Scalar macros
A scalar macro wraps a SQL expression. Think of it as a parameterized SQL alias:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_connection, DuckDBSuccess}; use quack_rs::error::ExtensionError; fn live_connection() -> duckdb_connection { std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), DuckDBSuccess); } con } fn query_i64(con: duckdb_connection, sql: &str) -> i64 { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); unsafe { chunk.reader(0).read_i64(0) } } use quack_rs::sql_macro::SqlMacro; fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { // clamp(x, lo, hi) → greatest(lo, least(hi, x)) SqlMacro::scalar("clamp", &["x", "lo", "hi"], "greatest(lo, least(hi, x))")? .register(con)?; // golden_ratio() → 1.61803398874989 SqlMacro::scalar("golden_ratio", &[], "1.61803398874989")? .register(con)?; // safe_div(a, b) → CASE WHEN b = 0 THEN NULL ELSE a / b END SqlMacro::scalar( "safe_div", &["a", "b"], "CASE WHEN b = 0 THEN NULL ELSE a / b END", )? .register(con)?; } Ok(()) } let con = live_connection(); register(con).unwrap(); assert_eq!(query_i64(con, "SELECT clamp(9, 1, 5)::BIGINT"), 5); assert_eq!(query_i64(con, "SELECT clamp(-3, 1, 5)::BIGINT"), 1); assert_eq!(query_i64(con, "SELECT (golden_ratio() * 1000)::BIGINT"), 1618); assert_eq!(query_i64(con, "SELECT count(*) FROM (SELECT safe_div(1, 0) AS v) WHERE v IS NULL"), 1); }
Use in DuckDB:
SELECT clamp(rating, 1, 5) FROM reviews;
SELECT safe_div(revenue, orders) FROM monthly_stats;
Table macros
A table macro wraps a SQL query that returns rows:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_connection, DuckDBSuccess}; use quack_rs::error::ExtensionError; use quack_rs::sql_macro::SqlMacro; fn live_connection() -> duckdb_connection { std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), DuckDBSuccess); } con } fn query_i64(con: duckdb_connection, sql: &str) -> i64 { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); unsafe { chunk.reader(0).read_i64(0) } } fn register(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { // active_users(tbl) → SELECT * FROM query_table(tbl) WHERE active = true SqlMacro::table( "active_users", &["tbl"], "SELECT * FROM query_table(tbl) WHERE active = true", )? .register(con)?; // recent_orders(days) → last N days of orders SqlMacro::table( "recent_orders", &["days"], "SELECT * FROM orders WHERE order_date >= current_date - INTERVAL (days) DAY", )? .register(con)?; } Ok(()) } let con = live_connection(); unsafe { quack_rs::query::execute(con, "CREATE TABLE users AS SELECT * FROM (VALUES (1, true), (2, false), (3, true)) t(id, active)") }.unwrap(); unsafe { quack_rs::query::execute(con, "CREATE TABLE orders AS SELECT current_date - 3 AS order_date UNION ALL SELECT current_date - 30") }.unwrap(); register(con).unwrap(); assert_eq!(query_i64(con, "SELECT count(*) FROM recent_orders(7)"), 1); assert_eq!(query_i64(con, "SELECT count(*) FROM active_users(users)"), 2); }
A table macro's body is bound when the macro is created, so every table it names
directly (orders above) must already exist when register runs, or registration
fails with a catalog error. To take a table as a parameter, read it through
query_table(tbl): a bare FROM tbl looks for a table literally named tbl.
Use in DuckDB:
SELECT * FROM active_users(users);
SELECT count(*) FROM recent_orders(7);
Inspecting the generated SQL
to_sql() returns the CREATE OR REPLACE MACRO statement without requiring a live connection.
Use it for logging, debugging, or assertions in tests:
#![allow(unused)] fn main() { use quack_rs::sql_macro::SqlMacro; let m = SqlMacro::scalar("add", &["a", "b"], "a + b")?; assert_eq!( m.to_sql(), r#"CREATE OR REPLACE MACRO "add"("a", "b") AS (a + b)"# ); let t = SqlMacro::table("active_users", &["tbl"], "SELECT * FROM query_table(tbl) WHERE active = true")?; assert_eq!( t.to_sql(), r#"CREATE OR REPLACE MACRO "active_users"("tbl") AS TABLE SELECT * FROM query_table(tbl) WHERE active = true"# ); Ok::<(), quack_rs::error::ExtensionError>(()) }
Name and parameter validation
Macro names are validated with
validate_function_name,
the same rules as function names, and parameter names with
validate_parameter_name.
A name must:
- start with an ASCII letter or underscore, followed by ASCII letters, digits or underscores;
- be at most 256 characters long;
- not be a DuckDB keyword that cannot be used in that position unquoted. For a
macro, that is a keyword it cannot be called by (
order,coalesce); for a parameter, one the body cannot refer to it by (order,left). The two lists differ: a parameter may be calledcolumns, which a macro may not.
Case is not restricted — DuckDB identifiers are case-insensitive, so a macro
registered as MyMacro is callable as mymacro(...) or MYMACRO(...).
#![allow(unused)] fn main() { use quack_rs::sql_macro::SqlMacro; assert!(SqlMacro::scalar("1f", &[], "1").is_err()); // ❌ starts with a digit assert!(SqlMacro::scalar("my-macro", &[], "1").is_err()); // ❌ hyphen assert!(SqlMacro::scalar("f", &["a b"], "1").is_err()); // ❌ space in param assert!(SqlMacro::scalar("f", &["order"], "1").is_err()); // ❌ reserved keyword as param assert!(SqlMacro::scalar("f", &["left"], "1").is_err()); // ❌ a body cannot refer to it assert!(SqlMacro::scalar("columns", &[], "1").is_err()); // ❌ cannot be called as a macro assert!(SqlMacro::scalar("f", &["columns"], "1").is_ok()); // ✅ fine as a parameter assert!(SqlMacro::scalar("MyMacro", &[], "1").is_ok()); // ✅ mixed case allowed assert!(SqlMacro::scalar("f", &["X"], "1").is_ok()); // ✅ mixed-case param allowed assert!(SqlMacro::scalar("f", &["_x"], "1").is_ok()); // ✅ underscore prefix allowed }
SQL injection safety
Macro and parameter names are restricted to ASCII letters, digits and underscores,
preventing SQL injection at the identifier level. to_sql() additionally emits every
name as a double-quoted identifier ("name") — the validated character set cannot
contain ", so no escaping is needed. Quoting does not make names case-sensitive
in DuckDB.
The body (expression or query) is your own extension code — it is included verbatim.
Never build macro bodies from untrusted user input. register does refuse a body
that turns the statement into several (1); DROP TABLE t; SELECT (1) — it counts
statements with DuckDB's own parser (duckdb_extract_statements) and executes
nothing if there is more than one — but a body can still change the meaning of the
single statement it is part of.
A scalar body containing -- gets a newline before the closing parenthesis, so a
trailing line comment ("x + 1 -- plus one") does not comment it out.
Where a macro lives
A macro is not a function registration: register runs CREATE OR REPLACE MACRO,
so the macro is an ordinary catalog object in the connection's default database and
schema — the user's database:
- It persists. In a database file the macro is still there in the next session,
even if the extension is never loaded again. Reloading is fine:
OR REPLACEreplaces it. - It needs a writable database. On a read-only database the
CREATEfails; an entry point that propagates that error with?makesLOADfail. Decide whether a macro is essential there. - It silently replaces a user's macro of the same name. Prefix macro names with your extension's name.
- It can shadow a built-in. A macro named
absin the default schema is found before the built-inabs, soabs(-1)calls the macro.
How it works under the hood
SqlMacro::register counts the statements in to_sql() with
duckdb_extract_statements, refuses more than one, and executes the
CREATE OR REPLACE MACRO statement via duckdb_query.
The query result is zero-initialized before duckdb_query, any error message is read
with duckdb_result_error, and duckdb_destroy_result runs on success and failure
alike.
Choosing between macros and scalar functions
| Scenario | Use |
|---|---|
| Logic expressible in SQL | SQL macro — simpler, no FFI |
| Logic needs Rust code (algorithms, external crates, etc.) | Scalar function |
| Simple expressions | SQL macro (expanded inline, no callback per chunk) |
| Type-specific overloads | Scalar function set (ScalarFunctionSetBuilder) |
| Returning a table | SQL table macro |
Copy Functions
A DuckDB copy function implements a custom file format for the COPY statement.
This page shows how to register one from a Rust extension with quack-rs'
CopyFunctionBuilder.
Requires the
duckdb-1-5feature flag (DuckDB 1.5.0+).
A format can support writing, reading, or both:
| Direction | You supply | DuckDB calls |
|---|---|---|
COPY t TO 'f' (FORMAT my_format) | bind + sink + finalize (and optionally global_init) | those callbacks |
COPY t FROM 'f' (FORMAT my_format) | copy_from(table_function) | your table function's bind, init and scan |
duckdb_register_copy_function decides which directions a format supports by
looking at the sink and the reader independently, so a read-only format leaves
the writing callbacks unset entirely. Set bind, sink and finalize together
or not at all: register refuses a builder with only some of them, and one that
implements neither direction.
Lifecycle (COPY … TO)
- Bind — called when the statement is bound: once for a plain
COPY, and again on everyEXECUTEof a prepared one. Inspect output columns, configure the export. - Global init — called once per output file: once for a plain
COPY, once per file withPER_THREAD_OUTPUTorPARTITION_BY. Open the file, allocate that file's global state. WithUSE_TMP_FILEthe path is a temporary name that DuckDB renames afterwards. - Sink — called for each data chunk; with
PER_THREAD_OUTPUTorPARTITION_BY, from several threads at once (see Threads). Write rows to the output. - Finalize — called once per output file, after its last sink. Flush buffers, close the file. It is not called when a sink reports an error, so release resources in the global state's destructor as well.
Builder API
#![allow(unused)] fn main() { use quack_rs::copy_function::CopyFunctionBuilder; fn demo( my_bind_fn: quack_rs::copy_function::CopyBindFn, my_global_init_fn: quack_rs::copy_function::CopyGlobalInitFn, my_sink_fn: quack_rs::copy_function::CopySinkFn, my_finalize_fn: quack_rs::copy_function::CopyFinalizeFn, ) -> Result<(), quack_rs::error::ExtensionError> { let builder = CopyFunctionBuilder::try_new("my_format")? .bind(my_bind_fn) .global_init(my_global_init_fn) .sink(my_sink_fn) .finalize(my_finalize_fn); // Register on a connection, for example in the entry point: // unsafe { builder.register(con)?; } Ok(()) } }
COPY … FROM
Reading is a table function, attached to the copy function rather than
registered on its own. Build it with
TableFunctionBuilder::build_handle,
then hand it to copy_from:
#![allow(unused)] fn main() { use quack_rs::copy_function::CopyFunctionBuilder; use quack_rs::table::TableFunctionBuilder; use quack_rs::types::TypeId; fn demo( bind: quack_rs::table::BindFn, init: quack_rs::table::InitFn, scan: quack_rs::table::ScanFn, ) -> Result<(), quack_rs::error::ExtensionError> { // SAFETY: the callbacks match their declared signatures. let reader = unsafe { TableFunctionBuilder::new("my_format") // name it after the format .param(TypeId::Varchar) // the file path — exactly one .named_param("skip_rows", TypeId::BigInt) // a COPY option .bind(bind) .init(init) .scan(scan) .build_handle() }?; let format = CopyFunctionBuilder::try_new("my_format")?.copy_from(reader)?; // unsafe { format.register(con)?; } Ok(()) } }
Four things about the reader are not like an ordinary table function:
- The file path is positional parameter 0, always
VARCHAR.duckdb.hrequires the function to declare exactly that one parameter, and DuckDB does not check it —copy_fromdoes, and returns an error naming the mismatch. - COPY options are named parameters.
(FORMAT my_format, SKIP_ROWS 1)arrives asskip_rows; matching is case-insensitive. An option the function never declared is a binder error before your bind callback runs — and the message names the table function, which is why the example gives it the format's name. - Those options arrive uncast. DuckDB passes each value as written, not
cast to the declared type:
SKIP_ROWS 'abc'reaches aBIGINTparameter as theVARCHAR'abc', andSKIP_ROWS 3as anINTEGER. An option written without a value ((FORMAT my_format, HEADER)) is not passed at all. CheckValue::type_id()before trusting a value —as_i64_or(0)on'abc'quietly returns the default. - The schema is already fixed, because
COPY … FROMloads into an existing table. The bind callback must not calladd_result_column(a typed reader built withwith_stateorwith_bind_initfails its bind if it does). Read the target's schema instead:
#![allow(unused)] fn main() { use quack_rs::table::BindInfo; fn demo(bind: &BindInfo) { for i in 0..bind.result_column_count() { let name = bind.result_column_name(i); // SAFETY: `i` is in range, and this runs during the bind callback. let ty = unsafe { bind.result_column_type(i) }; let _ = (name, ty); } } }
Registering a format name twice
duckdb_register_copy_function drops a copy function whose name already exists —
your earlier registration, another extension's, or a built-in format such as csv
— and still reports success. register checks first and returns an error. There is
no catalog view of copy functions, so the check asks the binder: it runs
COPY (SELECT <missing column>) TO '' (FORMAT '<name>'), where a binder error
means the format resolved and a catalog error means it did not. Nothing is written
and no copy callback runs; a format name owned by an autoloadable extension
(parquet, json) may be autoloaded by that lookup, exactly as the same COPY
typed by a user would. Copy functions registered through the C API are never
persisted, so reloading into a database file does not trip the check.
Options
CopyBindInfo::options() returns the COPY … TO options as one STRUCT value
(None only if DuckDB returns a null handle). How DuckDB 1.5.5 builds it:
- Option names are upper-cased:
compression 'zstd'arrives asCOMPRESSION.FORMATitself is not among them. - With no options besides
FORMAT, the value is SQLNULL, not an emptySTRUCT— checkis_sql_null()first. - An option given without a value (
HEADER) is aNULLfield. - Several values (
LST (1, 2)) arrive as aLIST, or an unnamedSTRUCTwhen their types differ. - An explicit
NULLvalue is rejected by the binder before your callback runs. - Field order follows DuckDB's internal hash map, not the statement — look fields up by name:
#![allow(unused)] fn main() { use quack_rs::copy_function::CopyBindInfo; fn demo(bind: &CopyBindInfo) -> Option<String> { let options = bind.options()?; if options.is_sql_null() { return None; // COPY ... (FORMAT my_format) with no other options } let names = options.struct_field_names(); let idx = names.iter().position(|n| n == "COMPRESSION")?; options.struct_child(idx)?.as_str().ok() } }
Threads
The data pointers are untyped, so the compiler cannot check this for you:
extra_infolives as long as the database and is read from every connection's thread — treat it asT: Send + Sync.- Bind data is shared by every sink call.
COPY … TOwithPER_THREAD_OUTPUTorPARTITION_BYruns the sink on several threads at once — treat it asT: Send + Syncand never mutate it without a lock. - Global state is one per output file: per thread with
PER_THREAD_OUTPUT, per partition withPARTITION_BY(reachable from more than one thread). Guard any mutation with aMutexunless you know neither option is in use.
Callback signatures
| Phase | Signature |
|---|---|
| Bind | unsafe extern "C" fn(info: duckdb_copy_function_bind_info) |
| Global init | unsafe extern "C" fn(info: duckdb_copy_function_global_init_info) |
| Sink | unsafe extern "C" fn(info: duckdb_copy_function_sink_info, chunk: duckdb_data_chunk) |
| Finalize | unsafe extern "C" fn(info: duckdb_copy_function_finalize_info) |
Callback info wrappers
Each phase provides an ergonomic wrapper type around its raw info handle. Wrap the handle at the top of your callback to access helper methods:
CopyBindInfo
| Method | Description |
|---|---|
column_count() | Number of output columns |
column_type(index) | LogicalType of the column at index, or None if out of range |
options() | The COPY … TO options, as one STRUCT Value (see Options) |
get_extra_info() | Extra-info pointer set on the copy function |
set_bind_data(data, destroy) | Store bind data and its destructor |
set_error(message) | Report a bind-time error |
get_client_context() | Returns a ClientContext for catalog/config access |
CopyGlobalInitInfo
| Method | Description |
|---|---|
get_bind_data() | Retrieve the bind data pointer |
get_extra_info() | Extra-info pointer set on the copy function |
get_file_path() | Output file path for the COPY operation; an error if it is not valid UTF-8 |
get_file_path_bytes() | The same path as its exact bytes, for a path that is not UTF-8 |
set_global_state(state, destroy) | Store global state and its destructor |
set_error(message) | Report an init-time error |
get_client_context() | Returns a ClientContext |
CopySinkInfo
| Method | Description |
|---|---|
get_bind_data() | Retrieve the bind data pointer |
get_extra_info() | Extra-info pointer set on the copy function |
get_global_state() | Retrieve the global state pointer |
set_error(message) | Report a sink-time error |
get_client_context() | Returns a ClientContext |
CopyFinalizeInfo
| Method | Description |
|---|---|
get_bind_data() | Retrieve the bind data pointer |
get_extra_info() | Extra-info pointer set on the copy function |
get_global_state() | Retrieve the global state pointer |
set_error(message) | Report a finalize-time error |
get_client_context() | Returns a ClientContext |
All four wrappers are re-exported from quack_rs::copy_function:
#![allow(unused)] fn main() { use quack_rs::copy_function::{CopyBindInfo, CopyGlobalInitInfo, CopySinkInfo, CopyFinalizeInfo}; }
Related modules
config_option— register custom settings for your formatclient_context— access the file system and catalog from callbackstable_description— inspect table metadatacatalog— look up catalog entriestable— build the table function aCOPY … FROMneeds
Reading & Writing Vectors
DuckDB passes data to and from your extension as vectors: columnar arrays of typed
values, each with a separate validity (NULL) bitmap. VectorReader and VectorWriter
give typed access to these vectors from scalar, aggregate and table function callbacks;
DataChunk, StructReader, StructWriter and ChunkWriter build on them.
VectorReader
Construction
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(input: duckdb_data_chunk, column_index: usize) { // In a scalar function callback: let reader = unsafe { VectorReader::new(input, column_index) }; // In an aggregate update callback: let reader = unsafe { VectorReader::new(input, 0) }; // first column } }
VectorReader::new takes the duckdb_data_chunk and a zero-based column index. The
reader holds raw pointers into the chunk, so it must not outlive the callback.
Row count
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader) { let n = reader.row_count(); // number of rows in this chunk } }
Chunk sizes vary. Always loop over 0..reader.row_count(); never assume a fixed size.
NULL check
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, writer: &mut VectorWriter) { for row in 0..reader.row_count() { if unsafe { !reader.is_valid(row) } { // row is NULL — skip or propagate NULL to output unsafe { writer.set_null(row) }; continue; } } } }
Always check is_valid before reading. Reading a fixed-width value from a
NULL row returns garbage data; reading a VARCHAR or BLOB from one can
follow a stale pointer into freed memory (see
NULL Handling & Strings).
Reading values
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, row: usize) { let i: i8 = unsafe { reader.read_i8(row) }; let i: i16 = unsafe { reader.read_i16(row) }; let i: i32 = unsafe { reader.read_i32(row) }; let i: i64 = unsafe { reader.read_i64(row) }; let u: u8 = unsafe { reader.read_u8(row) }; let u: u16 = unsafe { reader.read_u16(row) }; let u: u32 = unsafe { reader.read_u32(row) }; let u: u64 = unsafe { reader.read_u64(row) }; let f: f32 = unsafe { reader.read_f32(row) }; let f: f64 = unsafe { reader.read_f64(row) }; let b: bool = unsafe { reader.read_bool(row) }; // safe: uses u8 != 0 let s: &str = unsafe { reader.read_str(row) }; // handles inline + pointer format let iv = unsafe { reader.read_interval(row) }; // returns DuckInterval // Temporal and binary types (v0.10.0+): let d: i32 = unsafe { reader.read_date(row) }; // days since epoch let ts: i64 = unsafe { reader.read_timestamp(row) }; // microseconds since epoch let t: i64 = unsafe { reader.read_time(row) }; // microseconds since midnight let blob: &[u8] = unsafe { reader.read_blob(row) }; // binary data let uuid: u128 = unsafe { reader.read_uuid(row) }; // UUID's textual 128 bits } }
VectorWriter
Construction
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(output: duckdb_vector, result: duckdb_vector) { // In a scalar function callback: let mut writer = unsafe { VectorWriter::new(output) }; // In an aggregate finalize callback: let mut writer = unsafe { VectorWriter::new(result) }; } }
Writing values
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; use quack_rs::interval::DuckInterval; fn demo(writer: &mut VectorWriter, row: usize, s: &str, interval: DuckInterval, days_since_epoch: i32, micros_since_epoch: i64, micros_since_midnight: i64, bytes: Vec<u8>, uuid_bits: u128) { unsafe { writer.write_i8(row, -8) }; unsafe { writer.write_i16(row, -16) }; unsafe { writer.write_i32(row, -32) }; unsafe { writer.write_i64(row, -64) }; unsafe { writer.write_u8(row, 8) }; unsafe { writer.write_u16(row, 16) }; unsafe { writer.write_u32(row, 32) }; unsafe { writer.write_u64(row, 64) }; unsafe { writer.write_f32(row, 3.5) }; unsafe { writer.write_f64(row, 2.5) }; unsafe { writer.write_bool(row, true) }; unsafe { writer.write_varchar(row, s) }; // &str unsafe { writer.write_str(row, s) }; // alias for write_varchar unsafe { writer.write_interval(row, interval) }; // DuckInterval // Temporal and binary types (v0.10.0+): unsafe { writer.write_date(row, days_since_epoch) }; unsafe { writer.write_timestamp(row, micros_since_epoch) }; unsafe { writer.write_time(row, micros_since_midnight) }; unsafe { writer.write_blob(row, &bytes) }; unsafe { writer.write_uuid(row, uuid_bits) }; // UUID's textual 128 bits } }
write_varchar and write_blob panic for a value longer than
vector::string::MAX_STRING_LEN (u32::MAX bytes, DuckDB's string length
limit) rather than store a truncated one. Inside scalar_callback! and the
typed scalar constructors the panic becomes a SQL error. To handle the error
yourself, use try_write_varchar / try_write_blob, which return
Result<(), ExtensionError> and write nothing on error.
UUID is not stored as you'd expect
A UUID column is physically a HUGEINT, but the 128 bits in the vector are
not the bits you see in the text form: DuckDB flips the top bit so that
comparing the signed integers orders UUIDs the same way comparing their strings
does.
SELECT '11111111-2222-3333-4444-555555555555'::UUID
read_i128 (raw storage) : 0x91111111222233334444555555555555
read_uuid (textual bits): 0x11111111222233334444555555555555
read_uuid / write_uuid apply the flip for you and speak in textual bits
(u128) — the same convention as Value::uuid / Value::as_uuid and every
Rust Uuid type. Reach for read_i128 / write_i128 only when you want the raw
storage, and use quack_rs::vector::{uuid_from_storage, uuid_to_storage} to
convert.
Writing NULL
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(writer: &mut VectorWriter, row: usize) { unsafe { writer.set_null(row) }; } }
Pitfall L4:
set_nullcallsduckdb_vector_ensure_validity_writableautomatically beforeduckdb_vector_get_validity. A vector with no NULLs yet usually has no validity mask, so without that callget_validityreturns NULL andduckdb_validity_set_row_invalidsilently does nothing — the row you meant to be NULL reads back as a valid value.VectorWriter::set_nullhandles this correctly. See Pitfall L4.
NULL rows of STRUCT and ARRAY outputs
For a STRUCT output, set_null(row) (and set_null_range, and
DataChunk::propagate_nulls, which uses it) also nulls that row in every
field, recursively; for an ARRAY of size n it nulls child rows
row * n .. row * n + n. This mirrors DuckDB's internal FlatVector::SetNull,
and it matters: struct_extract / s.a reads the field vector without looking
at the parent, so a NULL struct row whose fields were left valid returns the
stale field value. StructWriter::set_row_null(row) does the same from a
StructWriter. LIST / MAP elements are not touched (as in DuckDB).
To reuse such a row, call set_valid(row) first: on a row that is NULL it also
marks valid everything set_null nulled below it. Then write the fields, and
any field NULLs after that. On a row that is already valid, set_valid leaves
the fields alone, so marking a row valid after writing its fields keeps their
NULLs.
Clearing NULL (v0.11.0+)
To undo a previous set_null call and mark a row as valid again:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(writer: &mut VectorWriter, row: usize) { unsafe { writer.set_valid(row) }; } }
Like set_null, set_valid calls ensure_validity_writable first.
DataChunk
DataChunk wraps a duckdb_data_chunk handle and gives access to its vectors
and row count without raw FFI calls:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::data_chunk::DataChunk; unsafe extern "C" fn my_scan(info: duckdb_function_info, output: duckdb_data_chunk) { let chunk = unsafe { DataChunk::from_raw(output) }; let mut writer = unsafe { chunk.writer(0) }; // VectorWriter for column 0 unsafe { writer.write_i64(0, 42) }; unsafe { chunk.set_size(1) }; // set output row count } }
Methods:
size()— current row countset_size(n)— set row count (0 = end of stream)column_count()— number of columnsvector(col)— rawduckdb_vectorhandlewriter(col)—VectorWriterfor a columnreader(col)—VectorReaderfor a columnstruct_writer(col, field_count)—StructWriterfor a STRUCT output columnstruct_reader(col, field_count)—StructReaderfor a STRUCT input columnstruct_field_reader(col, field)—VectorReaderfor a specific STRUCT fieldany_null(row)— whether any column is NULL atrowpropagate_nulls(&mut writer)— mark each output row NULL where any input column is NULLinto_chunk_writer()— convert toChunkWriter, which callsset_sizeon drop
StructWriter / StructReader
For STRUCT columns, creating a VectorWriter or VectorReader for each field by
hand is verbose. StructWriter and StructReader create one per field at
construction:
#![allow(unused)] fn main() { use quack_rs::data_chunk::DataChunk; struct Output { success: bool, data: String, count: i64, day: i32, payload: Vec<u8> } fn demo(chunk: &DataChunk, row: usize, result: &Output) { // Writing a 5-field STRUCT output: let mut sw = unsafe { chunk.struct_writer(0, 5) }; unsafe { sw.write_bool(row, 0, result.success); sw.write_varchar(row, 1, &result.data); sw.write_i64(row, 2, result.count); sw.write_date(row, 3, result.day); sw.write_blob(row, 4, &result.payload); } // Reading a 3-field STRUCT input: let sr = unsafe { chunk.struct_reader(0, 3) }; for row in 0..chunk.size() { let name = unsafe { sr.read_str(row, 0) }; let age = unsafe { sr.read_i32(row, 1) }; let active = unsafe { sr.read_bool(row, 2) }; } } }
ChunkWriter
ChunkWriter wraps an output duckdb_data_chunk and counts the rows handed out
by next_row. It calls set_size with that count on drop, so the row count
cannot be forgotten or set wrongly. next_row returns None once the chunk
holds duckdb_vector_size() rows:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::data_chunk::DataChunk; struct Item { name: String, value: i64 } fn demo(output: duckdb_data_chunk, data: &[Item]) { let mut cw = unsafe { DataChunk::from_raw(output).into_chunk_writer() }; for item in data { let Some(row) = cw.next_row() else { break }; // chunk is full unsafe { cw.writer(0).write_varchar(row, &item.name) }; unsafe { cw.writer(1).write_i64(row, item.value) }; } // set_size called automatically when `cw` is dropped } }
ValidityBitmap
For advanced NULL handling beyond VectorWriter::set_null, use ValidityBitmap
directly:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; fn demo(some_vector: duckdb_vector, row: usize) { use quack_rs::vector::ValidityBitmap; // Writing NULLs: let mut bitmap = unsafe { ValidityBitmap::ensure_writable(some_vector) }; unsafe { bitmap.set_row_invalid(row as u64) }; // mark as NULL unsafe { bitmap.set_row_valid(row as u64) }; // mark as non-NULL // Reading NULLs: let bitmap = unsafe { ValidityBitmap::get_read_only(some_vector) }; let is_valid = unsafe { bitmap.row_is_valid(row as u64) }; } }
ValidityBitmap is available in the prelude: use quack_rs::prelude::*.
Utility functions
The quack_rs::vector module provides two utility functions:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; fn demo(some_vector: duckdb_vector) { use quack_rs::vector::{vector_size, vector_get_column_type}; // Rows per data chunk: 2048 unless DuckDB was built with another STANDARD_VECTOR_SIZE. let size: u64 = vector_size(); // Returns the LogicalType of a vector (unsafe — requires a valid duckdb_vector). let lt = unsafe { vector_get_column_type(some_vector) }; } }
Memory layout details
DuckDB stores vector data as flat arrays. VectorReader and VectorWriter compute
element addresses as base_ptr + row * stride:
[value0][value1][value2]...[valueN] ← typed array
[validity bitmap] ← separate bit array, 1 bit per row
The validity bitmap is lazily allocated — it may be null if no NULLs have been written.
This is why duckdb_vector_ensure_validity_writable must be called before
duckdb_vector_get_validity when writing NULLs; VectorWriter and
ValidityBitmap::ensure_writable do so.
Complete scalar function pattern
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::vector::{VectorReader, VectorWriter}; fn transform(v: i64) -> i64 { v } unsafe extern "C" fn my_scalar( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..reader.row_count() { if unsafe { !reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } let value = unsafe { reader.read_i64(row) }; unsafe { writer.write_i64(row, transform(value)) }; } } }
Values & Parameter Extraction
DuckDB hands an extension single values — a table function's bind-time
parameters, the options of a COPY statement, a folded constant expression — as
duckdb_value handles, which are heap-allocated and must be destroyed after
use. quack-rs wraps them in Value, which destroys the handle on drop and reads
it through typed getters that return Option instead of aborting on SQL NULL.
This page covers reading parameters with Value, building values, and nested
(LIST, STRUCT, MAP) values.
The problem
Without Value, every parameter extraction requires three raw FFI calls and
careful manual cleanup:
#![allow(unused)] fn main() { use libduckdb_sys::*; unsafe fn demo(info: duckdb_bind_info) { // Before: raw FFI — easy to leak memory, and `duckdb_get_int64` aborts the // process if the argument is SQL NULL (DuckDB throws a C++ exception). let mut param = unsafe { duckdb_bind_get_parameter(info, 0) }; let n = unsafe { duckdb_get_int64(param) }; unsafe { duckdb_destroy_value(&mut param) }; // forget this → memory leak } }
The solution: Value
Value wraps a duckdb_value handle and calls duckdb_destroy_value on drop:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; use quack_rs::table::BindInfo; unsafe extern "C" fn my_bind(info: duckdb_bind_info) { let bind_info = unsafe { BindInfo::new(info) }; // Value is RAII — automatically destroyed when dropped. // `as_i64` is `None` for SQL NULL; `as_i64_or` supplies a default. let n = unsafe { bind_info.get_parameter_value(0) }.as_i64_or(0); // Named parameters work the same way let path = unsafe { bind_info.get_named_parameter_value("path") } .as_str() .unwrap_or_default(); } }
Typed extraction methods
| Method | Reads as | Rust type |
|---|---|---|
as_str() | any value, cast to text | Result<String, ExtensionError> |
as_blob() | BLOB only | Result<Vec<u8>, ExtensionError> |
as_i8() | TINYINT | Option<i8> |
as_i16() | SMALLINT | Option<i16> |
as_i32() | INTEGER | Option<i32> |
as_i64() | BIGINT | Option<i64> |
as_i128() | HUGEINT | Option<i128> |
as_u8() | UTINYINT | Option<u8> |
as_u16() | USMALLINT | Option<u16> |
as_u32() | UINTEGER | Option<u32> |
as_u64() | UBIGINT | Option<u64> |
as_u128() | UHUGEINT | Option<u128> |
as_f32() | FLOAT | Option<f32> |
as_f64() | DOUBLE | Option<f64> |
as_bool() | BOOLEAN | Option<bool> |
as_date(), as_time(), as_time_tz(), as_timestamp(), … | the temporal types | Option<i32> / Option<i64> / Option<u64> |
as_interval(), as_uuid() | INTERVAL, UUID | Option<DuckInterval>, Option<u128> |
as_decimal() | DECIMAL only (no cast) | Option<Decimal> |
as_enum_index() | ENUM only (no cast) | Option<u64> |
The scalar getters cast the way SQL's TRY_CAST does: a VARCHAR '42'
reads as Some(42) through as_i64(), a DOUBLE 1.5 as Some(2) through
as_i32(). They return None when:
- the value is SQL
NULL(for examplemy_func(n := NULL)), - the handle is null (a named parameter the caller did not supply),
- the type is not a scalar (
LIST,STRUCT,MAP,BLOB,ENUM, …), or - the cast fails (
'abc', or a number out of range for the target), or - for the temporal getters, SQL would refuse the conversion: the time of an
infinite timestamp (
as_time()of'infinity'::TIMESTAMP), aTIMESTAMPoutsideTIMESTAMP_NS's 1677–2262 range read withas_timestamp_ns(), or a result outside the target type's range.
No getter calls a duckdb_get_* function in the first three cases — those
functions abort the process on a SQL NULL and crash on a null handle — and
none modifies the value it reads (DuckDB's getters cast the value in place;
quack-rs reads from a copy). The temporal cases are checked before
the call too: DuckDB converts those pairs with a cast that throws a C++
exception instead of failing, which aborts the process from Rust.
as_blob() copies the bytes into an owned Vec<u8> without UTF-8 validation.
It accepts only a BLOB: DuckDB's conversion of anything else to BLOB can
throw. Use as_str() for text.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe fn demo(bind_info: BindInfo) -> Result<(), ExtensionError> { let bytes = unsafe { bind_info.get_parameter_value(0) }.as_blob()?; Ok(()) } }
Defaulting variants
The integer (except as_u128), float, bool and string getters have an
_or(default) variant that returns default wherever the plain getter returns
None (or Err); as_str_or_default() returns an empty string:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; // Values are DuckDB objects: fill the dispatch table first. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let val = Value::null_value(); let timeout = val.as_i64_or(30); // 30 if NULL, absent, or not a number let host = val.as_str_or("localhost"); // "localhost" if NULL or absent let port = val.as_u16_or(5432); // 5432 if NULL, absent, or out of range assert_eq!((timeout, host.as_str(), port), (30, "localhost", 5432)); }
Checking for NULL
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe fn demo(bind_info: BindInfo) { let val = unsafe { bind_info.get_named_parameter_value("limit") }; if val.is_null() { // the handle is null: the named parameter was not provided } if val.is_sql_null() { // the parameter was provided as SQL NULL } } }
Escape hatch
If you need the raw handle for an API quack-rs does not wrap, as_raw() borrows
it and into_raw() takes ownership:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; use libduckdb_sys::duckdb_value; unsafe fn demo(bind_info: BindInfo) { let val = unsafe { bind_info.get_parameter_value(0) }; let raw: duckdb_value = val.into_raw(); // takes ownership, no auto-destroy // ... use raw handle ... // caller must call duckdb_destroy_value manually } }
Building values
Value::bigint, Value::varchar, Value::date, Value::interval and the
other scalar constructors are infallible, with two exceptions that return
Result<Value, ExtensionError>:
- The temporal constructors that take a raw 64-bit payload —
time,time_tz,time_ns(withduckdb-1-5),timestamp,timestamp_tz,timestamp_s,timestamp_msandtimestamp_ns. DuckDB stores any payload unchecked, and rendering or casting an out-of-range one aborts, crashes or prints garbage, so quack-rs accepts exactly the range DuckDB's SQL produces (TIME00:00:00–24:00:00, theTIMESTAMPspan290309-12-22 (BC)–294247-01-10plus±infinity, and so on) and returns an error otherwise. Value::decimal(width, scale, unscaled), which checks thatwidthis1..=38,scale <= widthandunscaledhas at mostwidthdigits.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; // Values are DuckDB objects: fill the dispatch table first. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let noon = Value::time(12 * 3_600 * 1_000_000)?; assert!(Value::time(-1).is_err()); assert_eq!(noon.as_time(), Some(12 * 3_600 * 1_000_000)); Ok::<(), ExtensionError>(()) }
To build a nested value, Value::list_value and Value::array_value take the element
type and the items, Value::struct_value the STRUCT type and one value per
field, and Value::enum_value the ENUM type and an index; Value::map and
Value::union_value need duckdb-1-5. All return Result<Value, ExtensionError>.
Reading nested values
A parameter of type LIST, STRUCT or MAP is read element by element; each
accessor that returns a Value returns an owned one.
| Method | Returns |
|---|---|
list_len() | Number of elements of a LIST (0 for any other type) |
list_child(i) / list_items() | Element i (Option<Value>) / all elements (Vec<Value>) |
struct_field_names() | Field names of a STRUCT, in order |
struct_child(i) | Field i of a STRUCT (Option<Value>); fields are positional |
map_len(), map_key(i), map_value(i) | Number of pairs; key / value of pair i |
#![allow(unused)] fn main() { use quack_rs::value::Value; fn demo(options: &Value) -> Option<String> { let names = options.struct_field_names(); let idx = names.iter().position(|n| n == "compression")?; options.struct_child(idx)?.as_str().ok() } }
DataChunk, which wraps the chunk a table function's scan callback writes its
output to, is described in Reading & Writing Vectors.
Complex Types: STRUCT, LIST, MAP, ARRAY
DuckDB stores its nested types — STRUCT, LIST, MAP and ARRAY — as a parent
vector with one or more child vectors. This page shows how a quack-rs extension reads
and writes them: the four helper types in vector::complex reach the child vectors,
and ListBuilder writes LIST and MAP output without manual offset arithmetic.
Overview
| DuckDB type | Storage | quack-rs helper |
|---|---|---|
STRUCT{a T, b U, …} | Parent vector + N child vectors (one per field) | StructVector |
LIST<T> | Parent vector holds {offset, length} per row; flat child vector holds elements | ListVector |
MAP<K, V> | Stored as LIST<STRUCT{key K, value V}> | MapVector |
ARRAY<T>[N] | Fixed-size array; single child vector | ArrayVector |
Reading complex types (input vectors)
STRUCT
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; fn demo(parent_vec: duckdb_vector, row_count: usize) { use quack_rs::vector::{VectorReader, complex::StructVector}; // Inside a scalar function or aggregate update callback: // parent_vec comes from duckdb_data_chunk_get_vector(chunk, col_idx) let x_reader = unsafe { StructVector::field_reader(parent_vec, 0, row_count) }; let y_reader = unsafe { StructVector::field_reader(parent_vec, 1, row_count) }; for row in 0..row_count { // Each field has its own validity bitmap. if unsafe { x_reader.is_valid(row) && y_reader.is_valid(row) } { let x: f64 = unsafe { x_reader.read_f64(row) }; let y: f64 = unsafe { y_reader.read_f64(row) }; // process (x, y) … } } } }
LIST
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; fn demo(list_vec: duckdb_vector, row_count: usize) { use quack_rs::vector::{VectorReader, complex::ListVector}; let total_elements = unsafe { ListVector::get_size(list_vec) }; let elem_reader = unsafe { ListVector::child_reader(list_vec, total_elements) }; for row in 0..row_count { let entry = unsafe { ListVector::get_entry(list_vec, row) }; for i in 0..entry.length as usize { let elem_idx = entry.offset as usize + i; if unsafe { elem_reader.is_valid(elem_idx) } { let val: i64 = unsafe { elem_reader.read_i64(elem_idx) }; // process val … } } } } }
MAP
MAP is LIST<STRUCT{key, value}>. MapVector::key_reader and value_reader
read the two fields of the inner struct:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; fn demo(map_vec: duckdb_vector, row_count: usize) { use quack_rs::vector::complex::MapVector; let total = unsafe { MapVector::total_entry_count(map_vec) }; let key_reader = unsafe { MapVector::key_reader(map_vec, total) }; let value_reader = unsafe { MapVector::value_reader(map_vec, total) }; for row in 0..row_count { let entry = unsafe { MapVector::get_entry(map_vec, row) }; for i in 0..entry.length as usize { let idx = entry.offset as usize + i; let k = unsafe { key_reader.read_str(idx) }; // MAP keys are never NULL if unsafe { value_reader.is_valid(idx) } { let v: i64 = unsafe { value_reader.read_i64(idx) }; // process (k, v) … } } } } }
Writing complex types (output vectors)
STRUCT
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; fn demo(out_vec: duckdb_vector, batch_size: usize, x_values: &[f64], y_values: &[f64]) { use quack_rs::vector::{VectorWriter, complex::StructVector}; let mut x_writer = unsafe { StructVector::field_writer(out_vec, 0) }; let mut y_writer = unsafe { StructVector::field_writer(out_vec, 1) }; for row in 0..batch_size { unsafe { x_writer.write_f64(row, x_values[row]) }; unsafe { y_writer.write_f64(row, y_values[row]) }; } } }
Nested complex types inside STRUCT (v0.11.0+)
When a STRUCT field is itself a LIST, MAP or ARRAY, child_vector(field_idx) on
StructWriter or StructReader returns the field's raw vector handle, which the
ListVector, MapVector and ArrayVector helpers and ListBuilder accept:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; fn demo(struct_vec: duckdb_vector, row: usize) { use quack_rs::vector::{ListBuilder, StructWriter}; // STRUCT(name VARCHAR, services VARCHAR[], message VARCHAR) let mut sw = unsafe { StructWriter::new(struct_vec, 3) }; // Write scalar fields normally unsafe { sw.write_varchar(row, 0, "hello") }; unsafe { sw.write_varchar(row, 2, "ok") }; // The LIST field at index 1: ListBuilder appends after any elements // earlier rows already wrote to the child vector. let services = ["a", "b", "c"]; let mut builder = unsafe { ListBuilder::new(sw.child_vector(1)) }; unsafe { builder.push_row(row, services.len(), |writer, base| { for (i, s) in services.iter().enumerate() { writer.write_varchar(base + i, s); } }); builder.finish(); } } }
LIST — recommended: ListBuilder
ListBuilder tracks the running offset, writes each parent row's
{offset, length} entry, and — importantly — re-fetches the child writer after
every reserve:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; fn demo(list_vec: duckdb_vector, rows: &[Vec<i64>]) { use quack_rs::vector::ListBuilder; let mut builder = unsafe { ListBuilder::new(list_vec) }; for (row, elements) in rows.iter().enumerate() { unsafe { builder.push_row(row, elements.len(), |writer, base| { for (i, &val) in elements.iter().enumerate() { writer.write_i64(base + i, val); } }); } } unsafe { builder.finish() }; } }
Why the re-fetch matters.
duckdb_list_vector_reservetakes a total capacity, and when it grows it reallocates the child vector's data buffer. AVectorWriterobtained before that call is left holding a dangling pointer. The manual pattern below is safe only because it reserves exactly once, before any writer exists — which requires knowing the total element count up front.ListBuilderhas no such requirement. The same applies to writers on anything below the child, such as the fields of aLISTofSTRUCTs: fetch them again after every reserve that grows the list (see Pitfall L18).
push_map_row does the same for MAP, handing the closure a writer for the key
child and one for the value child.
DuckDB limits a child vector to 2^37 bytes per buffer
(MAX_LIST_CHILD_CAPACITY), and a reservation above that — or one the
allocator cannot satisfy — throws a C++ exception through the C API, which
aborts the process. vector::max_child_capacity(vec) turns the byte limit into
an element count for the child's type: 2^34 BIGINTs, 2^33 VARCHARs.
ListBuilder applies it by itself. When row lengths come from untrusted input,
also set a limit that fits in memory with with_element_limit(n): a row that
would exceed the limit is written as NULL instead, and overflowed() reports
that it happened.
LIST — manual
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; fn demo(list_vec: duckdb_vector, rows: &[Vec<i64>]) { use quack_rs::vector::{VectorWriter, complex::ListVector}; let total_elements: usize = rows.iter().map(|r| r.len()).sum(); // Must not exceed quack_rs::vector::max_child_capacity(list_vec); see above. unsafe { ListVector::reserve(list_vec, total_elements) }; let mut child_writer = unsafe { ListVector::child_writer(list_vec) }; let mut offset = 0usize; for (row, elements) in rows.iter().enumerate() { for (i, &val) in elements.iter().enumerate() { unsafe { child_writer.write_i64(offset + i, val) }; } unsafe { ListVector::set_entry(list_vec, row, offset as u64, elements.len() as u64) }; offset += elements.len(); } unsafe { ListVector::set_size(list_vec, total_elements) }; } }
MAP — manual
Writing a MAP follows the LIST pattern, but keys and values go into the two
fields of the inner STRUCT vector. Prefer ListBuilder::push_map_row unless you
know the total pair count before writing:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; fn demo(map_vec: duckdb_vector, total_pairs: usize, all_pairs: &[Vec<(String, i64)>]) { use quack_rs::vector::complex::MapVector; unsafe { MapVector::reserve(map_vec, total_pairs) }; let mut key_writer = unsafe { MapVector::key_writer(map_vec) }; let mut val_writer = unsafe { MapVector::value_writer(map_vec) }; let mut offset = 0usize; for (row, pairs) in all_pairs.iter().enumerate() { for (i, (k, v)) in pairs.iter().enumerate() { unsafe { key_writer.write_varchar(offset + i, k) }; unsafe { val_writer.write_i64(offset + i, *v) }; } unsafe { MapVector::set_entry(map_vec, row, offset as u64, pairs.len() as u64) }; offset += pairs.len(); } unsafe { MapVector::set_size(map_vec, total_pairs) }; } }
Constructing complex logical types
Use LogicalType constructors to define complex column types. Each constructor
has a variant that accepts TypeId values (for simple element types) and a
_from_logical variant (for nested complex types):
| Constructor | _from_logical variant | Creates |
|---|---|---|
LogicalType::list(TypeId) | list_from_logical(&LogicalType) | LIST<T> |
LogicalType::map(TypeId, TypeId) | map_from_logical(&LogicalType, &LogicalType) | MAP<K, V> |
LogicalType::struct_type(&[(&str, TypeId)]) | struct_type_from_logical(&[(&str, LogicalType)]) | STRUCT{...} |
LogicalType::union_type(&[(&str, TypeId)]) | union_type_from_logical(&[(&str, LogicalType)]) | UNION(...) |
LogicalType::array(TypeId, u64) | array_from_logical(&LogicalType, u64) | ARRAY<T>[N] |
LogicalType::enum_type(&[&str]) | — | ENUM(...) |
LogicalType::decimal(u8, u8) | — | DECIMAL(w, s) |
Each constructor also has a try_ form (try_list, try_struct_type_from_logical, …)
that returns Result<LogicalType, LogicalTypeError>; the plain forms panic where
the try_ form returns an error. Errors include a composite TypeId passed where
a _from_logical variant is needed, STRUCT field or UNION member names that
are equal ignoring ASCII case, and a UNION with more than MAX_UNION_MEMBERS
(255) members.
API reference
All helpers are in quack_rs::vector::complex (re-exported from quack_rs::prelude).
StructVector
| Method | Description |
|---|---|
get_child(vec, field_idx) | Returns the raw child vector for field field_idx |
field_reader(vec, field_idx, row_count) | Creates a VectorReader for a STRUCT field |
field_writer(vec, field_idx) | Creates a VectorWriter for a STRUCT field |
StructWriter / StructReader complex field access (v0.11.0+)
| Method | Description |
|---|---|
StructWriter::child_vector(field_idx) | Returns the raw duckdb_vector of a nested field (LIST, MAP, ARRAY) |
StructWriter::child_list_vector(field_idx) | Alias of child_vector for a LIST field |
StructReader::child_vector(field_idx) | Same as StructWriter::child_vector, for reading (unsafe) |
ListVector
| Method | Description |
|---|---|
get_child(vec) | Returns the flat element child vector |
get_size(vec) | Total number of elements across all rows |
set_size(vec, n) | Sets the number of elements after writing |
reserve(vec, capacity) | Reserves capacity in the child vector (at most max_child_capacity(vec)) |
get_entry(vec, row) | Returns {offset, length} for a row (reading) |
set_entry(vec, row, offset, length) | Sets {offset, length} for a row (writing) |
child_reader(vec, count) | Creates a VectorReader for the element vector |
child_writer(vec) | Creates a VectorWriter for the element vector |
MapVector
| Method | Description |
|---|---|
struct_child(vec) | Returns the inner STRUCT vector |
keys(vec) | Returns the key vector (STRUCT field 0) |
values(vec) | Returns the value vector (STRUCT field 1) |
total_entry_count(vec) | Total key-value pairs |
reserve(vec, n) | Reserves capacity for n pairs (at most max_child_capacity(vec)) |
set_size(vec, n) | Sets total entry count after writing |
get_entry(vec, row) | Returns {offset, length} for a row (reading) |
set_entry(vec, row, offset, length) | Sets {offset, length} for a row (writing) |
key_reader(vec, count) / value_reader(vec, count) | Creates a VectorReader for the keys / values |
key_writer(vec) / value_writer(vec) | Creates a VectorWriter for the keys / values |
ArrayVector
| Method | Description |
|---|---|
get_child(vec) | Returns the child vector of a fixed-size ARRAY vector |
NULL Handling & Strings
This page covers checking for NULL before reading a DuckDB vector, writing NULL
output, and reading and writing VARCHAR and BLOB values. The two topics
belong together: reading a string from a NULL row is undefined behaviour, not
just a wrong value.
NULL checks
Every row in a DuckDB vector may be NULL. Always check validity before reading:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, writer: &mut VectorWriter) { for row in 0..reader.row_count() { if unsafe { !reader.is_valid(row) } { // Propagate NULL to output unsafe { writer.set_null(row) }; continue; } // Safe to read let value = unsafe { reader.read_str(row) }; } } }
Reading a fixed-width value from a NULL row returns garbage. The vector's data buffer is not zeroed at NULL positions, and no error is raised: you get whatever bytes the buffer holds at that position.
For VARCHAR and BLOB it is worse than garbage. A NULL row's 16-byte entry
is left as it was, and DuckDB reuses vector buffers between chunks, so the
entry can still be the pointer-format record of a string from an earlier chunk
whose memory may since have been freed or reused. read_str / read_blob on such a row
follow that pointer: undefined behaviour, not merely a wrong answer. Their
# Safety sections require the row to be valid.
Writing NULL
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(writer: &mut VectorWriter, row: usize) { unsafe { writer.set_null(row) }; } }
Pitfall L4:
VectorWriter::set_nullcallsduckdb_vector_ensure_validity_writablebefore accessing the validity bitmap. Without it, a vector that has no mask yet makesduckdb_vector_get_validityreturn NULL, and the NULL you write is silently dropped. Never write NULL manually; always useset_null— which, for aSTRUCTorARRAYoutput, also nulls the fields / elements of that row the way DuckDB expects. See Pitfall L4.
Clearing NULL (v0.11.0+)
To mark a row as valid after a previous set_null:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(writer: &mut VectorWriter, row: usize) { unsafe { writer.set_valid(row) }; } }
VARCHAR reading
Read VARCHAR columns with VectorReader::read_str:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, row: usize) { let s: &str = unsafe { reader.read_str(row) }; } }
The returned &str borrows from the DuckDB vector — it must not outlive the
callback. Do not store it in a struct; clone it to a String if you need to
keep it.
The duckdb_string_t format
Pitfall P7: the Rust bindings do not document the layout of
duckdb_string_t.quack-rsdecodes it for you; the details below are for reference. See Pitfall P7.
DuckDB stores VARCHAR values in a 16-byte duckdb_string_t struct with two
representations, selected at runtime based on string length:
| Format | Condition | Layout |
|---|---|---|
| Inline | length ≤ 12 | [len: u32][data: [u8; 12]] |
| Pointer | length > 12 | [len: u32][prefix: [u8; 4]][ptr: *const u8] |
On a 32-bit target (DuckDB-WASM) the pointer is 4 bytes and the last 4 bytes are unused. The length and the pointer are in the target's byte order.
VectorReader::read_str and the underlying read_duck_string function handle
both formats, so you do not need to inspect the raw struct. A value that is not
valid UTF-8 is returned as ""; use read_blob to get its bytes.
Empty strings vs NULL
An empty string ("") and NULL are distinct values:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, row: usize) { // NULL: is_valid returns false // Empty string: is_valid returns true, read_str returns "" if unsafe { !reader.is_valid(row) } { // This is NULL } else { let s = unsafe { reader.read_str(row) }; if s.is_empty() { // This is an empty string, not NULL } } } }
Writing VARCHAR
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(writer: &mut VectorWriter, row: usize, my_str: &str) { unsafe { writer.write_varchar(row, my_str) }; // &str } }
write_varchar copies the string bytes into DuckDB's managed storage, so the
&str need not outlive the call. write_blob does the same for &[u8].
Both panic for a value longer than MAX_STRING_LEN (u32::MAX bytes, the
most a duckdb_string_t length field can hold) instead of storing a truncated
value; inside scalar_callback! and the typed scalar constructors the panic
becomes a SQL error. try_write_varchar and try_write_blob return
Result<(), ExtensionError> instead and write nothing on error:
#![allow(unused)] fn main() { use quack_rs::vector::VectorWriter; use quack_rs::error::ExtensionError; fn demo(writer: &mut VectorWriter, row: usize, my_str: &str) -> Result<(), ExtensionError> { unsafe { writer.try_write_varchar(row, my_str)? }; Ok(()) } }
Reading BLOB values
BLOB uses the same inline/pointer layout as VARCHAR, but may contain any
bytes. Use read_blob so the data is not interpreted as UTF-8:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, row: usize) { let bytes: &[u8] = unsafe { reader.read_blob(row) }; } }
Like read_str, the returned slice borrows from the DuckDB vector and must not
outlive the callback.
Complete NULL-safe VARCHAR pattern
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::vector::{VectorReader, VectorWriter}; unsafe extern "C" fn my_scalar( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..reader.row_count() { if unsafe { !reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } let s = unsafe { reader.read_str(row) }; let upper = s.to_uppercase(); unsafe { writer.write_varchar(row, &upper) }; } } }
DuckStringView
For advanced use cases where you need access to the raw string bytes or the
inline/pointer distinction, quack_rs::vector::string::DuckStringView is
available:
#![allow(unused)] fn main() { fn demo(data: *const u8, idx: usize) { use quack_rs::vector::string::{DuckStringView, DUCK_STRING_SIZE}; // From raw 16-byte data (inside a vector callback). `from_raw` is unsafe because it // follows the pointer of a string longer than 12 bytes; untrusted bytes go through // `DuckStringView::inline_from_bytes`, which refuses that format instead. let raw: &[u8; 16] = unsafe { &*data.add(idx * DUCK_STRING_SIZE).cast() }; let view = unsafe { DuckStringView::from_raw(raw) }; println!("length: {}", view.len()); println!("is_empty: {}", view.is_empty()); if let Some(s) = view.as_str() { println!("content: {s}"); } } }
In practice, prefer reader.read_str(row). DuckStringView is needed only when
you have a raw data pointer rather than a VectorReader. Unlike read_str, its
as_str returns None, not "", for a value that is not valid UTF-8.
Constants
| Constant | Value | Meaning |
|---|---|---|
DUCK_STRING_SIZE | 16 | Size of one duckdb_string_t in bytes |
DUCK_STRING_INLINE_MAX_LEN | 12 | Longest value stored inline (no heap pointer), in bytes |
MAX_STRING_LEN | u32::MAX (4,294,967,295) | Longest VARCHAR or BLOB value DuckDB can store, in bytes |
All three are in quack_rs::vector::string.
INTERVAL Type
DuckDB's INTERVAL type represents a duration with three independent components:
months, days and microseconds. The quack_rs::interval module provides the
DuckInterval struct, which matches DuckDB's in-memory layout, and overflow-safe
conversions to microseconds.
Why a custom struct?
Pitfall P8: the Rust bindings do not document the
INTERVALlayout or how DuckDB converts intervals.DuckIntervaland the functions below encode both. See Pitfall P8.
DuckDB's C duckdb_interval struct is 16 bytes with this exact layout:
offset 0: months (i32) — calendar months
offset 4: days (i32) — calendar days
offset 8: micros (i64) — microseconds (not limited to one day)
total: 16 bytes
DuckInterval is #[repr(C)] with the same field order, and a compile-time
assertion checks that it is exactly 16 bytes.
Reading INTERVAL values
#![allow(unused)] fn main() { use quack_rs::interval::DuckInterval; use quack_rs::vector::VectorReader; fn demo(reader: &VectorReader, row: usize) { let iv: DuckInterval = unsafe { reader.read_interval(row) }; println!("{} months, {} days, {} µs", iv.months, iv.days, iv.micros); } }
VectorReader::read_interval handles the raw pointer arithmetic and alignment
using read_interval_at internally.
DuckInterval fields
#![allow(unused)] fn main() { use quack_rs::interval::DuckInterval; let iv = DuckInterval { months: 1, // 1 calendar month days: 15, // 15 calendar days micros: 3_600_000_000, // 1 hour in microseconds }; }
The fields are public, so a DuckInterval can be built directly.
The derived PartialEq, Eq and Hash compare the three fields, so
{ months: 1, .. } and { days: 30, .. } are different values here, while in
SQL INTERVAL '1 month' = INTERVAL '30 days' is true.
Zero interval
#![allow(unused)] fn main() { use quack_rs::interval::DuckInterval; let zero = DuckInterval::zero(); // { months: 0, days: 0, micros: 0 } let zero = DuckInterval::default(); // same assert_eq!(zero, DuckInterval::zero()); }
Converting to microseconds
Months and days have no fixed length in wall-clock time, so an interval has no
single exact length. When you need one number, for ordering or bucketing,
convert to microseconds with the approximation DuckDB uses when it compares
intervals and in epoch_us(interval): 1 month = 30 days.
This is not date arithmetic: DuckDB adds an interval to a date by calendar
months (DATE '2024-01-31' + INTERVAL 1 MONTH is 2024-02-29). It also matches
SQL comparison only when the three fields share a sign: INTERVAL '1 month' - INTERVAL '1 day' converts to the same total as INTERVAL '29 days' but compares
greater in SQL.
Checked conversion (returns Option)
#![allow(unused)] fn main() { use quack_rs::interval::DuckInterval; use quack_rs::interval::interval_to_micros; let iv = DuckInterval { months: 0, days: 1, micros: 500_000 }; match interval_to_micros(iv) { Some(us) => println!("{us} microseconds"), None => println!("overflow"), } // Method form: let us: Option<i64> = iv.to_micros(); assert_eq!(us, Some(86_400_000_000 + 500_000)); }
Returns None if the total does not fit in an i64, which takes extreme values
such as months: i32::MAX, days: i32::MAX, micros: i64::MAX. The sum is computed
exactly, so large fields of opposite signs that cancel out still convert.
Saturating conversion (returns i64)
#![allow(unused)] fn main() { use quack_rs::interval::DuckInterval; use quack_rs::interval::interval_to_micros_saturating; let iv = DuckInterval { months: i32::MAX, days: i32::MAX, micros: i64::MAX }; let us: i64 = interval_to_micros_saturating(iv); // i64::MAX // Method form: let us: i64 = iv.to_micros_saturating(); assert_eq!(us, i64::MAX); assert_eq!(interval_to_micros_saturating(iv), i64::MAX); }
The saturating form clamps an out-of-range total to i64::MAX or i64::MIN.
Neither form panics; use the checked form when an overflow must be reported
rather than clamped.
Conversion constants
| Constant | Value | Meaning |
|---|---|---|
MICROS_PER_DAY | 86_400_000_000 | Microseconds in 24 hours |
MICROS_PER_MONTH | 2_592_000_000_000 | Microseconds in 30 days |
#![allow(unused)] fn main() { use quack_rs::interval::{MICROS_PER_DAY, MICROS_PER_MONTH}; assert_eq!(MICROS_PER_DAY, 86_400 * 1_000_000); assert_eq!(MICROS_PER_MONTH, 30 * MICROS_PER_DAY); }
Low-level: read_interval_at
If you have a raw data pointer (e.g., from duckdb_vector_get_data), you can
read an interval directly:
#![allow(unused)] fn main() { fn demo(data_ptr: *const u8, row_idx: usize) { use quack_rs::interval::read_interval_at; // SAFETY: data is a valid DuckDB INTERVAL vector data pointer, idx is in bounds. let iv = unsafe { read_interval_at(data_ptr, row_idx) }; } }
In practice, use VectorReader::read_interval(row), which computes the data
pointer for you; its remaining unsafe contract is the row index and the
column type.
Complete example: aggregate over INTERVAL
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_data_chunk, duckdb_function_info}; use quack_rs::aggregate::{AggregateState, FfiState}; use quack_rs::vector::VectorReader; #[derive(Default)] struct TotalDurationState { total_micros: i64, } impl AggregateState for TotalDurationState {} unsafe extern "C" fn update( _info: duckdb_function_info, input: duckdb_data_chunk, states: *mut duckdb_aggregate_state, ) { let reader = unsafe { VectorReader::new(input, 0) }; for row in 0..reader.row_count() { if unsafe { !reader.is_valid(row) } { continue; } let iv = unsafe { reader.read_interval(row) }; let us = iv.to_micros_saturating(); let state_ptr = unsafe { *states.add(row) }; if let Some(st) = unsafe { FfiState::<TotalDurationState>::with_state_mut(state_ptr) } { st.total_micros = st.total_micros.saturating_add(us); } } } }
Memory layout verification
A compile-time assertion checks that DuckInterval is 16 bytes with at least
4-byte alignment, the layout of DuckDB's duckdb_interval. If it fails, the
crate does not compile, so a layout change is caught at build time rather than
at run time.
Dates, Times and Timestamps
VectorReader and VectorWriter read and write DuckDB's DATE, TIME,
TIMETZ, TIMESTAMP and INTERVAL values as the raw integers DuckDB stores;
quack_rs::datetime converts them to and from calendar fields. This page also
covers DECIMAL and HUGEINT, whose conversions live in the same module.
| SQL type | Storage | Accessor |
|---|---|---|
DATE | i32 — days since 1970-01-01 | read_date / write_date |
TIME | i64 — microseconds since midnight | read_time / write_time |
TIMETZ | packed u64 | read_time_tz / write_time_tz |
TIMESTAMP | i64 — microseconds since the epoch | read_timestamp / write_timestamp |
TIMESTAMPTZ | i64 — microseconds since the epoch, UTC | read_timestamp_tz / write_timestamp_tz |
TIMESTAMP_S | i64 — seconds since the epoch | read_timestamp_s / write_timestamp_s |
TIMESTAMP_MS | i64 — milliseconds since the epoch | read_timestamp_ms / write_timestamp_ms |
TIMESTAMP_NS | i64 — nanoseconds since the epoch | read_timestamp_ns / write_timestamp_ns |
INTERVAL | { months: i32, days: i32, micros: i64 } | read_interval / write_interval (see INTERVAL Type) |
Turning those integers into year/month/day means implementing the proleptic
Gregorian calendar, and getting it to agree with DuckDB's SQL semantics exactly
rather than approximately. DuckDB already exposes the conversions, and they are
in the stable prefix of the C API, so quack_rs::datetime
wraps them instead of reimplementing them. They need no feature flag.
Decomposing and composing
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, writer: &mut VectorWriter, row: usize) { use quack_rs::datetime; // DATE -> calendar date let days = unsafe { reader.read_date(row) }; let date = unsafe { datetime::date_from_days(days) }; println!("{:04}-{:02}-{:02}", date.year, date.month, date.day); // …and back. `None` means DuckDB cannot represent the date. match unsafe { datetime::date_to_days(date) } { Some(days) => unsafe { writer.write_date(row, days) }, None => unsafe { writer.set_null(row) }, } } }
Time, TimeTz and Timestamp work the same way:
#![allow(unused)] fn main() { use quack_rs::datetime; use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, writer: &mut VectorWriter, rows: usize) { for row in 0..rows { let Some(ts) = (unsafe { datetime::timestamp_from_micros(reader.read_timestamp(row)) }) else { // ±infinity (or the first ~4 hours of the i64 range): no calendar form. unsafe { writer.set_null(row) }; continue; }; assert!((0..1_000_000).contains(&ts.time.micros)); // ts.date and ts.time are plain structs let micros = unsafe { datetime::timestamp_to_micros(ts) }; // Option<i64> } } }
Invalid input is None, not an abort
Several of DuckDB's conversions throw a C++ exception on bad input, and the
C API does not catch it — so calling them directly with, say, month 13 aborts
the whole process ("Rust cannot catch foreign exceptions"). The wrappers apply
DuckDB's own conditions first and return None instead:
| Function | Returns None when |
|---|---|
date_to_days | month not 1–12, day not in that month (leap years included), or the date is outside 5877642-06-25 BC – 5881580-07-10; datetime::is_valid_date is the same check |
timestamp_from_micros | the value is ±infinity, or below -106_751_991 * MICROS_PER_DAY (which includes i64::MIN) |
timestamp_to_micros | the date is invalid, the result overflows i64, or it lands on ±infinity |
time_from_micros | the value is outside 0..=MICROS_PER_DAY (00:00:00–24:00:00) |
time_tz_bits | the time is outside 0..=MICROS_PER_DAY, or the offset beyond ±15:59:59 (TIME_TZ_MAX_OFFSET_SECONDS) |
time_tz_from_bits | the bits decode to a time or offset that time_tz_bits would refuse |
decimal_to_f64 | width > 38 or scale > width |
time_from_micros and time_tz_from_bits guard an assertion rather than an
exception: a release build of DuckDB decomposes an out-of-range time into
out-of-range fields, and a build with assertions enabled aborts.
time_to_micros does no range check, exactly like DuckDB: an hour of 25
simply gives a TIME past midnight.
TIMETZ is a packed 64-bit value, not a plain integer — build and read it
through the helpers rather than by hand:
#![allow(unused)] fn main() { use quack_rs::datetime; use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, writer: &mut VectorWriter, row: usize) { let bits = unsafe { datetime::time_tz_bits(12 * 3_600 * 1_000_000, -5 * 3_600) } .expect("noon, UTC-5, is in range"); unsafe { writer.write_time_tz(row, bits) }; let decoded = unsafe { datetime::time_tz_from_bits(reader.read_time_tz(row)) } .expect("DuckDB wrote a valid TIMETZ"); assert_eq!(decoded.offset_seconds, -5 * 3_600); } }
Infinity
DuckDB reserves two values of DATE and of TIMESTAMP for infinity and
-infinity. Decomposing one into a calendar date is meaningless, so check first:
#![allow(unused)] fn main() { use quack_rs::datetime; use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, writer: &mut VectorWriter, row: usize) { let days = unsafe { reader.read_date(row) }; if unsafe { datetime::is_finite_date(days) } { let date = unsafe { datetime::date_from_days(days) }; // … } } }
Note the exact values, which are easy to get wrong:
| Constant | Value |
|---|---|
DATE_INFINITY_DAYS | i32::MAX |
DATE_NEGATIVE_INFINITY_DAYS | -i32::MAX |
TIMESTAMP_INFINITY_MICROS | i64::MAX |
TIMESTAMP_NEGATIVE_INFINITY_MICROS | -i64::MAX |
Negative infinity is -i32::MAX, not i32::MIN. i32::MIN is an ordinary
(if absurd) finite date, and treating it as infinity would silently drop real
rows.
DECIMAL
DECIMAL is stored in the narrowest integer that fits its declared width, so the
width has to travel with the value:
| Declared width | Physical storage |
|---|---|
| 1 – 4 | i16 |
| 5 – 9 | i32 |
| 10 – 18 | i64 |
| 19 – 38 | i128 |
read_decimal / write_decimal take the width and pick the right one. Get it
from the column's LogicalType:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_vector; use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(vec: duckdb_vector, reader: &VectorReader, writer: &mut VectorWriter, row: usize) { let logical = unsafe { quack_rs::vector::vector_get_column_type(vec) }; let width = unsafe { logical.decimal_width() }; let scale = unsafe { logical.decimal_scale() }; let unscaled = unsafe { reader.read_decimal(row, width) }; // The represented number is unscaled / 10^scale. The doubled value must // still fit in `width` digits; write_decimal does not check. unsafe { writer.write_decimal(row, width, unscaled * 2) }; } }
datetime::f64_to_decimal and datetime::decimal_to_f64 convert through
DuckDB's own routines when a floating-point view is what you want.
decimal_to_f64 returns None for a width above 38 or a scale above the
width: DuckDB would index its powers-of-ten table out of bounds.
Wide integers
HUGEINT is { lower: u64, upper: i64 } and UHUGEINT is two u64s.
read_i128 / write_i128 and read_u128 / write_u128 handle the halves;
datetime::hugeint_to_f64, f64_to_hugeint, uhugeint_to_f64 and
f64_to_uhugeint convert through DuckDB's own routines, so they round as
DuckDB does.
Running SQL from an Extension
Extensions often need to run SQL against the database that is loading them:
checking whether a table exists before registering a replacement scan, creating a
helper view, reading a setting, or looking up a credential through
duckdb_secrets(). This page covers queries, prepared statements with bound
parameters, keeping a connection for use after loading, and cancelling a
running query.
The C API has everything for that — duckdb_query, duckdb_prepare,
duckdb_bind_*, duckdb_fetch_chunk — and all of it is in the stable
prefix, so it needs no feature flag. What it does not have
is any help releasing the handles: every one of them has a matching destroy
that must run exactly once, including on the error paths, which is where
hand-written FFI usually leaks.
quack_rs::query wraps them:
| Type | Owns | Released by |
|---|---|---|
QueryResult | duckdb_result | duckdb_destroy_result |
OwnedDataChunk | duckdb_data_chunk | duckdb_destroy_data_chunk |
PreparedStatement | duckdb_prepared_statement | duckdb_destroy_prepare |
OwnedConnection | duckdb_connection | duckdb_disconnect |
During registration
Connection (from entry_point_v2!) can run SQL directly:
#![allow(unused)] fn main() { use quack_rs::connection::Connection; use quack_rs::error::ExtensionError; fn register(con: &Connection) -> Result<(), ExtensionError> { // Create a helper view the extension's functions rely on. unsafe { con.execute("CREATE OR REPLACE VIEW my_ext_config AS SELECT 1 AS version") }?; // Read something back. The cast makes the column VARCHAR, which read_str requires. let mut result = unsafe { con.query("SELECT current_setting('threads')::VARCHAR") }?; if let Some(chunk) = result.next_chunk()? { // The reader must outlive the `&str` it hands out, so bind it first. let reader = unsafe { chunk.reader(0) }; let threads = unsafe { reader.read_str(0) }; eprintln!("DuckDB is using {threads} threads"); } Ok(()) } }
Results arrive a chunk at a time — at most duckdb_vector_size() rows each — so
call next_chunk until it returns Ok(None):
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; use quack_rs::query::OwnedConnection; // `Connection` (from entry_point_v2!) only exists during an extension load. An // `OwnedConnection` has the same query / execute / prepare methods. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let mut db = std::ptr::null_mut(); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); } let con = unsafe { OwnedConnection::open(db) }.unwrap(); let run = || -> Result<(), ExtensionError> { let mut result = unsafe { con.query("SELECT i FROM range(10000) t(i)") }?; let mut total: i64 = 0; while let Some(chunk) = result.next_chunk()? { let reader = unsafe { chunk.reader(0) }; for row in 0..chunk.size() { total += unsafe { reader.read_i64(row) }; } } assert_eq!(total, 49_995_000); Ok(()) }; run().unwrap(); }
next_chunk returns Result<Option<OwnedDataChunk>, ExtensionError> because a
result can stop early. A streaming result
(PreparedStatement::execute_streaming, with the duckdb-1-5 feature) produces
rows as it runs, so a runtime error part-way through, an interrupt, or another
statement run on the same connection (which invalidates the stream) surfaces at
next_chunk. The C API reports that the same way as the end of the rows (a
null chunk); quack-rs reads the error DuckDB recorded and returns it, so a
partial result cannot pass for a complete one. The ? above is what keeps it
from being silently truncated. After Ok(None) or an error, later calls return
the same thing.
Inspecting a result
| Method | Returns |
|---|---|
column_count() | Number of columns |
column_name(i) | Name of column i (Option<String>) |
column_type(i) | Top-level TypeId of column i |
column_logical_type(i) | Full LogicalType of column i, keeping STRUCT fields, LIST element type, DECIMAL width and scale |
result_kind() | ResultKind::Rows, ChangedRows, Nothing or Invalid |
rows_changed() | Rows changed by an INSERT / UPDATE / DELETE; 0 for other statements |
is_streaming() | Whether the result is streaming (duckdb-1-5) |
Several statements in one string
query and execute accept several ;-separated statements, and DuckDB runs
every one, in order. The result you get back is the first statement that
produces rows — or, when none does, the last statement's; the results of later
row-producing statements are discarded. So "SELECT 1; INSERT …" runs the
INSERT but execute reports 0 rows changed. The first failing statement
fails the call, after the ones before it have run (and, outside an explicit
transaction, committed). An empty string, or just ;, succeeds with an empty
result. prepare takes exactly one statement.
Bind values, do not interpolate them
Anything that did not come from your own source text — a table name from a function argument, a path from a config option — goes through a parameter. Parameters are 1-indexed, matching the C API.
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; use quack_rs::query::OwnedConnection; // `Connection` (from entry_point_v2!) only exists during an extension load. An // `OwnedConnection` has the same query / execute / prepare methods. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let mut db = std::ptr::null_mut(); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); } let con = unsafe { OwnedConnection::open(db) }.unwrap(); con.execute("CREATE TABLE audit (name VARCHAR, n BIGINT)").unwrap(); let user_supplied_name = "O'Brien'); DROP TABLE audit; --"; let run = || -> Result<(), ExtensionError> { let stmt = unsafe { con.prepare("INSERT INTO audit VALUES (?, ?)") }?; stmt.bind_str(1, user_supplied_name)?; // safe even if it contains quotes stmt.bind_i64(2, 42)?; stmt.execute()?; let mut r = con.query("SELECT count(*) FROM audit WHERE name LIKE 'O''Brien%'")?; assert_eq!(unsafe { r.next_chunk()?.unwrap().reader(0).read_i64(0) }, 1); Ok(()) }; run().unwrap(); }
bind_str passes the length explicitly, so embedded NUL bytes are preserved and
no CString conversion can fail.
There is a typed bind for every integer width (bind_i8 … bind_u128),
bind_f32 / bind_f64, bind_bool, bind_blob, bind_null, bind_decimal,
bind_date, bind_time, bind_timestamp, bind_timestamp_tz and
bind_interval. bind_value takes any Value,
which covers the composite types. Like the Value constructors, bind_decimal
validates width, scale and digit count, and bind_time and the timestamp
binds refuse a payload outside the range DuckDB's SQL produces; the C API's
duckdb_bind_* functions check nothing.
Named parameters resolve by name:
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; use quack_rs::query::OwnedConnection; // `Connection` (from entry_point_v2!) only exists during an extension load. An // `OwnedConnection` has the same query / execute / prepare methods. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let mut db = std::ptr::null_mut(); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); } let con = unsafe { OwnedConnection::open(db) }.unwrap(); con.execute("CREATE TABLE t AS SELECT range AS id FROM range(10)").unwrap(); let id = 7; let run = || -> Result<(), ExtensionError> { let stmt = unsafe { con.prepare("SELECT * FROM t WHERE id = $needle") }?; let index = stmt.parameter_index("needle").expect("named parameter"); stmt.bind_i64(index, id)?; let mut r = stmt.execute()?; assert_eq!(unsafe { r.next_chunk()?.unwrap().reader(0).read_i64(0) }, 7); Ok(()) }; run().unwrap(); }
Reuse a statement by clearing its bindings between executions:
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; use quack_rs::query::OwnedConnection; // `Connection` (from entry_point_v2!) only exists during an extension load. An // `OwnedConnection` has the same query / execute / prepare methods. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let mut db = std::ptr::null_mut(); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); } let con = unsafe { OwnedConnection::open(db) }.unwrap(); let run = || -> Result<(), ExtensionError> { let stmt = con.prepare("SELECT ?::BIGINT * 2")?; let ids = [1_i64, 2, 3]; for id in ids { stmt.clear_bindings()?; stmt.bind_i64(1, id)?; let mut result = stmt.execute()?; // … } Ok(()) }; run().unwrap(); }
After registration
The connection DuckDB passes to your entry point is borrowed — the entry
point disconnects it when your closure returns. Inside a scalar, table or
aggregate callback you have no connection at all: the C API gives you a
duckdb_client_context, and there is no duckdb_client_context_get_connection.
If a callback or a background thread needs to run SQL, open your own connection
during registration and keep it. A duckdb_connection holds its own reference to
the database instance, so it stays valid after loading finishes:
#![allow(unused)] fn main() { use quack_rs::connection::Connection; use quack_rs::error::ExtensionError; use quack_rs::query::OwnedConnection; use std::sync::{Mutex, OnceLock}; // `OwnedConnection` is `Send` but not `Sync` (see below), so a `static` must // hold it behind a `Mutex`. static CONN: OnceLock<Mutex<OwnedConnection>> = OnceLock::new(); fn register(con: &Connection) -> Result<(), ExtensionError> { let owned = unsafe { con.open_connection() }?; let _ = CONN.set(Mutex::new(owned)); Ok(()) } }
OwnedConnection is Send but deliberately not Sync: DuckDB permits moving a
connection between threads, not using one concurrently. Open one connection per
thread, or guard it with a mutex.
Cancelling a query and reading its progress
OwnedConnection::interrupt_handle returns an InterruptHandle, which is
Send + Sync and borrows the connection, so it cannot outlive it. Another
thread can call its cancel() to stop the running query, which then fails with
an interrupt error at DuckDB's next check, or its progress() to read a
QueryProgress (percentage, rows_processed, total_rows_to_process).
percentage is -1.0 when DuckDB cannot report progress, for example when
the progress bar is disabled (SET enable_progress_bar = true).
OwnedConnection::interrupt and progress do the same on the calling thread.
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; use quack_rs::query::{OwnedConnection, QueryResult}; use std::sync::mpsc::{channel, RecvTimeoutError}; use std::time::Duration; fn query_with_timeout(con: &OwnedConnection, sql: &str) -> Result<QueryResult, ExtensionError> { let watchdog = con.interrupt_handle(); let (done, finished) = channel::<()>(); std::thread::scope(|scope| { scope.spawn(move || { // Cancel the query if it has not finished within 30 seconds. if let Err(RecvTimeoutError::Timeout) = finished.recv_timeout(Duration::from_secs(30)) { watchdog.cancel(); } }); let result = con.query(sql); drop(done); // wakes the watchdog result }) } }
Errors
Failures carry DuckDB's own message:
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; use quack_rs::query::OwnedConnection; // `Connection` (from entry_point_v2!) only exists during an extension load. An // `OwnedConnection` has the same query / execute / prepare methods. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let mut db = std::ptr::null_mut(); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); } let con = unsafe { OwnedConnection::open(db) }.unwrap(); let err = unsafe { con.query("SELECT * FROM no_such_table") }.unwrap_err(); assert!(err.as_str().contains("no_such_table")); assert_eq!(con.execute("SELECT 1").unwrap(), 0); }
The connection stays usable afterwards.
Bulk Appender
Appender is an RAII wrapper around DuckDB's appender, the fastest way for an
extension to bulk-insert rows into an existing table: rows are buffered and
written in batches instead of going through one INSERT statement each. This
page covers appending row by row and chunk by chunk, error handling, and what
happens when a row fails half-way.
No feature flag required. DuckDB has kept duckdb_appender_* in the frozen
stable prefix of the extension API (slots 281–291 and 330–356) since v1.2.0, so
using the appender does not push your extension onto the version-pinned unstable
ABI. See ABI Compatibility for why that distinction
matters. Three methods are the exception and need duckdb-1-5: error_data,
clear, and append_default_to_chunk.
Row at a time
Call one append_* per column, then finish the row with end_row. row calls
end_row for you when its closure succeeds, so a forgotten end_row cannot
leave the table silently short:
#![allow(unused)] fn main() { use quack_rs::appender::{AppendError, Appender}; use libduckdb_sys::duckdb_connection; unsafe fn demo(con: duckdb_connection) -> Result<(), AppendError> { // SAFETY: `con` is a valid, open connection (e.g. from an entry point). let appender = unsafe { Appender::new(con, None, c"measurements") }?; for (sensor, reading) in [("a", 1.5_f64), ("b", 2.5)] { appender.row(|row| { row.append_str(sensor)?; row.append_f64(reading) })?; } appender.close()?; Ok(()) } fn live_connection() -> libduckdb_sys::duckdb_connection { std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess); } con } /// First column of the first row, as BIGINT; `None` for NULL. fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); let reader = unsafe { chunk.reader(0) }; unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) } } let con = live_connection(); unsafe { quack_rs::query::execute(con, "CREATE TABLE measurements (sensor VARCHAR, reading DOUBLE)") }.unwrap(); unsafe { demo(con) }.unwrap(); assert_eq!(query_i64(con, "SELECT count(*) FROM measurements"), Some(2)); assert_eq!(query_i64(con, "SELECT (sum(reading) * 10)::BIGINT FROM measurements"), Some(40)); }
append_str uses duckdb_append_varchar_length, so interior NUL bytes
survive; the NUL-terminated duckdb_append_varchar would stop at the first
one. append_str and append_bytes refuse a value longer than u32::MAX
bytes, which DuckDB would silently truncate.
append_time and append_timestamp refuse, before calling DuckDB, a payload
outside the range DuckDB's SQL produces (TIME 00:00:00–24:00:00; a
TIMESTAMP from 290309-12-22 (BC), plus -infinity). DuckDB stores such a
value unchecked, and reading the row back later fails or crashes.
A chunk at a time
append_chunk takes a whole DataChunk whose column types match the
appender's active columns. It makes fewer FFI calls than appending row by row
and suits data that is already in vectors; it is refused while a row appended
by hand is still open.
#![allow(unused)] fn main() { use quack_rs::appender::{AppendError, Appender}; use quack_rs::data_chunk::DataChunk; use libduckdb_sys::duckdb_connection; unsafe fn load(con: duckdb_connection, chunks: &[DataChunk]) -> Result<(), AppendError> { // SAFETY: `con` is a valid, open connection. let appender = unsafe { Appender::new(con, None, c"events") }?; for chunk in chunks { appender.append_chunk(chunk)?; } appender.close()?; Ok(()) } }
Pass a schema (or a fully-qualified catalog + schema) when the default schema is not what you want:
#![allow(unused)] fn main() { use quack_rs::appender::{AppendError, Appender}; use libduckdb_sys::duckdb_connection; unsafe fn demo(con: duckdb_connection) -> Result<(), AppendError> { let a = unsafe { Appender::new(con, Some(c"main"), c"events") }?; let b = unsafe { Appender::with_catalog(con, Some(c"mydb"), Some(c"main"), c"events") }?; let _ = (a, b); Ok(()) } }
Appending a subset of columns
add_column narrows the active column list; the omitted columns take their
DEFAULT (or NULL). Both add_column and clear_columns flush every row
appended so far.
#![allow(unused)] fn main() { use quack_rs::appender::{AppendError, Appender}; unsafe fn demo(appender: &Appender) -> Result<(), AppendError> { appender.add_column(c"id")?; // now only `id` is expected appender.row(|row| row.append_i32(7))?; appender.clear_columns()?; // back to every column Ok(()) } }
Use TableDescription::column_has_default to find out whether a column has
a default before relying on one. append_default() fails for a default that is
not a constant, such as nextval('seq') or now(): DuckDB's appender
evaluates defaults once, when it is created.
Errors arrive late, and invalidate the batch
Appended rows are buffered. A constraint violation therefore surfaces at flush
or close, not at the append_* call that caused it, and it invalidates
every buffered row.
With duckdb-1-5, clear discards the offending buffer so you can carry on
without re-appending rows that were already committed:
#![allow(unused)] fn main() { use quack_rs::appender::Appender; fn demo(appender: &Appender) { if let Err(err) = appender.flush() { eprintln!("flush failed: {err}"); let _ = appender.clear(); // drop the offending buffered rows } } }
A row that fails half-way loses the batch
DuckDB counts the values of the current row and cannot take one back. If a
row closure fails or panics after its first value went in, the half-written
row can be neither finished nor dropped, and DuckDB's close then returns
success while writing nothing: every row buffered since the last flush is
gone. (DuckDB also flushes by itself each time 204,800 rows accumulate; rows
written by such a flush are safe.)
quack-rs does not let that pass silently. The appender becomes poisoned:
every later row, append, end_row, flush and close returns an error
saying how many buffered rows were not written. With duckdb-1-5, clear()
discards them and makes the appender usable again; without it, create a new
appender. close() with a row that was started but not ended is an error
too (finish the row and close again).
A value that fails as the first of its row loses nothing, and when you
append by hand you can retry a rejected value in the same column. If losing a
batch is not acceptable, flush() at the points you can afford to go back to.
After a successful close() the appender refuses all further work.
Row order and schema changes
- Row-at-a-time rows wait in their own buffer until a full chunk (2,048 rows)
accumulates, while
append_chunkadds its rows to the table-bound buffer directly, so a chunk lands ahead of rows appended before it that are still buffered. Flush first if insertion order matters. - Buffered rows are written by column position at flush time. If another connection drops a column and adds one in between, a value lands in the new column with no error.
API
| Method | Description |
|---|---|
Appender::new(con, schema, table) (unsafe) | Create for table in schema (None = default) |
Appender::with_catalog(con, catalog, schema, table) (unsafe) | Create fully qualified |
column_count() / column_type(i) | The active column list |
add_column(name) / clear_columns() | Narrow / reset the active column list |
row(closure) | Append one row, calling end_row if the closure succeeds |
end_row() | Finish the current row explicitly |
append_bool/_i8/_i16/_i32/_i64/_i128 | Signed integers and BOOLEAN |
append_u8/_u16/_u32/_u64/_u128 | Unsigned integers |
append_f32/_f64 | FLOAT / DOUBLE |
append_str(&str) / append_bytes(&[u8]) | VARCHAR (NUL-safe) / BLOB |
append_date/_time/_timestamp/_interval | Temporal types |
append_value(&Value) | Anything else — LIST, STRUCT, MAP, UUID, DECIMAL, ENUM |
append_null() / append_default() | SQL NULL / the column's DEFAULT |
append_chunk(&chunk) | Append an entire DataChunk |
flush() / close() | Flush buffered rows / flush and close |
error_message() | Message from the last failed operation (Option<String>) |
append_default_to_chunk(&chunk, col, row) ¹ | Write a column's DEFAULT into a chunk cell |
clear() ¹ | Discard buffered, unflushed rows (and un-poison the appender) |
error_data() ¹ | Structured ErrorData from the last failed operation |
¹ Requires the
duckdb-1-5feature flag.
Every fallible method returns Result<_, AppendError>. AppendError is
ErrorData (message and machine-readable category) when duckdb-1-5 is
enabled, and ExtensionError (message only) otherwise — enabling the feature
upgrades the error type in place without changing any method's shape.
Safety
new and with_catalog are unsafe: you must pass a valid, open
duckdb_connection (such as the one provided to your extension's entry point).
Both return Result for a reason: DuckDB's duckdb_append_* functions do not
check whether the appender was created successfully before dereferencing it,
so an appender whose creation failed must never be used. A failed create
returns Err and no Appender, so such an appender cannot be reached.
Drop
Dropping an Appender closes (and so flushes) it, but the result is ignored
— DuckDB's own header notes that after destruction "it is no longer possible to
obtain the specific error message". Call close() explicitly whenever the
outcome matters.
Related chapters
- Reading & Writing Vectors — building the
DataChunks you append - Table Metadata — column names and
DEFAULTs - Structured Errors — the
ErrorDatareturned withduckdb-1-5 - The Entry Point — where you obtain a connection
Table Metadata
TableDescription reads the column metadata of an existing DuckDB table from
inside your extension: column names, whether a column has a DEFAULT, and,
with duckdb-1-5, the column count and types. Replacement scans, table
functions and copy functions use it to inspect a table before deciding what to
do.
No feature flag required for creating a description or reading column names
and defaults: duckdb_table_description_* has been in the frozen stable prefix
of the extension API (slots 292–297) since v1.2.0. Two accessors, column_count
and column_type, were added in DuckDB 1.5 in the unstable region and need
duckdb-1-5.
#![allow(unused)] fn main() { use quack_rs::table_description::TableDescription; use libduckdb_sys::duckdb_connection; unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { // SAFETY: `con` is a valid, open connection. let desc = unsafe { TableDescription::create(con, "main", "events") }?; assert_eq!(desc.column_name(0).as_deref(), Some("id")); assert_eq!(desc.column_has_default(0), Some(false)); Ok(()) } }
with_catalog addresses a table in another catalog, and takes None to mean
"the default" for either the catalog or the schema:
#![allow(unused)] fn main() { use quack_rs::table_description::TableDescription; use libduckdb_sys::duckdb_connection; unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { let desc = unsafe { TableDescription::with_catalog(con, Some("mydb"), None, "events") }?; let _ = desc; Ok(()) } }
API
| Method | Description |
|---|---|
TableDescription::create(con, schema, table) (unsafe) | Describe schema.table; Err if the table does not exist |
TableDescription::with_catalog(con, catalog, schema, table) (unsafe) | Describe a fully-qualified table (None = default) |
column_name(i) | Column name, or None if i is out of range (or the name is not valid UTF-8) |
column_has_default(i) | Whether the column has a DEFAULT, or None if i is out of range |
column_count() ¹ | Number of columns |
column_type(i) ¹ | Column LogicalType, or None if i is out of range |
¹ Requires the
duckdb-1-5feature flag.
Out-of-range indices return None rather than panicking, so a description can
be walked without knowing the width up front when duckdb-1-5 is off.
Related chapters
- Bulk Appender —
append_defaultfills a column without aDEFAULTwithNULL, and fails for aDEFAULTthat is not a constant (nextval(...),random()), whichcolumn_has_defaultreports astrue - Type System — what a
LogicalTypedescribes
Structured Errors
Requires the
duckdb-1-5feature flag (DuckDB 1.5.0+).
ErrorData is an RAII wrapper around DuckDB's duckdb_error_data handle, the
structured error type that several DuckDB 1.5 C API functions return. Unlike a
bare error string, an ErrorData carries both a human-readable message and a
machine-readable category (DuckDbErrorType), so your extension can branch
on the kind of failure (for example, distinguishing Io from OutOfMemory).
Expression::fold, the virtual file system,
the Arrow bridge and the appender all report
failures as an ErrorData.
Inspecting an error
#![allow(unused)] fn main() { use quack_rs::error_data::{DuckDbErrorType, ErrorData}; fn handle(err: ErrorData) { if err.has_error() { match err.error_type() { DuckDbErrorType::Io => eprintln!("I/O failure"), DuckDbErrorType::OutOfMemory => eprintln!("out of memory"), other => eprintln!("{other:?}: {}", err.message().unwrap_or_default()), } } } }
Constructing an error
Build a structured error to hand back to DuckDB (for example from a callback):
#![allow(unused)] fn main() { // ErrorData is a DuckDB object: fill the dispatch table first. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); use quack_rs::error_data::{DuckDbErrorType, ErrorData}; let err = ErrorData::new(DuckDbErrorType::InvalidInput, "row index out of range"); assert!(err.has_error()); assert_eq!(err.error_type(), DuckDbErrorType::InvalidInput); }
Propagating with ?
into_extension_error converts an ErrorData into the SDK's ExtensionError,
so a structured DuckDB error can flow through the ? operator in your
registration or callback logic:
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; use quack_rs::file_system::{FileOpenOptions, FileSystem}; use quack_rs::client_context::ClientContext; use quack_rs::error_data::ErrorData; fn read_header(ctx: &ClientContext) -> Result<(), ExtensionError> { let fs = FileSystem::from_client_context(ctx) .ok_or_else(|| ExtensionError::new("no file system"))?; let handle = fs .open(c"data.bin", &FileOpenOptions::read_only()) .map_err(ErrorData::into_extension_error)?; let _ = handle; Ok(()) } }
API
| Item | Description |
|---|---|
ErrorData::new(error_type, message) | Construct a structured error |
ErrorData::from_raw(raw) (unsafe) | Take ownership of a raw duckdb_error_data |
has_error() | true if the handle represents an actual error |
error_type() | The DuckDbErrorType category |
message() | Option<String> — the error text |
into_extension_error() | Consume into an ExtensionError |
is_null() / as_raw() / into_raw() | Handle inspection / escape hatches |
DuckDbErrorType is a #[non_exhaustive] enum mirroring duckdb_error_type
(Io, OutOfMemory, Conversion, Catalog, Constraint, Permission, …).
Unknown or future categories map to DuckDbErrorType::Invalid.
UTF-8 validation
The free function check_valid_utf8 exposes DuckDB's own UTF-8 validator. Its
rules match Rust's exactly: it accepts every Unicode scalar value and rejects
surrogates, overlong forms, code points above U+10FFFF, truncated sequences
and stray continuation bytes, just as std::str::from_utf8 does. Use it when
you want DuckDB's structured ErrorData for the failure; for a yes/no answer,
std::str::from_utf8 is equivalent and needs no database:
#![allow(unused)] fn main() { // ErrorData is a DuckDB object: fill the dispatch table first. std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); use quack_rs::error_data::check_valid_utf8; fn demo(bytes: &[u8]) { match check_valid_utf8(bytes) { Ok(()) => { /* safe to pass to DuckDB */ } Err(err) => eprintln!("invalid UTF-8: {}", err.message().unwrap_or_default()), } } assert!(check_valid_utf8("héllo".as_bytes()).is_ok()); for bad in [&b"\xff"[..], b"\xed\xa0\x80", b"\xc0\xaf", b"\xf4\x90\x80\x80", b"\xe2\x82"] { assert!(check_valid_utf8(bad).is_err()); assert!(std::str::from_utf8(bad).is_err()); } }
Ownership
ErrorData calls duckdb_destroy_error_data on drop. An ErrorData received
from a fallible 1.5 API owns its handle: let it drop, or call
into_extension_error() / into_raw() to move the data out.
Related modules
- Bound Expressions —
Expression::foldreturnsErrorData - Virtual File System — file operations return
ErrorData - Arrow Interop — the conversion functions return
ErrorData - Bulk Appender —
Appender::error_datareturnsErrorData - Error Handling — the SDK's primary
ExtensionErrortype
Bound Expressions
Requires the
duckdb-1-5feature flag (DuckDB 1.5.0+).
Expression is an RAII wrapper around DuckDB's duckdb_expression handle. You
obtain one from a scalar function's bind callback via
ScalarBindInfo::argument, which lets the bind phase inspect each argument's
static type and — when the argument is a constant — fold it to a concrete
Value.
Use it for scalar functions whose behaviour depends on a constant argument (a format string, a precision, a regex) that should be validated or pre-computed once at bind time rather than on every row.
Folding a constant argument at bind time
#![allow(unused)] fn main() { use quack_rs::scalar::{RawScalarBindInfo, ScalarBindInfo}; unsafe extern "C" fn my_bind(info: RawScalarBindInfo) { let bind = unsafe { ScalarBindInfo::new(info) }; if let Some(arg) = unsafe { bind.argument(0) } { // Inspect the argument's static return type. let _ty = arg.return_type(); // If the argument is constant, evaluate it once here instead of // recomputing it for every row in the execute callback. if arg.is_foldable() { let ctx = unsafe { bind.get_client_context() }; match arg.fold(&ctx) { Ok(value) => { // Stash `value` as bind data for the execute phase. let _ = value; } Err(err) => bind.set_error(&err.message().unwrap_or_default()), } } } } }
API
| Method | Description |
|---|---|
return_type() | Option<LogicalType> — the expression's static type |
is_foldable() | true if the expression is constant and can be folded |
fold(&client_context) | Result<Value, ErrorData> — evaluate a constant expression |
from_raw(raw) (unsafe) / as_raw() / is_null() | Handle inspection / escape hatches |
fold only succeeds when is_foldable returns true; otherwise it
returns an ErrorData of type InvalidInput. A successful fold can still
yield a SQL NULL (for example NULL::INTEGER), whose typed getters return
None.
When evaluation itself fails, DuckDB (checked on 1.5.5) hands the C API the
exception's JSON form ({"exception_type":"Conversion","exception_message":"...",...}) and always
tags it INVALID_INPUT. fold unpacks it: err.message() is the plain message
and err.error_type() the type the exception named (Conversion, OutOfRange,
...). Text that is not that JSON is passed through unchanged.
Obtaining an Expression
ScalarBindInfo (the wrapper around a scalar bind callback's duckdb_bind_info)
provides two accessors:
| Method | Returns |
|---|---|
argument(index) (unsafe) | Option<Expression> — RAII, the ergonomic path |
get_argument(index) (unsafe) | raw duckdb_expression — escape hatch |
Use argument_count() to bound the index.
Ownership
Expression calls duckdb_destroy_expression on drop. The handle returned by
argument() is owned by the caller, so the wrapper cleans it up automatically.
Related modules
- Scalar Functions — registering the function whose bind callback yields these expressions
- Values & Parameter Extraction — working
with the
Valueproduced byfold - Structured Errors — the
ErrorDatareturned on failure
Virtual File System
Requires the
duckdb-1-5feature flag (DuckDB 1.5.0+).
The file_system module exposes DuckDB's virtual file system (VFS) to your
extension, so a custom table function, replacement scan, or copy function can
read and write files through the same abstraction DuckDB uses internally. Paths
then resolve through every registered file system — httpfs (s3://,
http://) when it is loaded, in-memory files, and so on — where std::fs only
ever sees local disk.
Obtaining a FileSystem
A FileSystem comes from a ClientContext, which most function callbacks can
hand you (for example via BindInfo::get_client_context() or
ScalarBindInfo::get_client_context()).
Reading a file
#![allow(unused)] fn main() { use quack_rs::client_context::ClientContext; use quack_rs::file_system::{FileOpenOptions, FileSystem}; fn read_all(ctx: &ClientContext) -> Option<Vec<u8>> { let fs = FileSystem::from_client_context(ctx)?; let handle = fs.open(c"s3://bucket/data.csv", &FileOpenOptions::read_only()).ok()?; let mut buf = Vec::new(); handle.read_to_end(&mut buf).ok()?; Some(buf) } }
Writing a file
#![allow(unused)] fn main() { use quack_rs::client_context::ClientContext; use quack_rs::file_system::{FileOpenOptions, FileSystem}; fn write_report(ctx: &ClientContext, bytes: &[u8]) -> Result<(), quack_rs::error_data::ErrorData> { let fs = FileSystem::from_client_context(ctx).expect("file system"); let handle = fs.open(c"report.bin", &FileOpenOptions::write_create())?; handle.write_all(bytes)?; handle.sync()?; handle.close()?; Ok(()) } }
Open options and flags
FileOpenOptions describes how a file is opened. Two convenience constructors
cover the common cases; use set_flag for anything else.
| Constructor / method | Effect |
|---|---|
FileOpenOptions::read_only() | Open for reading |
FileOpenOptions::write_create() | Open for writing, creating if absent |
FileOpenOptions::new() | Empty; configure with set_flag |
set_flag(flag, value) | Set an individual FileFlag; returns true on success |
FileFlag variants: Read, Write, Create, CreateNew, Append.
Create,CreateNewandAppendneedWriteas well.CreateNew("create, failing if the file exists") also setsCreate. DuckDB maps it toFILE_FLAGS_EXCLUSIVE_CREATE, which only has that meaning together withFILE_FLAGS_FILE_CREATE; on its own it neither creates a missing file nor refuses an existing one. On Windows an existing file is not refused: DuckDB's local file system there ignores the exclusive flag and opens withOPEN_ALWAYS(checked in DuckDB 1.5.5), so the existing file is opened, untruncated, without an error.set_flag(flag, false)does not clear a flag: the C API ORs flags in and ignoresvalue. Build a freshFileOpenOptionsinstead.
FileHandle operations
| Method | Returns | Description |
|---|---|---|
read(&mut buf) | Result<usize, ErrorData> | Read up to buf.len() bytes (0 = EOF) |
read_exact(&mut buf) | Result<(), ErrorData> | Read exactly buf.len() bytes, or fail |
read_to_end(&mut vec) | Result<usize, ErrorData> | Append the rest of the file |
write(&buf) | Result<usize, ErrorData> | Write up to buf.len() bytes |
write_all(&buf) | Result<(), ErrorData> | Write all of buf, or fail |
seek(position) | Result<(), ErrorData> | Seek to an absolute byte offset |
tell() | Result<u64, ErrorData> | Current byte offset |
size() | Result<u64, ErrorData> | Total file size in bytes |
sync() | Result<(), ErrorData> | Flush buffered writes to durable storage |
close() | Result<(), ErrorData> | Close the file |
error_data() | ErrorData | Structured error from the last failed operation |
Short reads and short writes are real
duckdb_file_handle_read and duckdb_file_handle_write return "the number of
bytes actually read/written" — not a promise that the whole buffer moved. On
a local file the two almost always agree; over httpfs a short transfer is
common.
Prefer read_exact / read_to_end / write_all, which loop for you, and reach
for the raw read / write only when a partial transfer is what you want.
tell() and size() are Result for the same reason: the C API reports failure
with a negative return value, which is far too easy to clamp to zero and then
silently treat as an empty file.
FileSystem exposes open(path, options) and error_data(). Both FileSystem
and FileHandle are RAII: they are destroyed (and the handle closed) on drop.
Lifetimes
A FileSystem<'ctx> borrows the ClientContext it came from, and a
FileHandle<'fs> borrows the FileSystem that opened it. A handle refers to the
database's file system, so using it after the database is closed is a
use-after-free (valgrind: invalid read in duckdb::FileHandle::Read); the borrow
makes that a compile error instead:
#![allow(unused)] fn main() { use quack_rs::client_context::ClientContext; use quack_rs::file_system::{FileOpenOptions, FileSystem}; fn demo(ctx: &ClientContext) { let fs = FileSystem::from_client_context(ctx).unwrap(); let handle = fs.open(c"data.csv", &FileOpenOptions::read_only()).unwrap(); drop(fs); // error[E0505]: cannot move out of `fs` because it is borrowed let _ = handle.size(); } }
Related modules
- Replacement Scans —
SELECT * FROM 'file.xyz'handlers that read through this file system - Copy Functions —
COPY TOhandlers that write through it - Structured Errors — the
ErrorDatareturned on failure
Selection Vectors
Requires the
duckdb-1-5feature flag (DuckDB 1.5.0+).
A SelectionVector is a list of row indices used to logically reorder or filter
a data vector without copying its payload — the building block behind
DuckDB's zero-copy filtering. Extensions that implement custom filtering or
reordering in vectorized callbacks can allocate one, fill in the indices, and
hand its raw handle to the relevant DuckDB vector operations.
This is an advanced, low-level primitive; most extensions never need it.
Allocating and filling
#![allow(unused)] fn main() { use quack_rs::selection_vector::SelectionVector; fn demo() -> Result<(), quack_rs::error::ExtensionError> { // Select source rows 3, 1, 4, 1, 5 (in that order) — note repeats are allowed. let mut sel = SelectionVector::new(5)?; sel.as_mut_slice().copy_from_slice(&[3, 1, 4, 1, 5]); assert_eq!(sel.len(), 5); assert_eq!(sel.as_slice(), &[3, 1, 4, 1, 5]); Ok(()) } }
The indices are 32-bit (sel_t / u32) and are zeroed by new — DuckDB
itself leaves them uninitialised, so the wrapper clears them to keep
as_slice() from exposing stale heap contents. Fill them via as_mut_slice().
new returns Err for a length above selection_vector::MAX_LEN: 2^32 (the
range of sel_t) on 64-bit targets, and the largest u32 slice that fits in
isize::MAX bytes on 32-bit ones. DuckDB does not check the length itself:
a large enough length overflows its allocation-size arithmetic into a tiny
buffer, and a request of 2^48 bytes or more makes its allocator throw, which
would abort the extension.
API
| Method | Description |
|---|---|
SelectionVector::new(size) | Allocate size zeroed indices; Err above MAX_LEN |
len() / is_empty() | Number of indices |
as_slice() | &[u32] — read the indices |
as_mut_slice() | &mut [u32] — fill the indices |
as_raw() | The raw duckdb_selection_vector handle for DuckDB vector ops |
SelectionVector is RAII: it is destroyed on drop.
Related modules
- Reading & Writing Vectors — the data vectors a selection vector reorders or filters
- Complex Types — STRUCT / LIST / MAP / ARRAY vectors
Instance Cache
Requires the
duckdb-1-5feature flag (DuckDB 1.5.0+).
An InstanceCache wraps DuckDB's duckdb_instance_cache, which lets several
database handles share one underlying DuckDB instance per database path.
Opening the same path twice through the cache returns handles backed by the
same instance, which avoids the "database is already open in another
instance" conflict and saves the cost of re-initialising the database.
This is primarily useful for extensions or host integrations that open secondary databases on behalf of a query.
Opening through the cache
#![allow(unused)] fn main() { use quack_rs::instance_cache::InstanceCache; fn demo() -> Result<(), quack_rs::error::ExtensionError> { let cache = InstanceCache::new(); // Returns a duckdb_database the caller OWNS and must close with duckdb_close. let db = cache.get_or_create(c"analytics.db", None)?; let _ = db; Ok(()) } }
Pass a DbConfig to control how a freshly created instance is configured.
When an instance already exists for the path, the config must match the one it
was created with — a different one is an error ("Can't open a connection to same
database file with a different configuration than existing connections"), not
silently ignored. None means DuckDB's defaults, so it too conflicts with an
instance created with a custom config. An empty path or :memory: is never
cached: each call creates a separate in-memory database.
#![allow(unused)] fn main() { use quack_rs::instance_cache::InstanceCache; use quack_rs::config::DbConfig; fn demo() -> Result<(), quack_rs::error::ExtensionError> { let cache = InstanceCache::new(); let config = DbConfig::new()?; // configure `config` as needed... let db = cache.get_or_create(c"analytics.db", Some(&config))?; let _ = db; Ok(()) } }
API
| Method | Description |
|---|---|
InstanceCache::new() | Create a new, empty cache |
get_or_create(path, config) | Open path, creating the instance if needed |
as_raw() | The raw duckdb_instance_cache handle |
get_or_create returns Result<duckdb_database, ExtensionError>; on failure
the error carries DuckDB's message.
Threads
InstanceCache is Send + Sync, so one cache can serve several threads: DuckDB's
DBInstanceCache guards its map with a mutex and serialises creation of each
database. One ordering hazard remains, and it is DuckDB's: while the last handle
to a cached file database is being closed, a concurrent get_or_create of the same
path can fail with "Unique file handle conflict" because the closing instance still
holds the file. Keep one handle open while other threads may open the path, or
retry.
Ownership
InstanceCache is RAII and destroys the cache on drop. The duckdb_database
returned by get_or_create is, however, owned by the caller: close it with
duckdb_close when finished. The cache holds only a weak reference to each
instance. An instance stays alive while at least one handle opened through the
cache is open, and a handle stays valid after the cache itself is dropped. Once
the last handle is closed, the next get_or_create for that path creates a
fresh instance.
Related modules
config—DbConfig, the RAII configuration builder accepted here
Arrow Interop
Requires the
duckdb-1-5-4feature flag.
DuckDB's C API has a family of conversion functions (present since 1.4.4) that
move data directly between a duckdb_data_chunk and the
Arrow C Data Interface,
without a query result in between. The quack_rs::arrow module wraps all eight
of them.
No arrow crate dependency
The Arrow C Data Interface is an ABI, not a library: ArrowSchema and
ArrowArray are plain #[repr(C)] records with a release callback.
libduckdb-sys defines them directly — and asserts in its own test suite that
they match arrow-rs's FFI_ArrowSchema / FFI_ArrowArray field for field — so
quack-rs speaks Arrow without pulling in the arrow crate, and an extension
that does use arrow-rs bridges across with a pointer cast.
Why the feature is duckdb-1-5-4 and not duckdb-1-5
All eight C functions are already in duckdb_ext_api_v1 in DuckDB 1.4.4
(extension_api.hpp at v1.4.4, slots 410 to 434; they moved to 411 to 509 in
1.5.0). The floor comes from the bindings: libduckdb-sys declared both records
as opaque zero-sized bindgen placeholders (_unused: [u8; 0]) until
1.10504.0, and the caller-allocated structs these functions need cannot be
created from a zero-sized type. src/arrow.rs carries a const assertion that
fails with exactly that message against an older binding. Because the feature
implies duckdb-1-5, an extension built with it also needs a DuckDB 1.5.0+
engine.
The types
| Type | Wraps | Freed by |
|---|---|---|
ArrowOptions | duckdb_arrow_options | duckdb_destroy_arrow_options |
ArrowSchema | the ArrowSchema ABI record | release(schema) |
ArrowArray | the ArrowArray ABI record | release(array) |
ArrowConvertedSchema | duckdb_arrow_converted_schema | duckdb_destroy_arrow_converted_schema |
RawArrowSchema and RawArrowArray are the ABI records themselves, re-exported
for code that already speaks the raw interface.
Exporting a chunk
ArrowOptions carries the settings DuckDB renders Arrow with — the timezone for
TIMESTAMPTZ, the string offset width, registered extension types. Take them
from the result whose chunks you are exporting, so schema and data agree:
#![allow(unused)] fn main() { use quack_rs::arrow::{data_chunk_to_arrow, to_arrow_schema}; use quack_rs::query::QueryResult; fn demo(result: &mut QueryResult) -> Result<(), Box<dyn std::error::Error>> { // SAFETY: the connection that ran this query stays open while `options` is // in use — the options point at that connection's client context. let options = unsafe { result.arrow_options() }?; let columns: Vec<(String, quack_rs::types::LogicalType)> = (0..result.column_count()) .filter_map(|i| Some((result.column_name(i)?, result.column_logical_type(i)?))) .collect(); let pairs: Vec<(&str, &quack_rs::types::LogicalType)> = columns.iter().map(|(n, t)| (n.as_str(), t)).collect(); let schema = to_arrow_schema(&options, &pairs)?; assert_eq!(schema.format(), Some("+s")); // a record batch is a struct while let Some(chunk) = result.next_chunk()? { let array = data_chunk_to_arrow(&options, &chunk)?; // hand `array` (plus `schema`) to any Arrow consumer let _ = array; } Ok(()) } }
ArrowOptions must not outlive its connection
The options hold a raw pointer to the connection's client context, and the
conversion functions dereference it. Used after the connection is closed they
read freed memory. ArrowOptions<'conn> carries that lifetime:
ArrowOptions::from_connection(&con)is safe. It borrows theOwnedConnection, so the compiler rejects adrop(con)while the options are still in use.ArrowOptions::from_raw_connection(raw)(for a rawduckdb_connection) andQueryResult::arrow_options()/ArrowOptions::from_resultareunsafe. AQueryResultdoes not borrow the connection that ran it, so nothing checks that the connection is still open. The caller must keep it open for as long as the options are used.
Importing an array
Going the other way needs the Arrow schema translated into DuckDB's own type descriptors first. That translation is reusable — do it once, not per batch:
#![allow(unused)] fn main() { use quack_rs::arrow::{data_chunk_from_arrow, schema_from_arrow, ArrowArray, ArrowSchema}; fn demo( con: libduckdb_sys::duckdb_connection, schema: &mut ArrowSchema, array: ArrowArray, ) -> Result<(), Box<dyn std::error::Error>> { // SAFETY: `con` is a live DuckDB connection. let converted = unsafe { schema_from_arrow(con, schema) }?; // SAFETY: same connection; `array` was built against `schema`. let chunk = unsafe { data_chunk_from_arrow(con, array, &converted) }?; let _ = chunk.size(); Ok(()) } }
Note the asymmetry, which mirrors what DuckDB actually does:
schema_from_arrowborrows the schema. You still own it and it is released when itsArrowSchemadrops.data_chunk_from_arrowtakes the array by value. DuckDB setsarrow_array->release = nullptrbefore the conversion loop body — so it claims the array even when the conversion then fails. The by-value binding is still dropped on the way out, which releases the array in the one case where DuckDB does not claim it (a zero-column schema, where the loop never runs).
The resulting chunk keeps the Arrow buffers alive, so most columns share the data rather than copy it. Dictionary-encoded columns are the exception: the wrapper copies them into flat vectors.
What the wrapper refuses that DuckDB would not
duckdb_data_chunk_from_arrow indexes arrow_array->children[i] once per column
in the converted schema with no bounds check, dereferences each child without a
null check, reads offset + length rows from each child without comparing its
length, and dereferences an array without checking whether it was already
released. Those are segfaults or out-of-bounds reads, not errors.
data_chunk_from_arrow checks them first — which is why ArrowConvertedSchema
remembers the column count of the schema it was built from — and returns an
InvalidInput error instead.
It also refuses a zero-row array. DuckDB passes arrow_array->length
through as the chunk's capacity, and a capacity of zero trips a
D_ASSERT(size > 0) that aborts a debug build of DuckDB, while a release build
carries on. Skip empty batches, or create the empty chunk directly with
duckdb_create_data_chunk.
What it cannot check, and what data_chunk_from_arrow's # Safety section
therefore makes the caller's job:
- The array must conform to the converted schema. Nothing in an Arrow
array records its type, so DuckDB reads each child's buffers as the format
the schema declares. An
int32child imported under autf8schema has its values read as string offsets into a buffer that does not exist. Arrays exported withdata_chunk_to_arrowunder the schema you converted conform. - A fixed-width dictionary whose indices can be NULL needs one element of
padding past its values: DuckDB points NULL indices at an entry there,
and the copy
data_chunk_from_arrowmakes of every dictionary column reads it (docs/upstream-duckdb-reports.md, item 29). - The buffers must be as long as the lengths say, and
lengthmust be the true row count. DuckDB allocates the chunk forlengthrows before its error handling starts, so an absurd length is an allocation failure that aborts the process.
It also refuses valid Arrow layouts that DuckDB imports wrongly: it walks the
array alongside its schema and returns InvalidInput, naming the node, for
- an offset below the top level that DuckDB applies to the wrong rows: a
struct inside an offset struct or a list, a union's members, a run-end-encoded
array's value validity (
docs/upstream-duckdb-reports.md, item 24); - a dictionary with NULLs under a list that starts past element 0, or with more than 2048 rows and NULLs of its own or an enclosing struct's, which DuckDB copies past a 2048-row heap mask (items 9 and 25);
- a dictionary whose values are themselves dictionary-encoded (item 26);
- list views that overlap or leave gaps (item 27);
- a sparse union whose
+us:type codes are not0, 1, …(item 28), or whosenull_countis not 0 (item 32); - a dictionary whose
null_countis -1 ("not computed", item 31); - a
geoarrow.wkbcolumn read as more than 2048 rows (item 33); - a run-end-encoded array where DuckDB reads a plain one: a fixed-size list's child, or another run-end array's values (item 24).
Arrays that data_chunk_to_arrow produced, paired with the schema they were
produced with, never take these shapes. Arrays from other producers can: one
that slices a nested array without copying it may leave offsets below the top
level. Copying the slice before export avoids them. Every error DuckDB
reports from the conversion arrives as InvalidInput.
Round trips are not always exact
Two types come back different from an Arrow round trip through DuckDB's own converters:
TIMETZcomes back asTIMEwith the offset dropped:01:02:03+05:30returns as01:02:03.BITcomes back asBLOB.
Check the converted types (ArrowConvertedSchema) when a round trip must be
lossless.
DuckDB would export three kinds of value wrongly, with no error (checked on
1.4.4, 1.5.0 and 1.5.5), so data_chunk_to_arrow checks the chunk first and
refuses one that holds such a value, at any nesting depth:
- An
INTERVALwhose microseconds exceed about ±106,751 days (2,562,047 hours) would wrap, because Arrow counts nanoseconds in ani64and DuckDB multiplies by 1000 unchecked:INTERVAL 2562048 HOURwould export as a negative interval. - A 39-digit
UHUGEINTwould export as adecimal128(38, 0)it does not fit; from 2^127 it comes out negative (2^128 - 1becomes-1). - A 39-digit
HUGEINTwould export as adecimal128(38, 0)it does not fit, unlessarrow_lossless_conversionis set (then it exports as a 16-byte fixed-size binary and is not refused).
After the export, the array is also checked against the schema DuckDB
declares for the chunk's types. Before 1.5.5, BIGNUM (and, from 1.5.0,
GEOMETRY) exported under arrow_output_version = '1.4' is written as
binary views while the schema declares plain binary, which a consumer reads as
offsets (docs/upstream-duckdb-reports.md, item 34); such an export is
refused. Set arrow_output_version = '1.0' on those releases.
Bridging to arrow-rs
This sketch uses the arrow crate's FFI_ArrowArray, which quack-rs does not
depend on, so it is not compiled with the book:
// quack-rs -> arrow-rs
let ffi: FFI_ArrowArray = unsafe { std::mem::transmute(array.into_raw()) };
// arrow-rs -> quack-rs, neutralising the source so only one side releases
let array = unsafe { ArrowArray::take_from(std::ptr::from_mut(&mut ffi).cast()) };
take_from moves the record out and writes a released placeholder back, so the
foreign wrapper's own Drop becomes a no-op instead of a double free.
Thread safety
None of these types are Send or Sync. The Arrow C Data Interface says nothing
about which thread may call release, and duckdb_arrow_options wraps a
ClientProperties tied to the connection's client context.
TLS Configuration
Extensions that make outbound HTTPS connections (e.g., fetching remote data, calling REST APIs) need a way to inject TLS configuration — client certificates for mTLS, custom CA bundles, or restricted cipher suites.
The tls module
provides the TlsConfigProvider trait so that extensions can supply their TLS
setup through a uniform interface, regardless of which TLS library they use
(rustls, native-tls, etc.).
Design
The trait is type-erased via Arc<dyn Any + Send + Sync> so that quack-rs
does not depend on any specific TLS library. The code that consumes the provider
downcasts the returned Arc to the concrete config type, after checking
config_type_name(). TlsConfigProvider requires Send + Sync.
Implementing a TLS Provider
#![allow(unused)] fn main() { use quack_rs::tls::{TlsConfigProvider, TlsVersion}; use quack_rs::error::ExtensionError; use std::any::Any; use std::sync::Arc; struct MyTlsProvider { // In practice: Arc<rustls::ClientConfig> config: Arc<String>, mtls_enabled: bool, } impl TlsConfigProvider for MyTlsProvider { fn client_config(&self) -> Result<Arc<dyn Any + Send + Sync>, ExtensionError> { Ok(self.config.clone()) } fn provider_name(&self) -> &str { "my-extension-tls" } fn config_type_name(&self) -> &str { "String" } // in practice: "rustls::ClientConfig" fn min_tls_version(&self) -> TlsVersion { TlsVersion::Tls12 // Minimum recommended } fn supports_mtls(&self) -> bool { self.mtls_enabled } fn accepts_invalid_certs(&self) -> bool { false // MUST default to false } } }
Security Requirements
Implementations must:
- Return
falsefromaccepts_invalid_certs()unless explicitly configured otherwise by the user. Certificate validation bypass (CWE-295) should never be the default. - Return
TlsVersion::Tls12or higher frommin_tls_version(). TLS 1.0 and 1.1 are deprecated per RFC 8996. - Emit an
ExtensionWarningviaWarningCollectorwhen certificate validation is disabled or when using a TLS version below 1.2.
Auditing a Provider
audit_tls_provider() checks a provider for two common misconfigurations and
returns one ExtensionWarning per problem found:
- Certificate verification bypass (CWE-295): code
TLS_NO_VERIFY, severityHigh - A deprecated minimum version, TLS 1.0 or 1.1 (CWE-327): code
TLS_DEPRECATED_VERSION, severityMedium
Feed the result into a WarningCollector:
#![allow(unused)] fn main() { use quack_rs::error::ExtensionError; use quack_rs::tls::{audit_tls_provider, TlsConfigProvider, TlsVersion}; use quack_rs::warning::WarningCollector; use std::any::Any; use std::sync::Arc; struct InsecureProvider; impl TlsConfigProvider for InsecureProvider { fn client_config(&self) -> Result<Arc<dyn Any + Send + Sync>, ExtensionError> { Ok(Arc::new(())) } fn provider_name(&self) -> &str { "insecure" } fn config_type_name(&self) -> &str { "()" } fn min_tls_version(&self) -> TlsVersion { TlsVersion::Tls10 } fn supports_mtls(&self) -> bool { false } fn accepts_invalid_certs(&self) -> bool { true } } let collector = WarningCollector::new(); for w in audit_tls_provider(&InsecureProvider) { collector.emit(w); } let codes: Vec<&str> = collector.snapshot().iter().map(|w| w.code).collect(); assert_eq!(codes, ["TLS_NO_VERIFY", "TLS_DEPRECATED_VERSION"]); }
Downcasting Safely
Never use .unwrap() or .expect() when downcasting in FFI callback contexts
(see Pitfall L3).
Handle a failed downcast as an error:
#![allow(unused)] fn main() { use std::any::Any; use std::sync::Arc; use quack_rs::error::ExtensionError; // With rustls this would be `downcast_ref::<rustls::ClientConfig>()`. fn use_config(config: &Arc<dyn Any + Send + Sync>) -> Result<(), ExtensionError> { let text = config .downcast_ref::<String>() .ok_or_else(|| ExtensionError::new("expected a String TLS config"))?; let _ = text; Ok(()) } let config: Arc<dyn Any + Send + Sync> = Arc::new(String::from("pem bundle")); assert!(use_config(&config).is_ok()); let wrong: Arc<dyn Any + Send + Sync> = Arc::new(42_u32); assert!(use_config(&wrong).is_err()); }
Structured Warnings
Extensions that access external resources (network, files, credentials) should
emit structured warnings when potentially unsafe operations occur. The
warning module
provides a consistent, thread-safe API for collecting and surfacing security
warnings.
Core Types
ExtensionWarning
A structured warning with:
| Field | Type | Description |
|---|---|---|
code | &'static str | Machine-readable code (e.g., "TLS_NO_VERIFY") |
severity | WarningSeverity | Info / Low / Medium / High / Critical |
message | String | Human-readable description |
cwe | Option<u32> | Optional CWE identifier |
WarningSeverity
Five levels mirroring common security advisory severity. WarningSeverity
implements Ord in this order, so severity >= WarningSeverity::High selects
the warnings that need attention:
- Info — no security impact, but worth noting
- Low — minimal security impact
- Medium — potential security concern
- High — significant security risk
- Critical — immediate action recommended
WarningCollector
A thread-safe collector backed by Mutex<Vec<ExtensionWarning>>. Share it
across threads via Arc<WarningCollector>. A panic on another thread while it
held the lock does not lose warnings: the collector recovers the list from the
poisoned lock and keeps using it.
Usage
#![allow(unused)] fn main() { use quack_rs::warning::{ExtensionWarning, WarningSeverity, WarningCollector}; let collector = WarningCollector::new(); // Emit a warning when detecting an insecure configuration collector.emit(ExtensionWarning { code: "TLS_NO_VERIFY", severity: WarningSeverity::High, message: "TLS certificate verification is disabled".into(), cwe: Some(295), }); // Check warnings assert_eq!(collector.len(), 1); assert!(!collector.is_empty()); // Read without clearing let snapshot = collector.snapshot(); assert_eq!(snapshot.len(), 1); assert_eq!(collector.len(), 1); // still there // Consume all warnings let warnings = collector.drain(); assert_eq!(warnings.len(), 1); assert!(collector.is_empty()); // now empty }
Display Format
ExtensionWarning implements Display as [SEVERITY] CODE: message (CWE-nnn);
the CWE suffix is omitted when cwe is None:
[HIGH] TLS_NO_VERIFY: TLS certificate verification is disabled (CWE-295)
[MEDIUM] TLS_DEPRECATED_VERSION: TLS provider "my-tls" allows deprecated TLS 1.0 (RFC 8996) (CWE-327)
Integration with TLS Auditing
tls::audit_tls_provider() returns a Vec<ExtensionWarning> that can be fed
straight into a WarningCollector; see
Auditing a Provider for a complete example.
Best Practices
- Create a single
WarningCollectorper extension (typically in global or bind-data state) - Use
snapshot()for read-only diagnostics; usedrain()when consuming warnings for output - Include a CWE identifier whenever one applies
- Surface collected warnings through a table function of your own
(for example
SELECT * FROM __extension_warnings())
Secrets Management
Extensions that access external services (HTTP APIs, databases, cloud storage)
commonly need credentials. DuckDB has a native secrets system (CREATE SECRET),
but the extension C API exposes none of it: duckdb_ext_api_v1 has no
duckdb_secret_* functions, so an extension cannot ask DuckDB for a credential.
The secrets
module therefore offers two separate things:
list_duckdb_secretsreads the metadata of the secrets the user has configured, with sensitive fields redacted by DuckDB.SecretsManagerandSecretEntryare a trait and a type for the credential source an extension has to provide itself (an environment variable, a config option, a file, its own key store), with leak-resistant defaults already in place.
Listing DuckDB's secrets
list_duckdb_secrets queries the duckdb_secrets() table function and returns
one DuckDbSecretInfo per secret: name, secret_type, provider,
persistent, storage, scope (the URI prefixes it applies to) and
secret_string. DuckDB redacts sensitive fields in that table, so a secret
created with SECRET 'super-secret-value' comes back as ...;secret=redacted.
Use it to find out which secrets exist and what they cover, not to
authenticate:
#![allow(unused)] fn main() { use quack_rs::secrets::list_duckdb_secrets; use libduckdb_sys::duckdb_connection; unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { // SAFETY: `con` is a valid, open connection. let secrets = unsafe { list_duckdb_secrets(con) }?; if !secrets.iter().any(|s| s.secret_type == "s3") { eprintln!("no S3 secret configured; run CREATE SECRET (TYPE s3, ...)"); } Ok(()) } }
Core Types
SecretEntry
A single secret entry with metadata and key-value fields. Designed to minimize accidental credential leakage:
Debugredacts field values and the scope — field keys are shown, and every value and a non-empty scope are replaced with"[REDACTED]"Dropzeroizes sensitive data — every field key and value, the provider and the scope are overwritten with zeros usingstd::ptr::write_volatilebefore deallocation. This covers the buffers aSecretEntryowns, not aStringthe caller passed in and still holds- No
PartialEq— prevents accidental non-constant-time comparisons of secret material Cloneduplicates the secret — it is supported, but each clone is another copy of the credential in memory
SecretsManager
The trait an extension implements over its own credential source. It requires
Send + Sync:
#![allow(unused)] fn main() { use quack_rs::secrets::{SecretEntry, SecretsManager}; struct MySecrets { entries: Vec<SecretEntry>, } impl SecretsManager for MySecrets { fn get_secret(&self, name: &str, secret_type: &str) -> Option<SecretEntry> { self.entries.iter() .find(|e| e.name() == name && e.secret_type() == secret_type) .cloned() } fn list_secrets(&self, secret_type: Option<&str>) -> Vec<SecretEntry> { self.entries.iter() .filter(|e| secret_type.is_none() || secret_type == Some(e.secret_type())) .cloned() .collect() } fn remove_secret(&self, _name: &str, _secret_type: &str) -> bool { false // read-only example } } }
Building Secret Entries
Use the builder pattern:
#![allow(unused)] fn main() { use quack_rs::secrets::SecretEntry; let entry = SecretEntry::new("my_api_key", "bearer") .with_provider("config") .with_scope("https://api.example.com") .with_field("token", "sk-abc123") .with_field("refresh_token", "xyz789"); assert_eq!(entry.name(), "my_api_key"); assert_eq!(entry.secret_type(), "bearer"); assert_eq!(entry.get_field("token"), Some("sk-abc123")); }
Safe Diagnostics
Use field_keys() for logging without leaking secrets:
#![allow(unused)] fn main() { use quack_rs::secrets::SecretEntry; let entry = SecretEntry::new("key", "s3") .with_field("access_key", "AKIA...") .with_field("secret_key", "wJalr..."); // Safe for logging — returns keys only, no values let mut keys = entry.field_keys(); keys.sort_unstable(); // the fields are a HashMap, so the order is unspecified assert_eq!(keys, ["access_key", "secret_key"]); }
Debug Output
The Debug implementation redacts field values and the scope. For an entry
with a scope and one field, {:#?} prints:
SecretEntry {
name: "api_key",
secret_type: "bearer",
provider: "config",
scope: "[REDACTED]",
fields: {
"token": "[REDACTED]",
},
}
Security Best Practices
- Never log secret field values — use
field_keys()for diagnostics - Drop clones promptly — minimize the window during which sensitive data resides in memory
- Implement
remove_secretwith zeroization — don't just remove the reference; zeroize the data before deallocation - Thread safety —
SecretsManagerrequiresSend + Sync, because DuckDB may run the callbacks that consult it on several threads at once
Testing Guide
This page covers how to test a DuckDB extension written with quack-rs. The
strategy has two tiers: pure-Rust unit tests for business logic (no DuckDB
required), and SQLLogicTest end-to-end (E2E) tests that load the packaged
extension into a real DuckDB process. Between the two, the InMemoryDb helper
lets cargo test register and call your real callbacks against a bundled DuckDB.
Architectural limitation: the loadable-extension dispatch wall
This is the most important thing to understand before writing tests.
DuckDB loadable extensions use libduckdb-sys with
features = ["loadable-extension"]. This intentionally does not link the
DuckDB runtime into the extension binary. Instead, every DuckDB C API call
(duckdb_vector_get_data, duckdb_create_logical_type, etc.) goes through a
dispatch table: one global AtomicPtr per C API function, filled in only when
the extension's entry point calls duckdb_rs_extension_api_init as DuckDB
loads it.
In cargo test, no DuckDB process loads your extension. The dispatch table
is never initialized, and the first call to any DuckDB C API function panics:
DuckDB API not initialized or DuckDB feature omitted
Opening an InMemoryDb
fills the table for the whole test process, after which the APIs below work.
What this breaks
| API | Why it fails |
|---|---|
VectorReader::new | calls duckdb_vector_get_data |
VectorWriter::new | calls duckdb_vector_get_data |
Connection::register_* | calls the DuckDB registration C API |
LogicalType::new | calls duckdb_create_logical_type |
LogicalType::drop | calls duckdb_destroy_logical_type |
BindInfo::add_result_column | calls duckdb_bind_add_result_column |
What still works in cargo test
| API | Why it works |
|---|---|
AggregateTestHarness | pure Rust, zero DuckDB dependency |
MockVectorWriter / MockVectorReader | in-memory buffers, zero DuckDB dependency |
MockRegistrar | records registrations without calling the C API |
SqlMacro::to_sql() | generates SQL strings, no DuckDB needed |
interval_to_micros | pure arithmetic |
validate / scaffold | pure Rust |
InMemoryDb | links a real DuckDB via the duckdb crate (bundled-test or bundled-test-prebuilt feature) |
Mock types for callback logic
Keep the per-row computation in a plain Rust function and test that function
directly — it needs no vectors at all. The FFI callback is then a thin loop
around it, and for the common shapes the typed constructors
(ScalarFunctionBuilder::map1, map1_str, …) write that loop for you, NULL
handling included.
#![allow(unused)] fn main() { // The logic: plain Rust, tested with plain `#[test]`s. fn shout(s: &str) -> String { s.to_uppercase() } #[test] fn shout_uppercases() { assert_eq!(shout("hello"), "HELLO"); } }
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::vector::{VectorReader, VectorWriter}; fn shout(s: &str) -> String { s.to_uppercase() } // The callback: a thin loop over the real vectors. unsafe extern "C" fn shout_callback( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let rows = usize::try_from(unsafe { libduckdb_sys::duckdb_data_chunk_get_size(input) }) .unwrap_or(0); let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..rows { if unsafe { reader.is_valid(row) } { let out = shout(unsafe { reader.read_str(row) }); unsafe { writer.write_varchar(row, &out) }; } else { unsafe { writer.set_null(row) }; } } } }
To test the loop itself against real vectors, use InMemoryDb (below): it
runs the callback inside a real DuckDB.
MockVectorReader and MockVectorWriter are in-memory stand-ins with the
same method names as VectorReader and VectorWriter. They are separate
types, so a function written against the mocks cannot be handed the real
reader and writer; they are for prototyping and checking row-loop logic
without a database. They do reproduce the behaviour of a real vector that a
more forgiving mock would hide:
set_nullclears a validity bit and a laterwrite_*does not set it again (a real vector keeps returning NULL for that row);- a row that is never written is valid, not NULL — use
is_writtento check a loop wrote every row; - writing past the capacity given to
MockVectorWriter::newpanics.
#![allow(unused)] fn main() { use quack_rs::testing::{MockVectorReader, MockVectorWriter}; let reader = MockVectorReader::from_strs([Some("hello"), None, Some("world")]); let mut writer = MockVectorWriter::new(3); for i in 0..reader.row_count() { match reader.try_get_str(i) { Some(s) => writer.write_varchar(i, &s.to_uppercase()), None => writer.set_null(i), } } assert_eq!(writer.try_get_str(0), Some("HELLO")); assert!(writer.is_null(1)); assert_eq!(writer.try_get_str(2), Some("WORLD")); assert!((0..3).all(|i| writer.is_written(i) || writer.is_null(i))); }
Testing registration with MockRegistrar
MockRegistrar implements the Registrar trait without calling any DuckDB C API.
Use it to verify your registration function registers the right set of functions:
#![allow(unused)] fn main() { use quack_rs::connection::Registrar; use quack_rs::testing::MockRegistrar; use quack_rs::scalar::ScalarFunctionBuilder; use quack_rs::types::TypeId; use quack_rs::error::ExtensionError; use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; unsafe extern "C" fn upper(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} unsafe extern "C" fn lower(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} fn register_all(reg: &impl Registrar) -> Result<(), ExtensionError> { let upper = ScalarFunctionBuilder::new("upper_ext") .param(TypeId::Varchar) .returns(TypeId::Varchar) .function(upper); let lower = ScalarFunctionBuilder::new("lower_ext") .param(TypeId::Varchar) .returns(TypeId::Varchar) .function(lower); unsafe { reg.register_scalar(upper)?; reg.register_scalar(lower)?; } Ok(()) } #[test] fn test_register_all() { let mock = MockRegistrar::new(); register_all(&mock).unwrap(); assert_eq!(mock.total_registrations(), 2); assert!(mock.has_scalar("upper_ext")); assert!(mock.has_scalar("lower_ext")); } }
MockRegistrar refuses, with the same error, what the real registration
refuses before it calls DuckDB: a missing return type or callback, an empty
function set, a copy function with neither direction, a config option without a
type or default, a composite or literal TypeId in any slot, an ANY return
type. Checks that need DuckDB (a name or signature already taken, a type the
running DuckDB lacks, a config default that does not convert) are not run.
Limitation:
MockRegistrarcannot be used with builders that holdLogicalTypevalues (created via.returns_logical()or.param_logical()), becauseLogicalType::dropcallsduckdb_destroy_logical_type, which panics while the dispatch table is uninitialised. UseTypeIdparameters withMockRegistrar.
SQL-level testing with InMemoryDb (bundled-test feature)
For SQL-level assertions — verifying that a SQL macro produces the correct output,
or that a CREATE TABLE + INSERT + SELECT pipeline works — enable the bundled-test
Cargo feature. This provides InMemoryDb, which wraps the duckdb crate's bundled
DuckDB and automatically initialises the loadable-extension dispatch table before
opening a connection (see Pitfall P9).
Two features expose InMemoryDb; pick the one that fits your build-time budget:
# Zero-config but slow: compile libduckdb from C++ source (~5–10 min cold).
[dev-dependencies]
quack-rs = { version = "0.18", features = ["bundled-test"] }
# Fast: link against a pre-built libduckdb. Set DUCKDB_DOWNLOAD_LIB=1 at build
# time and libduckdb-sys downloads the upstream release zip (~40 MB, cached
# under target/); or set DUCKDB_LIB_DIR=/path/to/libduckdb if you already have
# one extracted. Header discovery needs libduckdb-sys >= 1.10503; with an older
# one, also set DUCKDB_INCLUDE_DIR.
[dev-dependencies]
quack-rs = { version = "0.18", features = ["bundled-test-prebuilt"] }
Both keep duckdb out of a plain cargo test and out of your published
crate's dependency tree — it is pulled in only when one of these features is on.
If your tests need to LOAD your own locally built .duckdb_extension
file, use InMemoryDb::open_unsigned instead of open(): the
allow_unsigned_extensions option can only be set at startup, not with SET
after the database is running.
#![allow(unused)] fn main() { use quack_rs::testing::InMemoryDb; use quack_rs::sql_macro::SqlMacro; #[test] fn test_clamp_macro_sql() { let db = InMemoryDb::open().unwrap(); // Generate and execute the CREATE MACRO SQL let m = SqlMacro::scalar("clamp", &["x", "lo", "hi"], "greatest(lo, least(hi, x))").unwrap(); db.execute_batch(&m.to_sql()).unwrap(); // Verify correct output let result: i64 = db.query_one("SELECT clamp(5, 1, 10)").unwrap(); assert_eq!(result, 5); let clamped: i64 = db.query_one("SELECT clamp(15, 1, 10)").unwrap(); assert_eq!(clamped, 10); } }
Testing your FFI callbacks for real
Opening an InMemoryDb also populates the loadable-extension dispatch table —
for the whole process, not just that handle. After that the entire C API works,
so you can register a real function and call it from SQL inside cargo test:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_connection, DuckDBSuccess}; use quack_rs::data_chunk::DataChunk; use quack_rs::query::query; use quack_rs::scalar::ScalarFunctionBuilder; use quack_rs::testing::InMemoryDb; use quack_rs::types::TypeId; use quack_rs::vector::VectorWriter; quack_rs::scalar_callback!(triple_it, |_info, input, output| { let chunk = unsafe { DataChunk::from_raw(input) }; let reader = unsafe { chunk.reader(0) }; let mut writer = unsafe { VectorWriter::from_vector(output) }; for row in 0..chunk.size() { unsafe { writer.write_i64(row, reader.read_i64(row) * 3) }; } }); #[test] fn triple_it_works() { // 1. Initialise the dispatch table. let _dispatch = InMemoryDb::open().unwrap(); // 2. Open a raw connection. let mut db = std::ptr::null_mut(); let mut con: duckdb_connection = std::ptr::null_mut(); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), DuckDBSuccess); } // 3. Register with the usual builder. unsafe { ScalarFunctionBuilder::try_new("triple_it").unwrap() .param(TypeId::BigInt) .returns(TypeId::BigInt) .function(triple_it) .register(con) .unwrap(); } // 4. Run SQL and assert on the answer. let mut result = unsafe { query(con, "SELECT triple_it(14)") }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); assert_eq!(unsafe { chunk.reader(0).read_i64(0) }, 42); } }
This is the most valuable coverage available inside cargo test: it exercises
the builder, DuckDB's planner, your extern "C" callback, and the vector
accessors' pointer arithmetic in one go. tests/ffi_roundtrip.rs in the quack-rs
repository does this for every vector type.
The mocks are still the right tool for unit-testing callback logic without a
database, and are the only option when the bundled-test features are off.
Why two tiers?
Pitfall P3 — Unit tests are insufficient. 435 unit tests passed in duckdb-behavioral while the extension had three critical bugs: a SEGFAULT on load, 6 of 7 functions not registering, and wrong results from a combine bug. E2E tests caught all three.
| Test tier | What it catches | What it misses |
|---|---|---|
| Unit tests | Logic bugs in state structs | FFI wiring, registration failures, SEGFAULT |
| E2E tests | FFI wiring, registration, load-time crashes, wrong results | Only the inputs you did not write a test for |
Both tiers are required. Unit tests give fast, deterministic feedback. E2E tests prove the extension actually works inside DuckDB.
Unit tests with AggregateTestHarness
AggregateTestHarness<S> simulates the DuckDB aggregate lifecycle in pure Rust
without any DuckDB dependency:
flowchart LR
N["new()"] --> U["update() × N"]
U --> C["combine() <i>(optional)</i>"]
C --> F["finalize()"]
Basic usage
#![allow(unused)] fn main() { use quack_rs::testing::AggregateTestHarness; use quack_rs::aggregate::AggregateState; #[derive(Default, Debug, PartialEq)] struct SumState { total: i64 } impl AggregateState for SumState {} #[test] fn test_sum() { let mut h = AggregateTestHarness::<SumState>::new(); h.update(|s| s.total += 10); h.update(|s| s.total += 20); h.update(|s| s.total += 5); assert_eq!(h.finalize().total, 35); } }
Convenience: aggregate
For testing over a collection of inputs:
#![allow(unused)] fn main() { use quack_rs::aggregate::AggregateState; use quack_rs::testing::AggregateTestHarness; #[derive(Default)] struct WordCountState { count: usize } impl AggregateState for WordCountState {} fn count_words(s: &str) -> usize { s.split_whitespace().count() } #[test] fn test_word_count() { let result = AggregateTestHarness::<WordCountState>::aggregate( ["hello world", "one", "two three four", ""], |s, text| s.count += count_words(text), ); assert_eq!(result.count, 6); // 2 + 1 + 3 + 0 } }
Testing combine (Pitfall L1)
DuckDB creates fresh target states — set up by state_init, which with
FfiState<T> means T::default() — and calls combine to merge into them.
combine must propagate all fields, including configuration fields, not
just accumulated data. Test this explicitly:
#![allow(unused)] fn main() { use quack_rs::aggregate::AggregateState; use quack_rs::testing::AggregateTestHarness; #[derive(Default)] struct MyState { window_size: i64, count: i64 } impl AggregateState for MyState {} #[test] fn combine_propagates_config() { let mut h1 = AggregateTestHarness::<MyState>::new(); h1.update(|s| { s.window_size = 3600; // config field s.count += 5; // data field }); // h2 simulates a fresh target state: `state_init` gave it `MyState::default()` let mut h2 = AggregateTestHarness::<MyState>::new(); h2.combine(&h1, |src, tgt| { tgt.window_size = src.window_size; // MUST propagate config tgt.count += src.count; }); let result = h2.finalize(); assert_eq!(result.window_size, 3600); // Would be 0 if forgotten assert_eq!(result.count, 5); } }
Inspecting intermediate state
#![allow(unused)] fn main() { use quack_rs::aggregate::AggregateState; use quack_rs::testing::AggregateTestHarness; #[derive(Default)] struct SumState { total: i64 } impl AggregateState for SumState {} let mut h = AggregateTestHarness::<SumState>::new(); h.update(|s| s.total += 5); assert_eq!(h.state().total, 5); // borrow without consuming h.update(|s| s.total += 3); assert_eq!(h.state().total, 8); }
Resetting
#![allow(unused)] fn main() { use quack_rs::aggregate::AggregateState; use quack_rs::testing::AggregateTestHarness; #[derive(Default)] struct SumState { total: i64 } impl AggregateState for SumState {} let mut h = AggregateTestHarness::<SumState>::new(); h.update(|s| s.total = 999); h.reset(); assert_eq!(h.state().total, 0); // back to S::default() }
Pre-populating state
#![allow(unused)] fn main() { use quack_rs::aggregate::AggregateState; use quack_rs::testing::AggregateTestHarness; #[derive(Default)] struct MyState { window_size: i64, count: i64 } impl AggregateState for MyState {} let initial = MyState { window_size: 3600, count: 0 }; let h = AggregateTestHarness::with_state(initial); assert_eq!(h.finalize().window_size, 3600); }
Unit tests for scalar functions
Scalar logic is pure Rust — test it directly:
#![allow(unused)] fn main() { // From examples/hello-ext/src/lib.rs — scalar function logic pub fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } #[test] fn first_word_basic() { assert_eq!(first_word("hello world"), "hello"); assert_eq!(first_word(" padded "), "padded"); assert_eq!(first_word(""), ""); assert_eq!(first_word(" "), ""); } }
Unit tests for SQL macros
SqlMacro::to_sql() is pure Rust — no DuckDB connection needed:
#![allow(unused)] fn main() { use quack_rs::sql_macro::SqlMacro; #[test] fn scalar_macro_sql() { let m = SqlMacro::scalar("double_it", &["x"], "x * 2").unwrap(); assert_eq!(m.to_sql(), r#"CREATE OR REPLACE MACRO "double_it"("x") AS (x * 2)"#); } #[test] fn table_macro_sql() { let m = SqlMacro::table("recent", &["n"], "SELECT * FROM events LIMIT n").unwrap(); assert_eq!(m.to_sql(), r#"CREATE OR REPLACE MACRO "recent"("n") AS TABLE SELECT * FROM events LIMIT n"#); } }
E2E testing with SQLLogicTest
Community extensions are tested using DuckDB's SQLLogicTest format, which runs SQL directly in DuckDB and compares the output line by line.
File location
test/sql/my_extension.test
Format
# my_extension tests
require my_extension
statement ok
LOAD my_extension;
query I
SELECT my_function('hello world');
----
2
Directives:
| Directive | Meaning |
|---|---|
require | Load the extension; skip the file if it is not available |
statement ok | SQL must succeed |
statement error | SQL must fail |
query I | Query returning one INTEGER column |
query II | Query returning two INTEGER columns |
query T | Query returning one TEXT column |
---- | Expected output follows |
Installing DuckDB (1.4.x or 1.5.x)
E2E testing needs the DuckDB CLI. Download it with curl; no system package
manager is needed. Every 1.4.x and 1.5.x release loads an extension stamped with
C API version v1.2.0 (1.4.4 through 1.5.5 declare that version; 1.5.6 declares
v1.5.6 and accepts every earlier one). The quack-rs CI extension-load job
loads its example extension into 1.4.4, 1.5.0, 1.5.5 and the latest release
(currently 1.5.6). Develop against the current release, 1.5.6:
# DuckDB 1.5.6 (current release)
curl -fsSL https://github.com/duckdb/duckdb/releases/download/v1.5.6/duckdb_cli-linux-amd64.zip \
-o /tmp/duckdb.zip \
&& unzip -o /tmp/duckdb.zip -d /tmp/ \
&& chmod +x /tmp/duckdb \
&& /tmp/duckdb --version
# → v1.5.6
To match a pinned CI engine instead, substitute v1.4.4, v1.5.0 or v1.5.5
in the URL.
For macOS, replace linux-amd64 with osx-universal. For Windows, use
windows-amd64 and unzip to a directory on %PATH%.
Running E2E tests
# Build the extension
cargo build --release
# Append the metadata footer DuckDB's loader requires. append_metadata ships
# with quack-rs: cargo install quack-rs --bin append_metadata
append_metadata \
target/release/libmy_extension.so \
/tmp/my_extension.duckdb_extension \
--abi-type C_STRUCT \
--extension-version v0.1.0 \
--duckdb-version v1.2.0 \
--platform linux_amd64
# Load it in the DuckDB CLI (-unsigned allows an unsigned extension)
/tmp/duckdb -unsigned -c "
LOAD '/tmp/my_extension.duckdb_extension';
SELECT my_function('hello world');
"
The community extension CI runs these SQLLogicTest files automatically. Give each function at least one test, covering NULL, empty and typical input:
# Test NULL handling
query I
SELECT my_function(NULL);
----
NULL
# Test empty input
query I
SELECT my_function('');
----
0
# Test normal case
query I
SELECT my_function('hello world');
----
2
Pitfall P5 — SQLLogicTest does exact string matching. Copy expected values directly from DuckDB CLI output. NULL is represented as
NULL(uppercase). Floats must match to the number of decimal places DuckDB outputs.
Property-based testing with proptest
The proptest crate checks a property over arbitrary inputs, which suits
arithmetic and aggregate logic:
#![allow(unused)] fn main() { use quack_rs::interval::{interval_to_micros_saturating, DuckInterval}; use proptest::prelude::*; proptest! { #[test] fn saturating_never_panics(months: i32, days: i32, micros: i64) { let iv = DuckInterval { months, days, micros }; // Must not panic for any input let _ = interval_to_micros_saturating(iv); } } }
quack-rs's own test suite uses proptest for interval conversion and
AggregateTestHarness properties.
What to test
| Scenario | Unit | E2E |
|---|---|---|
| NULL input → NULL output | ✓ | |
| Empty string | ✓ | ✓ |
| Unicode strings | ✓ | |
| Numeric edge cases (0, MAX, MIN) | ✓ | |
| Combine propagates config | ✓ | |
| Multi-group aggregation | ✓ | |
| Function registration success | ✓ | |
| Extension loads without crash | ✓ | |
| SQL macro produces correct output | ✓ (to_sql) | ✓ |
Dev dependencies
[dependencies]
quack-rs = "0.18"
[dev-dependencies]
proptest = "1"
# Only for InMemoryDb; enables the feature for test builds alone.
quack-rs = { version = "0.18", features = ["bundled-test"] }
The testing module is compiled unconditionally (not #[cfg(test)]), so crates
that depend on quack-rs can use it in their own tests. InMemoryDb additionally
needs the bundled-test or bundled-test-prebuilt feature.
Community Extensions
DuckDB's community extensions repository lets anyone publish a loadable
extension that DuckDB users install with INSTALL … FROM community. This page
covers scaffolding, description.yml, naming, versioning, platforms, build
settings and submission for a community extension built with quack-rs.
Prerequisites
- A working extension that passes local E2E tests
- A GitHub repository (the community build runs from it)
- All functions tested with SQLLogicTest format
- A globally unique extension name
Scaffolding a new project
quack_rs::scaffold::generate_scaffold generates the project files in one call.
It returns paths relative to the project root and writes nothing itself, so join
them under the project directory — never write them
relative to the current directory, which would overwrite its Cargo.toml and
src/lib.rs:
#![allow(unused)] fn main() { use quack_rs::scaffold::{ScaffoldConfig, generate_scaffold}; use std::path::Path; let config = ScaffoldConfig { name: "my_extension".to_string(), description: "Does something useful".to_string(), version: "0.1.0".to_string(), license: "MIT".to_string(), maintainer: "Your Name".to_string(), github_repo: "yourorg/duckdb-my-extension".to_string(), excluded_platforms: vec![], // `target_duckdb_version`, `use_unstable_c_api` and `git_ref` default to the // stable-ABI settings; see `concepts/abi.md` for when to change them. ..ScaffoldConfig::default() }; let files = generate_scaffold(&config).expect("scaffold failed"); let root = Path::new(&config.name); // ./my_extension/ for file in &files { let path = root.join(&file.path); std::fs::create_dir_all(path.parent().unwrap()).unwrap(); std::fs::write(&path, &file.content).unwrap(); } }
In a new repository, add the build tooling submodule once with
git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools
(git submodule update --init does nothing until then — Pitfall P4).
This generates:
my_extension/
├── Cargo.toml
├── Makefile
├── extension_config.cmake
├── src/lib.rs
├── src/wasm_lib.rs
├── description.yml
├── test/sql/my_extension.test
├── .github/workflows/extension-ci.yml
├── .gitmodules
├── .gitignore
└── .cargo/config.toml
description.yml
The file the scaffold generates, with the fields a community submission uses:
extension:
name: my_extension
description: One-line description of what your extension does
version: 0.1.0
language: Rust
build: cargo
license: MIT
requires_toolchains: rust;python3
# excluded_platforms: "wasm_mvp;wasm_eh;wasm_threads" # optional
maintainers:
- Your Name
repo:
github: yourorg/duckdb-my-extension
# Must be a commit hash, not a branch: the community repository builds
# exactly this revision and signs the result, so a moving reference would
# make the build unreproducible. DuckDB's docs: "Provide the hash of the
# latest commit on the branch targeting stable as `ref`".
ref: 0a11ddc058beb2d480ccbfa83e16a68400c5d076
# ref_next: <hash> # optional: a revision compatible with DuckDB main,
# # used while a new DuckDB release is being prepared.
docs:
hello_world: |
SELECT my_extension_hello('world');
extended_description: |
A longer description, rendered on the community-extensions site.
Use quack_rs::validate to pre-validate fields before submission:
use quack_rs::validate::{ validate_extension_name, validate_extension_version, validate_spdx_license, validate_excluded_platforms_str, }; fn main() -> Result<(), quack_rs::error::ExtensionError> { validate_extension_name("my_extension")?; validate_extension_version("0.1.0")?; validate_spdx_license("MIT")?; validate_excluded_platforms_str("wasm_mvp;wasm_eh")?; Ok(()) }
Naming rules
Extension names must satisfy all of the following:
- Match
^[a-z][a-z0-9_]*$(lowercase, digits, underscores — no hyphens, because DuckDB looks up the entry point as<name>_init_c_api) - Not exceed 64 characters
- Be globally unique across the entire DuckDB community extensions ecosystem
Check existing names at community-extensions.duckdb.org before choosing. Use vendor-prefixed names to avoid collisions:
myorg_analytics ✓
analytics ✗ (likely taken or too generic)
Pitfall P1 — The
[lib] nameinCargo.tomlMUST exactly match the extension name. If your crate name isduckdb-my-ext(producinglibduckdb_my_ext.so) butdescription.ymlsaysname: my_ext, the community build fails withFileNotFoundError.
Versioning
| Format | Example | Meaning |
|---|---|---|
| 7–40 lowercase hex chars (a git hash) | 690bfc5 | Unstable — no guarantees |
0.y.z | 0.1.0 | Pre-release — working toward stability |
x.y.z (x > 0) | 1.0.0 | Stable — full semver guarantees |
validate_extension_version is deliberately permissive: it accepts these three
formats and anything else made of [A-Za-z0-9._+-], such as the date-based build ids
some published extensions use. classify_extension_version accepts only the three
formats above and returns the stability tier:
use quack_rs::validate::semver::{classify_extension_version, ExtensionStability}; fn main() -> Result<(), quack_rs::error::ExtensionError> { // Returns the tier and the version string it classified. let (stability, _version) = classify_extension_version("0.1.0")?; match stability { ExtensionStability::Unstable => println!("git hash"), ExtensionStability::PreRelease => println!("0.y.z"), ExtensionStability::Stable => println!("x.y.z, x>0"), } Ok(()) }
Platform targets
Community extensions are built for:
| Platform | Description | Opt-in? |
|---|---|---|
linux_amd64 | Linux x86_64 (glibc) | |
linux_amd64_musl | Linux x86_64 (musl) | yes |
linux_arm64 | Linux AArch64 (glibc) | |
linux_arm64_musl | Linux AArch64 (musl) | yes |
osx_amd64 | macOS x86_64 | |
osx_arm64 | macOS Apple Silicon | |
windows_amd64 | Windows x86_64 | |
windows_amd64_mingw | Windows x86_64 (MinGW) | |
windows_arm64 | Windows AArch64 | yes |
wasm_mvp | WebAssembly (MVP) | |
wasm_eh | WebAssembly (exception handling) | |
wasm_threads | WebAssembly (threads) |
An opt-in platform is not built unless an extension asks for it, so listing
one in excluded_platforms has no effect. validate::platform::is_opt_in_platform
reports which these are.
linux_amd64_gcc4 used to appear in this table and no longer exists: DuckDB
retired the legacy CXX ABI target, and DuckDBPlatform() now raises a compile
error rather than emitting a _gcc4 suffix. validate_platform rejects it with
that explanation.
This table is derived from config/distribution_matrix.json in
duckdb/extension-ci-tools, and scripts/check-platform-table.py
fails CI when quack-rs's copy drifts from it.
If your extension cannot be built for a platform (e.g., it uses a
platform-specific system library), add it to excluded_platforms:
#![allow(unused)] fn main() { use quack_rs::scaffold::ScaffoldConfig; let config = ScaffoldConfig { excluded_platforms: vec![ "wasm_mvp".to_string(), "wasm_eh".to_string(), "wasm_threads".to_string(), ], ..ScaffoldConfig::default() }; }
Validate individual platform names with validate_platform:
#![allow(unused)] fn main() { use quack_rs::validate::validate_platform; assert!(validate_platform("linux_amd64").is_ok()); assert!(validate_platform("invalid").is_err()); }
Cargo.toml requirements
[package]
name = "my_extension"
version = "0.1.0"
edition = "2021"
[lib]
name = "my_extension" # Must match description.yml `name`
crate-type = ["cdylib"]
[dependencies]
quack-rs = "0.18"
libduckdb-sys = { version = ">=1.4.4, <2", features = ["loadable-extension"] }
[profile.release]
panic = "unwind" # Required — "abort" disables quack-rs's panic guards
opt-level = 3
lto = true
codegen-units = 1
strip = true
This matches the scaffold's Cargo.toml (which also declares a staticlib example
target for WebAssembly builds).
ADR-4 (in
LESSONS.md) — Do NOT use theduckdbcrate'sbundledfeature. A loadable extension must call into the DuckDB that loads it, not bundle its own copy.libduckdb-syswithloadable-extensionprovides function pointers that are filled in at load time from the API struct DuckDB passes in.
Release profile check
validate_release_profile checks the four release-profile settings. Only
panic = "unwind" is required; lto = true, opt-level = 3 and codegen-units = 1
are recommended, and the returned ReleaseProfileCheck reports each one:
#![allow(unused)] fn main() { use quack_rs::validate::validate_release_profile; // Pass all four release profile settings from your Cargo.toml assert!(validate_release_profile("unwind", "true", "3", "1").is_ok()); // Err — see below assert!(validate_release_profile("abort", "true", "3", "1").is_err()); }
panic must be "unwind". quack-rs wraps every extern "C" entry point in
catch_unwind so a panic in your code becomes a DuckDB error rather than a
crash, and catch_unwind cannot catch anything under panic = "abort": the
runtime aborts before unwinding starts, killing the user's DuckDB session.
The older advice to set abort came from panics escaping an extern "C"
boundary once being undefined behavior. They no longer are — Rust defines that
as an abort — and quack-rs catches them before the boundary anyway.
CI workflow
The scaffold generates .github/workflows/extension-ci.yml, which:
- Runs on pushes and pull requests to
main - Runs
cargo fmt --checkandcargo clippyon Linux, andcargo teston Linux, macOS and Windows - Runs
make configureandmake release, which useextension-ci-toolsto build the.duckdb_extensionfile - Runs the SQLLogicTests in
test/sqlwithmake test
After scaffolding:
cd my_extension
git init
git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools
git submodule update --init --recursive
make configure
make release
Pitfall P4 — The
extension-ci-toolssubmodule must be initialized.make configurefails if the submodule is missing.
Submitting to the community registry
- Create a pull request against the community-extensions repository
- Add your
description.ymlunderextensions/my_extension/description.yml - CI runs automatically to verify the build
- Once approved, users can install your extension:
INSTALL my_extension FROM community;
LOAD my_extension;
Binary compatibility
An extension that uses only the stable C API (the default quack-rs features) is
stamped C_STRUCT with C API version v1.2.0, and one binary loads into every
DuckDB 1.4.x and 1.5.x release for its platform. The libduckdb-sys = ">=1.4.4, <2"
range above is correct for such an extension.
An extension that enables the duckdb-1-5* features should be stamped
C_STRUCT_UNSTABLE with the exact DuckDB release it was built against, so that DuckDB
loads it only into that release. (Stamped C_STRUCT instead, it loads anywhere and
quack-rs's runtime layout check refuses a mismatched release at LOAD.) Pin libduckdb-sys to that release's bindings
(~1.10505.0 for DuckDB 1.5.5) and rebuild for each new DuckDB release. See
ABI Compatibility.
The community build pipeline rebuilds extensions for each DuckDB release.
Pitfall P2 — For
C_STRUCT, the-dvflag toappend_extension_metadata.pymust be the C API version (v1.2.0), not the DuckDB release version (v1.4.4). Usequack_rs::DUCKDB_API_VERSIONto avoid hardcoding this.
Security considerations
The DuckDB team does not audit community extensions for security, so the responsibility is yours:
- Never let a panic escape an FFI boundary: quack-rs's callbacks catch panics and
report them as SQL errors, which requires
panic = "unwind" - Validate user inputs at system boundaries (extension entry point is the boundary)
- Do not include secrets, API keys, or credentials in your binary
- Dynamic SQL in SQL macros must not construct queries from unsanitized user data
Pitfall Catalog
The 31 known pitfalls of writing a DuckDB extension in Rust against the C extension API, with the symptom, root cause and fix for each. The first were found while building duckdb-behavioral, a production DuckDB community extension; the rest while building and auditing quack-rs. Most of them affect any Rust extension that calls the C API directly, and quack-rs prevents most of them. The summary at the end lists each one with its status.
L1: COMBINE must propagate ALL config fields
Status: Testable with AggregateTestHarness.
Symptom: Aggregate function returns wrong results. No error, no crash.
Root cause: DuckDB's segment tree creates fresh target states, initialised
by state_init (with FfiState<T>, a T::default()), then calls combine to
merge source states into them. If your combine only propagates data fields
(count, sum) but omits configuration fields (window_size, mode), the
configuration is still its state_init default at finalize time, silently
corrupting results.
This bug passed 435 unit tests before being caught by E2E tests.
Fix:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { window_size: i64, mode: u8, count: i64 } impl AggregateState for MyState {} unsafe extern "C" fn combine( _info: duckdb_function_info, source: *mut duckdb_aggregate_state, target: *mut duckdb_aggregate_state, count: idx_t, ) { for i in 0..count as usize { let src_ptr = unsafe { *source.add(i) }; let tgt_ptr = unsafe { *target.add(i) }; if let (Some(src), Some(tgt)) = ( FfiState::<MyState>::with_state(src_ptr), FfiState::<MyState>::with_state_mut(tgt_ptr), ) { tgt.window_size = src.window_size; // config — MUST copy tgt.mode = src.mode; // config — MUST copy tgt.count += src.count; // data — accumulate } } } }
Test this with AggregateTestHarness::combine — see Testing Guide.
L2: State destroy double-free
Status: Made impossible by FfiState<T>.
Symptom: Crash or memory corruption on extension unload.
Root cause: If state_destroy frees the inner Box but does not null the
pointer, a second state_destroy call (common in error paths) frees
already-freed memory → undefined behavior.
Fix: FfiState<T>::destroy_callback clears the slot's tag before dropping
the T, and drops only a slot whose tag matches, so a second call is a no-op.
Use it instead of writing your own destructor:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { window_size: i64, mode: u8, count: i64 } impl AggregateState for MyState {} unsafe extern "C" fn state_destroy(states: *mut duckdb_aggregate_state, count: idx_t) { unsafe { FfiState::<MyState>::destroy_callback(states, count) }; } }
You rarely need even this wrapper: .ffi_state::<MyState>() on the aggregate
builder installs destroy_callback together with FfiState<MyState>'s size and
init callbacks.
L3: No panic across FFI boundaries
Status: Made impossible by init_extension and the callback guards (which require panic = "unwind").
Symptom: The whole DuckDB process aborts when an extension callback panics.
Root cause: a panic cannot unwind out of an extern "C" function. Since
Rust 1.81 the runtime aborts the process when one tries (before 1.81 it was
undefined behaviour), so an uncaught panic!() or .unwrap() in a callback
takes down the user's whole DuckDB session.
Fix: Use Result and ? inside init_extension. Never use unwrap() in
FFI callbacks. FfiState::with_state_mut returns Option, not Result, so
callers use if let:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default)] struct MyState { window_size: i64, mode: u8, count: i64 } impl AggregateState for MyState {} unsafe fn demo(state_ptr: duckdb_aggregate_state) { // Safe pattern — no unwrap in FFI callback if let Some(st) = unsafe { FfiState::<MyState>::with_state_mut(state_ptr) } { st.count += 1; } // Dangerous — never do this in an FFI callback let st = unsafe { FfiState::<MyState>::with_state_mut(state_ptr) }.unwrap(); // panics if None } }
quack-rs's callback macros and typed builders catch a panic and report it as a
SQL error. That requires panic = "unwind" in the release profile, which is
what the scaffold generates and what validate_release_profile insists on:
under panic = "abort" nothing can be caught.
L4: ensure_validity_writable is required before NULL output
Status: Made impossible by VectorWriter::set_null.
Symptom: NULLs you write are silently lost — the row reads back as a valid value (whatever is in the data buffer).
Root cause: a vector that has never held a NULL usually has no validity
mask at all, and duckdb_vector_get_validity then returns NULL (as duckdb.h
documents). duckdb_validity_set_row_invalid returns early on a NULL mask, so
nothing is written and nothing crashes. duckdb_vector_ensure_validity_writable
allocates the mask, after which get_validity returns it. (Dereferencing the
NULL pointer yourself, instead of going through the C API helpers, would
crash.)
Fix: Always call duckdb_vector_ensure_validity_writable before accessing
the validity bitmap on the write path. VectorWriter::set_null does this
automatically:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe fn demo(writer: &mut VectorWriter, row: usize) { // Correct — handled by set_null unsafe { writer.set_null(row) }; // Wrong — validity bitmap may not be allocated yet // let validity = duckdb_vector_get_validity(output); // NULL // duckdb_validity_set_row_invalid(validity, row); // silently ignored } }
For STRUCT and ARRAY outputs set_null also nulls the children at that
row, as DuckDB's internal FlatVector::SetNull does; a bare
duckdb_validity_set_row_invalid on the parent leaves the fields valid, and
struct_extract on the NULL row returns their stale values.
L5: Boolean reading must use u8 != 0, not *const bool
Status: Made impossible by VectorReader::read_bool.
Symptom: Undefined behavior; Rust requires bool to be exactly 0 or 1.
Root cause: DuckDB's C API does not guarantee that boolean values in vectors
are exactly 0 or 1. Values of 2, 255, etc. cast to Rust bool is undefined
behavior.
Fix: Read as u8 and compare with != 0. VectorReader::read_bool always
does this:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe fn demo(reader: &VectorReader, row: usize) { let b: bool = unsafe { reader.read_bool(row) }; // safe: uses u8 != 0 internally } }
L6: Function set name must be set on EACH member
Status: Made impossible by AggregateFunctionSetBuilder.
Symptom: Functions are silently not registered. No error returned.
Root cause: When using duckdb_register_aggregate_function_set, the function
name must be set on EACH individual duckdb_aggregate_function using
duckdb_aggregate_function_set_name, not just on the set.
This is completely undocumented. Discovered by reading DuckDB's C++ test code
at test/api/capi/test_capi_aggregate_functions.cpp.
In duckdb-behavioral, 6 of 7 functions failed to register silently due to this bug.
Fix: AggregateFunctionSetBuilder calls duckdb_aggregate_function_set_name
on every individual function before adding it to the set. Use it instead of
managing the set manually.
L7: LogicalType memory leak
Status: Made impossible by LogicalType RAII wrapper.
Symptom: Memory leak proportional to number of registered functions.
Root cause: duckdb_create_logical_type allocates memory that must be freed
with duckdb_destroy_logical_type. Forgetting leaks memory.
Fix: LogicalType implements Drop and calls duckdb_destroy_logical_type
automatically when it goes out of scope.
L8: DEFAULT_NULL_HANDLING does not propagate NULLs for scalar functions
Status: Made impossible by ScalarFunctionBuilder::map1 / map2 /
map1_str / map2_str. DataChunk::propagate_nulls fixes it in one line for
hand-written callbacks.
Symptom: A scalar function returns a value where SQL requires NULL — but only
for arguments that come from a column. SELECT f(NULL) looks correct, because a
literal NULL is constant-folded before the function is reached, so the bug
survives review and ships.
Root cause: The name suggests DuckDB returns NULL on your behalf. For a
scalar function registered through the C API it does not, at run time:
CAPIScalarFunction calls the callback for every row including NULL ones and
checks only the error flag, and the one NULL check in ExpressionExecutor —
VerifyNullHandling — has its entire body inside #ifdef DEBUG. Every DuckDB a
user installs is a release build.
Fix: use the typed closure constructors, which skip NULL rows and write NULL
for them, or call DataChunk::propagate_nulls(&mut writer) at the end of a
hand-written callback. map1_opt / map2_opt and
NullHandling::SpecialNullHandling are for functions that genuinely mean to see
NULLs. Aggregates are no different: update receives NULL rows under either
setting too, so check is_valid before reading (see L12).
L9: duckdb_data_chunk_from_arrow takes the array even when it fails
Status: Made impossible by arrow::data_chunk_from_arrow, which takes the
ArrowArray by value.
Symptom: One of two opposite bugs, depending on which way you guessed. Treat the array as still yours after a failed conversion and you double-release it. Treat it as gone in every case and a zero-column conversion leaks the whole Arrow buffer tree.
Root cause: duckdb.h says "Data ownership is passed on to DuckDB's
DataChunk", which reads like a success-path statement. arrow-c.cpp sets
arrow_array->release = nullptr inside the per-column loop, before the work
that can throw — so the array is claimed on the error path too, but only if the
loop runs at all. A zero-column converted schema leaves release intact and the
array still belongs to the caller.
Fix: own the record in a wrapper whose Drop releases only if release
survived, and consume it by value. The by-value binding drops on the way out: a
no-op when DuckDB nulled release, a correct release when it did not. The
mirror case is handled by the same rule — ToArrowSchema / ToArrowArray
install release last, so a failed export leaves nothing to free.
L10: Scalar bind data is dropped when DuckDB copies the expression
Status: Fixable only from the extension, and now possible:
ScalarBindInfo::set_bind_data_copy.
Symptom: A scalar function that allocates per-query state in its bind
callback reads null from duckdb_scalar_function_get_bind_data during
execution, for some queries and not others. Nothing crashes and nothing is
reported: the callback simply runs without the state it bound, so the answer is
quietly wrong.
Root cause: duckdb_scalar_function_set_bind_data registers the pointer and
its destructor, but not how to duplicate it. DuckDB copies a bound
expression whenever it duplicates a plan, and CScalarFunctionBindData::Copy()
in src/main/capi/scalar_function-c.cpp (read at v1.5.5; byte-identical in
v1.5.4) only fills the copy in when a copy callback exists:
unique_ptr<FunctionData> Copy() const override {
auto copy = make_uniq<CScalarFunctionBindData>(info);
if (copy_callback) {
copy->bind_data = copy_callback(bind_data);
copy->delete_callback = delete_callback;
copy->copy_callback = copy_callback;
}
return std::move(copy); // bind_data stays null without a callback
}
With no callback the copy carries bind_data = nullptr, and the original is
untouched — which is why the failure is intermittent rather than total, and why
it survives a test suite that only ever executes the first-bound expression.
Fix: use ScalarBindData::set, which registers a generated, panic-safe
copy callback (it requires T: Clone + Send + Sync). With the raw API, register
a copy callback alongside the bind data, in the same bind callback and after
set_bind_data:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; use std::os::raw::c_void; use quack_rs::scalar::ScalarBindInfo; #[derive(Clone)] struct MyBindData; unsafe extern "C" fn destroy(p: *mut c_void) { drop(unsafe { Box::from_raw(p.cast::<MyBindData>()) }); } unsafe fn demo(bind_info: ScalarBindInfo, boxed: Box<MyBindData>) { unsafe extern "C" fn copy(data: *mut c_void) -> *mut c_void { if data.is_null() { return std::ptr::null_mut(); } let src = unsafe { &*data.cast::<MyBindData>() }; Box::into_raw(Box::new(src.clone())).cast() } unsafe { bind_info.set_bind_data(Box::into_raw(boxed).cast(), Some(destroy)); bind_info.set_bind_data_copy(Some(copy)); } } }
The duplicate is freed with the same destructor as the original, so copy
must return an independently owned allocation — returning the pointer it was
given is a double free.
Related: the copy callback runs across the FFI boundary like any other, so
it must not unwind. Wrap anything that can panic in
callback::catch_ffi_panic and return null.
L11: C API aggregates crash under agg(x) OVER () and agg(x ORDER BY y)
Status: A DuckDB defect, reported upstream as
duckdb/duckdb#26109. Cannot be prevented or detected from an extension;
documented on AggregateFunctionBuilder, AggregateFunctionSetBuilder and
FfiState.
Symptom: An aggregate that works under SELECT agg(x) FROM t and GROUP BY
segfaults (or corrupts memory, or returns a wrong answer) when used as a window
over a whole-partition frame — agg(x) OVER (), OVER (PARTITION BY p) — or as
an ordered aggregate, agg(x ORDER BY y).
Root cause: CAPIAggregateUpdate (src/main/capi/aggregate_function-c.cpp)
flattens the input vectors but not the state vector, then hands the callback
FlatVector::GetDataUnsafe(state). The C API registers no simple_update, so
two executors fall back to calling update with a constant state vector and
count > 1: WindowConstantAggregatorLocalState (statep(Value::POINTER(0)))
and SortedAggregateFunction (agg_state_vec.SetVectorType(CONSTANT_VECTOR)).
The callback reads states[i] for every row, as the C API contract says it
may; only states[0] exists. Reproduced with a plain C aggregate (no quack-rs)
against DuckDB 1.4.4, 1.5.0 and 1.5.5; AddressSanitizer places the fault in the
callback, called from CAPIAggregateUpdate.
Fix: none on the extension side — the callback receives a raw
duckdb_aggregate_state * and cannot tell a constant vector from a flat one,
and reading states[1] to find out is itself the out-of-bounds read. Until
DuckDB fixes it, document for your users that the aggregate must not be used
in those two query shapes. Frames that are not whole-partition (ROWS BETWEEN 5 PRECEDING AND CURRENT ROW, segment-tree windows) and DISTINCT windows were
checked and work.
L12: Aggregate update receives NULL rows under DEFAULT_NULL_HANDLING
Status: Documented on NullHandling, UpdateFn and both aggregate builders'
null_handling. Pinned by
aggregate_update_receives_null_rows_under_either_null_handling in
tests/ffi_roundtrip/lifecycle.rs.
Symptom: An aggregate that reads every row — state.sum += reader.read_i64(row)
— returns a wrong answer, with no error, as soon as its input column contains a
NULL. The value read for a NULL row is whatever the data buffer happens to hold.
Root cause: quack-rs used to document (in NullHandling, the builders and
this book) that DuckDB's aggregate executor filters NULL rows out before
update unless SpecialNullHandling is set. It does not. CAPIAggregateUpdate
(src/main/capi/aggregate_function-c.cpp) flattens each input vector and
passes the whole chunk, validity and all. For an aggregate the setting is
read in one place, BoundAggregateExpression::PropagatesNullValues, which only
the correlated-subquery decorrelator (flatten_dependent_join.cpp) consults to
pick an INNER or LEFT join; the aggregate VerifyNullHandling check is
compiled only under #ifdef DEBUG. Checked against DuckDB 1.5.5: update saw
every NULL row, ungrouped and under GROUP BY, under both settings, and no
correlated subquery tried answered differently under the two.
Fix: in update, skip rows where VectorReader::is_valid(row) is false,
whatever the null handling. Use SpecialNullHandling to declare that the
aggregate returns non-NULL for NULL input; it does not change which rows arrive.
A related trap under either setting: in a correlated subquery,
(SELECT my_count(x) FROM t2 WHERE t2.k = t1.k) is NULL, not my_count of an
empty input, for an outer row with no match. DuckDB rewrites that NULL to 0 only
for its own count. Wrap the subquery in coalesce(..., 0) if it matters.
L13: A C API aggregate without a destructor is wrong in a running window
Status: Fixed in quack-rs: every aggregate builder registers a destructor,
a no-op when none is given. Pinned by
an_aggregate_without_a_destructor_is_right_in_a_running_window in
tests/ffi_roundtrip/agg_window.rs. Reported in
docs/upstream-duckdb-reports.md, item 7.
Symptom: agg(x) OVER (ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW)
with no PARTITION BY or ORDER BY returns the wrong running value, with no
error: a sum over 1..5 reads 1 2 3 4 5. The same aggregate is right with
ORDER BY, ungrouped, grouped, and in DuckDB's own sum.
Root cause: DuckDB streams such a window only for an aggregate with no
destructor (PhysicalStreamingWindow::IsStreamingFunction). Streaming calls
update once per row, count 1, on a one-row dictionary slice it moves along,
and CAPIAggregateUpdate flattens that input vector in place, so the slice
becomes row 0's value for the rest of the chunk. Checked against 1.4.4, 1.5.0
and 1.5.5.
Fix: register a destructor, even an empty one. quack-rs does this for you;
with the raw C API, call duckdb_aggregate_function_set_destructor.
L14: The C API behaves differently across the releases one build loads into
Status: Four cases fixed in quack-rs (ScalarBindInfo::argument,
LogicalType::try_decimal, LogicalType::try_new, the scalar collision
check); CI job test-older-engines runs the suite against DuckDB 1.4.4 and
1.4.5 (default features), 1.5.0 (duckdb-1-5), 1.5.3 (duckdb-1-5-3) and
1.5.4 (duckdb-1-5-4).
Symptom: code tested against the release Cargo.lock pins (1.5.5) aborts
or misbehaves in an older release the same binary loads into. A default-feature
extension loads into every release from 1.4.4; a duckdb-1-5 one built against
the 1.5.4 bindings has the 546-slot layout of 1.5.2 to 1.5.6, so the ABI guard
rightly lets it load into all five.
Root cause: a C function's contract can change in a release while its
slot stays put. duckdb_scalar_function_bind_get_argument gained its try
only in 1.5.5 (before it, a subquery argument throws through the extension's
callback: an abort in Rust); duckdb_create_decimal_type its width/scale
check only in 1.5.4 (before it, DECIMAL(0, 0) comes back as a type);
duckdb_register_scalar_function its ALTER_ON_CONFLICT only in 1.5.0 (before
it, no existing name can take another overload); and 1.4.x's C API reports
TIME_NS, which its SQL has, as INVALID. All four were found only by running
the whole suite against the oldest release of each range.
Fix: when a wrapper relies on C API behaviour, find the release that
introduced it (the source at each tag answers that), and either check the
engine version at run time (abi::engine_version) or make the check in Rust.
Test against the oldest release a build can load into, not only the pinned one.
L15: combine must leave its source states unchanged
Status: Documented on CombineFn and in the aggregate chapter; pinned by
combine_must_leave_its_source_unchanged (tests/ffi_roundtrip/agg_window.rs).
Not preventable by the SDK: combine receives raw state pointers.
Symptom: an aggregate is right in GROUP BY and wrong in a sliding window
(ROWS BETWEEN n PRECEDING AND CURRENT ROW): in the regression test a sum
whose combine moved its value out of the source gave 4985 of 5000 rows wrong.
Root cause: a window's segment tree keeps one state per tree node and
combines each node's state into every frame that covers it, from several
threads at once (WindowSegmentTreePart::WindowSegmentValue,
window_segment_tree.cpp). A combine that consumes its source (mem::take,
zeroing a counter) is right for the first frame and wrong for the rest, and a
write to a source another thread reads is a data race.
Fix: read the source; copy or clone what the target needs. AggregateState
requires Sync for the same reason.
L16: A valid Arrow array is not always one DuckDB imports correctly
Status: Fixed in quack-rs: data_chunk_from_arrow walks the array with its
schema and refuses the layouts DuckDB 1.4.4 to 1.5.5 mishandles
(src/arrow/import_layout.rs, tests/ffi_roundtrip/arrow_layout.rs;
docs/upstream-duckdb-reports.md, items 9, 24 to 29 and 31 to 33).
Symptom: an array that arrow-rs or another producer built, valid by the
Arrow specification, imports with values from the wrong rows, reads past a
buffer, or corrupts the heap. Arrays DuckDB exported itself never show it,
which is why round-trip tests pass.
Root cause: DuckDB's importer tracks where a node's rows start with two
parameters, parent_offset and nested_offset, and some paths pass the wrong
one: a struct gives its children only its own offset, union members start at
row 0, a dictionary's validity ignores the list's offset. List views, recoded
union type ids and nested dictionaries are mishandled too.
Fix: never assume a producer's layout matches the one DuckDB writes.
Test an importer with hand-built arrays that put offsets at every level, and
refuse what the engine cannot import rather than return wrong values.
L17: A COPY … FROM reader must not declare result columns
Status: Refused for typed table functions (their bind fails with a
message); documented for raw ones on CopyFunctionBuilder::copy_from and
BindInfo::add_result_column. Pinned by
tests/ffi_roundtrip/copy_from_columns.rs;
docs/upstream-duckdb-reports.md, item 37.
Symptom: on a DuckDB built with assertions, COPY t FROM … fails with
chunk.ColumnCount() == types.size() and the database is invalidated. A
release build silently drops the extra column, so the bug hides in testing.
Root cause: CCopyFromBind hands the reader's bind the INSERT's own list
of expected types as its result types, and duckdb_bind_add_result_column
appends to that list, so every chunk the INSERT receives is wider than the
table. duckdb.h says the reader "should not" declare columns; nothing
enforces it.
Fix: in a COPY … FROM reader's bind, read the target's columns with
BindInfo::result_column_count and its siblings, and declare none.
L18: A LIST reserve moves every buffer below its child
Status: Documented in the # Safety sections of VectorWriter::from_vector,
StructWriter::new, StructVector::field_writer and
ValidityBitmap::ensure_writable; measured by
tests/ffi_roundtrip/nested_reserve.rs.
Symptom: a writer on a STRUCT field of a list's elements writes into freed memory after the list is grown, although it was never a direct child of the list.
Root cause: duckdb_list_vector_reserve resizes the child with
Vector::Resize, which reallocates the data and validity buffers of the child
and of every STRUCT field and ARRAY element vector below it, down to the next
LIST (whose child has its own buffer). Writers cache both pointers.
Fix: fetch every writer and bitmap below a list's child again after each
reserve on that list (a ListBuilder row that grows it counts).
L19: Addresses of constants are not identities
Status: Fixed in FfiState (0.18.0): its per-type tag salt is a hash of
TypeId::of::<T>(), not the address of type_name::<T>(). AUDIT.md 10.2
(High) and 10.8.
Symptom: a per-type tag compared across callbacks mismatches only in
release builds of the user's crate — FfiState::with_state returns None,
destroy_callback skips states — while debug and fat-LTO builds hide it.
Root cause: core::any::type_name::<T>().as_ptr() (and a function
pointer, and the address of any &'static constant) can differ between
codegen units: rustc emits a private copy of a constant in each codegen unit
that uses it, so with codegen-units > 1 and no fat LTO (Cargo's default
release profile) two uses of the same constant can have different addresses.
Rust makes no address-identity guarantee for functions or constants.
Fix: derive identity from a value, e.g. hash TypeId::of::<T>().
Evidence: a standalone crate taking type_name::<T>().as_ptr() for one
T in 8 modules: release profile (codegen-units = 16, lto = false) → 6
distinct addresses; dev profile → 1; codegen-units = 1, lto = true → 1.
In quack-rs itself, cargo test --release --lib aggregate:: with
CARGO_PROFILE_RELEASE_CODEGEN_UNITS=16 CARGO_PROFILE_RELEASE_LTO=false
failed 5 tests on the address salt and passes on the TypeId one; the
end-to-end suite under the same profile failed 10 of 279 (every aggregate:
NULL or garbage results) and passes all 279. Every other
CI build is debug or uses the repository's codegen-units = 1, fat-LTO
release profile, which is why the bug got past them; CI's test job now
runs this build.
P1: Library name must match extension name
Status: Must be configured in Cargo.toml. Scaffold handles this.
Symptom: Community build fails with FileNotFoundError.
Root cause: The community build expects lib{extension_name}.so. If the
Cargo crate name produces a different .so filename, the build fails.
Fix: Set name explicitly in [lib]:
[lib]
name = "my_extension" # Must match description.yml `name: my_extension`
crate-type = ["cdylib", "rlib"]
P2: Metadata version is C API version, not DuckDB version
Status: The DUCKDB_API_VERSION constant holds the correct value.
Symptom: The metadata script succeeds, and LOAD then refuses the file:
"The file was built for DuckDB C API version 'v1.5.5', but we can only load
extensions built for DuckDB C API 'v1.2.0' and lower" (verified on DuckDB 1.4.4
and 1.5.5 with a file stamped -dv v1.5.5).
Root cause: The -dv flag to append_extension_metadata.py must be the
C API version (v1.2.0), not the DuckDB release version (v1.4.4). These are
different strings. DuckDB 1.4.x and 1.5.0 – 1.5.5 declare C API version
v1.2.0; 1.5.6 declares v1.5.6 and still loads v1.2.0 extensions.
Fix: Use quack_rs::DUCKDB_API_VERSION ("v1.2.0") in init_extension,
and use the same version with append_extension_metadata.py -dv v1.2.0.
This holds only for the C_STRUCT ABI type. For C_STRUCT_UNSTABLE and CPP,
-dv is the exact DuckDB release: with USE_UNSTABLE_C_API=1 (required when
you use quack-rs's duckdb-1-5 features; see P10),
TARGET_DUCKDB_VERSION must be a real release such as v1.5.6, and v1.2.0
would pin the binary to DuckDB v1.2.0. ScaffoldConfig validates this pairing.
P3: E2E testing is mandatory
Status: Documented. See Testing Guide.
Symptom: All unit tests pass but the extension is completely broken.
Root cause: Unit tests cannot detect SEGFAULTs on load, silent registration failures, or wrong results from combine bugs.
Fix: Always run E2E tests using an actual DuckDB binary. The scaffold generates a complete SQLLogicTest skeleton.
P4: extension-ci-tools submodule must be initialized
Status: Build-time check.
Symptom: make configure or make release fails.
Fix: In a new project (for example one fresh from the scaffold) the
submodule has never been added: the scaffold writes .gitmodules, but a file
cannot create the gitlink git needs, so git submodule update --init finds
nothing to do and exits 0 without cloning anything. Add it once:
git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools
In a clone of a repository that already has the submodule:
git submodule update --init --recursive
The generated Makefile checks for the checkout before it includes anything
from it and prints both commands if it is missing.
P5: SQLLogicTest expected values must match exactly
Status: Test-authoring care required.
Symptom: Tests fail in CI but pass locally (or vice versa).
Root cause: SQLLogicTest does exact string matching. Output format (decimal places, NULL representation, column separators) must match character-for-character.
Fix: Generate expected values by running the SQL in DuckDB CLI and copying
the output. NULL is NULL (uppercase). Integers have no decimal places.
P6: duckdb_register_aggregate_function_set silently fails
Status: Builder returns Err. Also see L6.
Symptom: Function appears registered but is not found in SQL.
Root cause: The return value of duckdb_register_aggregate_function_set is
often ignored. When it returns DuckDBError, the function set is not registered.
Fix: The builder checks the return value and propagates it as Err.
P7: duckdb_string_t format is undocumented
Status: Handled by VectorReader::read_str, VectorReader::read_blob, and
DuckStringView.
Symptom: VARCHAR reading produces garbage, empty strings, or crashes; BLOB reading silently drops bytes that are not valid UTF-8.
Root cause: DuckDB stores strings in a 16-byte struct with two formats
(inline ≤ 12 bytes, pointer > 12 bytes) that are not documented in
libduckdb-sys. The length and the pointer are in the target's own byte
order, so a decoder that reads them as little-endian misreads every string on
a big-endian target.
Fix: Use VectorReader::read_str(row) for UTF-8 text and
VectorReader::read_blob(row) for arbitrary binary data. See
NULL Handling & Strings.
P8: INTERVAL struct layout is undocumented
Status: Handled by DuckInterval and read_interval_at.
Symptom: Interval calculations produce wrong results or crashes.
Root cause: DuckDB's INTERVAL is { months: i32, days: i32, micros: i64 }
(16 bytes total). This is not documented in libduckdb-sys. Month conversion
uses 1 month = 30 days (DuckDB's approximation).
Fix: Use VectorReader::read_interval(row) and DuckInterval. See
INTERVAL Type.
P9: loadable-extension dispatch table uninitialised in cargo test
Status: Fixed. InMemoryDb::open() initialises the dispatch table
automatically.
Symptom: All three InMemoryDb unit tests panic at runtime:
thread 'testing::in_memory_db::tests::in_memory_db_opens' panicked at
'DuckDB API not initialized or DuckDB feature omitted'
This failure appears only when running cargo test --features bundled-test.
Regular cargo test (no feature) does not exercise this code path, so CI can
miss it entirely.
Root cause: Cargo's feature-unification merges loadable-extension (from
the main libduckdb-sys dependency) and bundled (pulled in by the
duckdb crate's features = ["bundled"]) into a single libduckdb-sys build
with both features active. In loadable-extension mode every DuckDB C API
call is routed through a dispatch table of one AtomicPtr per function, which
is normally populated at load time, when DuckDB calls the extension's entry
point and the entry point calls duckdb_rs_extension_api_init. In
cargo test, no DuckDB host process loads the extension, so the table stays
uninitialised and every call panics.
Discovery: This was triggered by the crates.io release workflow (which runs
cargo test --all-targets --all-features) failing on macOS. Regular CI at the time
(cargo test --all-targets, no --all-features) never compiled the bundled-test path, so the bug was hidden
during development and code review.
Fix (implemented in quack-rs 0.6.0):
-
src/testing/bundled_api_init.cpp— a thin C++ shim that wraps DuckDB's internalCreateAPIv1()(fromduckdb/main/capi/extension_api.hpp) as a C-linkage symbol:#include "duckdb/main/capi/extension_api.hpp" extern "C" duckdb_ext_api_v1 quack_rs_create_api_v1() { return CreateAPIv1(); } -
build.rs— compiles the shim (via thecccrate) only when thebundled-testorbundled-test-prebuiltfeature is active. It finds the DuckDB headers throughDEP_DUCKDB_INCLUDE(published bylibduckdb-sys >= 1.10503), falling back to thelibduckdb-sysbuild output directory or, for a prebuilt library,DUCKDB_INCLUDE_DIR. -
InMemoryDb::open()— callsinit_dispatch_table_once()before opening the connection. That function callsquack_rs_create_api_v1()once and feeds the result throughduckdb_rs_extension_api_init, populating everyAtomicPtrslot in the dispatch table, one per field ofduckdb_ext_api_v1(546 with the 1.5.2 – 1.5.6 bindings, 459 with 1.4.x; see the table in ABI Compatibility). Astd::sync::Onceguard makes it safe to call from any number of threads and test cases. -
CI
test-bundledjob — runscargo test --all-targets --features bundled-testand then the release workflow's owncargo test --all-targets --all-featureson Linux, macOS and Windows on every PR. The second step was added after the v0.18.0 tag failed on Windows while PR CI was green: until then no PR job ran theduckdb-1-5*tests on macOS or Windows.
ABI compatibility note: DuckDB's duckdb_ext_api_v1 struct is defined
identically in both the public duckdb_extension.h (used by libduckdb-sys
bindgen) and the internal extension_api.hpp (used by CreateAPIv1()). Both
include the DUCKDB_EXTENSION_API_VERSION_UNSTABLE fields. CreateAPIv1() sets
every field. The Rust and C++ structs are produced from the same DuckDB
release and therefore stay in sync.
Risk table (using DuckDB's internal C++ API):
| Risk | Mitigation |
|---|---|
extension_api.hpp is renamed or moved | build.rs fails with a clear compile error |
CreateAPIv1() is renamed | Same — C++ compile error |
duckdb_ext_api_v1 gains new fields | CreateAPIv1() fills new fields too |
duckdb_ext_api_v1 field order changes | Both structs from same DuckDB release, stay in sync |
libduckdb-sys drops loadable-extension dispatch | Problem disappears; Once guard becomes cheap no-op |
P10: The C API struct has a stable prefix and an unstable tail
Status: Detected at load time by quack_rs::abi
(default AbiPolicy::Strict).
Symptom: An extension loads without complaint and then corrupts memory —
double free or corruption, a segfault, or silently wrong results — on a
DuckDB release other than the one it was built against. Nothing in the build
or the load warns you.
Root cause: DuckDB hands a loadable extension a pointer to a
duckdb_ext_api_v1 struct of function pointers, and the extension calls through
it at compiled-in offsets. The struct has two regions:
| Region | Slots | Guarantee |
|---|---|---|
| Stable | 0–356 | Frozen since v1.2.0 — same slots, order and signatures in every release through v1.5.6 (two slots, 114 and 138, were renamed varint → bignum in v1.4.0 with an identical struct layout) |
| Unstable | 357+ | DuckDB inserts entries in the middle, shifting every later slot |
duckdb_appender_clear landed at slot 410 in v1.5.0 and
duckdb_geometry_type_get_crs in the middle of v1.5.2's tail; each insertion
moves everything after it. An extension compiled against one layout and loaded
by another calls the wrong function through the right offset.
Your action: use init_extension, which verifies the layout and refuses a
mismatch. If you enable a duckdb-1-5* feature, also stamp the binary
C_STRUCT_UNSTABLE with the exact DuckDB release (USE_UNSTABLE_C_API=1 and a
real TARGET_DUCKDB_VERSION, or append_metadata --abi-type C_STRUCT_UNSTABLE --duckdb-version vX.Y.Z), so that DuckDB itself refuses to load it into any
other release. Two knobs matter:
QUACK_RS_TARGET_DUCKDB_VERSIONat build time stamps the release you built against, so aDuckDBnewer than quack-rs's table is still accepted when your build genuinely targeted it. The community-extension CI rebuilds per release, so this is the normal path.AbiPolicy::WarnorTrustif you would rather load anyway.Trustis the old behaviour, and the failure mode above is the reason it is no longer the default.
If your extension enables no duckdb-1-5* feature it only calls into the stable
prefix: the check reports AbiCheck::StableOnly, and the extension loads into
every release from v1.2.0 on.
P11: const char * returns are borrowed — freeing one corrupts the heap
Status: Fixed in quack-rs; documented here because extension authors calling the C API directly hit the same trap.
Symptom: corrupted size vs. prev_size in fastbins, free(): invalid pointer, or a SIGABRT at an unrelated later allocation. Nothing points at the
call that caused it.
Root cause: The C API returns strings two ways, and only one transfers ownership.
| Return type | Typical implementation | Caller must |
|---|---|---|
char * | strdup(...) or duckdb_malloc + memcpy | duckdb_free it |
const char * | some_std_string.c_str() | not free it |
duckdb_copy_function_global_init_get_file_path is the second kind: it returns
info_ref.file_path.c_str(), the interior pointer of a C++ std::string
DuckDB still owns and destroys itself. Calling duckdb_free on it hands the
allocator a pointer it never issued.
The trap in the rule: the signature alone is not enough.
duckdb_parameter_name is declared const char * and yet returns
strdup(identifier.c_str()) — it is owned, and not freeing it leaks. The
only reliable check is reading the implementation in DuckDB's
src/main/capi/.
Your action: before calling duckdb_free on anything the C API returned,
read the implementation. const char * is a strong hint that it is borrowed,
but duckdb_parameter_name proves it is only a hint. Every duckdb_free site
in quack-rs was audited this way; see
LESSONS.md P11
for the full table.
How it was found: by writing the first live test for copy functions. The
module had 16 unit tests and none of them registered a copy function against a
real DuckDB, so the corruption had never had a chance to happen. Unit tests
over an FFI wrapper test the wrapper's arithmetic, not its contract with the
library.
P12: duckdb_client_context_get_config_option aborts on a missing setting
Status: A DuckDB defect, not a quack-rs one. Documented on
ClientContext::config_option, with an abort-free alternative.
Symptom: Assertion 'scope != SettingScope::INVALID' failed and a
SIGABRT when asking for a configuration option that does not exist — but only
against a DuckDB built with debug assertions. Release builds return NULL
exactly as documented, so this never reproduces for end users and always
reproduces in a test suite that links a debug DuckDB.
Root cause (DuckDB 1.5.5):
// src/main/capi/config_options-c.cpp
switch (ctx.TryGetCurrentSetting(option_name, result).GetScope()) {
...
default: // <- INVALID is handled here
res_scope = DUCKDB_CONFIG_OPTION_SCOPE_INVALID;
// src/include/duckdb/main/setting_info.hpp
SettingScope GetScope() {
D_ASSERT(scope != SettingScope::INVALID); // <- but never reached in debug
return scope;
}
The default: arm shows the not-found case is meant to be tolerated; the code
just calls GetScope() before checking operator bool().
Your action: use ClientContext::config_option for settings you registered
or know exist. To ask whether a setting exists, use SQL — it has no assertion
on this path:
SELECT count(*) FROM duckdb_settings() WHERE name = 'my_setting';
Summary
| Pitfall | SDK status | Your action |
|---|---|---|
| L1: combine config fields | Testable | Test with AggregateTestHarness::combine |
| L2: state double-free | Prevented | Use FfiState::destroy_callback |
| L3: panic across FFI | Prevented | Use init_extension, no unwrap in callbacks |
| L4: NULL silently dropped (no validity mask) | Prevented | Use VectorWriter::set_null |
| L5: bool UB | Prevented | Use VectorReader::read_bool |
| L6: function set name | Prevented | Use AggregateFunctionSetBuilder |
| L7: LogicalType leak | Prevented | Use LogicalType (RAII) |
| L8: NULLs reach the callback anyway | Prevented | Use map1/map2, or DataChunk::propagate_nulls |
| L9: Arrow array taken on failure | Prevented | Use arrow::data_chunk_from_arrow (takes by value) |
| L10: bind data lost on expression copy | Prevented | Use ScalarBindData::set (or pair set_bind_data with set_bind_data_copy) |
L11: aggregate crash under OVER () / ORDER BY | DuckDB defect | Do not use C API aggregates in those query shapes |
L12: aggregate update sees NULL rows | Documented | Skip rows where is_valid is false |
| L13: running window without a destructor | Prevented | Every aggregate builder registers a destructor |
| L14: C API differs across releases | Prevented | Wrappers that rely on newer behaviour check the engine version or check in Rust |
L15: combine consumes its source | Documented | Read the source states; copy what the target needs |
| L16: Arrow layouts DuckDB misimports | Prevented | data_chunk_from_arrow refuses them |
L17: COPY … FROM reader declares columns | Prevented (typed) / Documented | Read the target's columns; declare none |
L18: LIST reserve moves nested buffers | Documented | Fetch writers again after each reserve |
| L19: constant addresses as identities | Fixed | FfiState salts its tag with a hash of TypeId |
| P1: lib name mismatch | Scaffold | Set [lib] name in Cargo.toml |
| P2: API version string | Constant | Use DUCKDB_API_VERSION |
| P3: unit tests insufficient | Documented | Write SQLLogicTest E2E tests |
| P4: submodule not initialized | Build-time | New project: git submodule add …; clone: git submodule update --init |
| P5: SQLLogicTest exact match | Documented | Copy output from DuckDB CLI |
| P6: register set silent fail | Prevented | Builder returns Err |
| P7: VARCHAR format undocumented | Prevented | Use VectorReader::read_str |
| P8: INTERVAL layout undocumented | Prevented | Use DuckInterval |
| P9: dispatch table uninitialised | Fixed | InMemoryDb::open() initialises it via C++ shim |
| P10: unstable ABI tail shifts | Prevented | Use init_extension; set QUACK_RS_TARGET_DUCKDB_VERSION when building |
P11: freeing a borrowed const char * | Fixed | Read the C++ impl before duckdb_free; prefer quack-rs wrappers |
| P12: config-option probe aborts (debug) | Documented | Ask duckdb_settings() in SQL instead |
TypeId Reference
quack_rs::types::TypeId is the enum of DuckDB column types that the quack-rs
builder APIs accept. Each variant names one of the C API's DUCKDB_TYPE_*
integer constants, which libduckdb-sys exposes as
DUCKDB_TYPE_DUCKDB_TYPE_* (for example
libduckdb_sys::DUCKDB_TYPE_DUCKDB_TYPE_BIGINT).
Full variant table
| Variant | SQL name | C API constant | Notes |
|---|---|---|---|
TypeId::Boolean | BOOLEAN | DUCKDB_TYPE_BOOLEAN | true/false, stored as a u8 |
TypeId::TinyInt | TINYINT | DUCKDB_TYPE_TINYINT | 8-bit signed |
TypeId::SmallInt | SMALLINT | DUCKDB_TYPE_SMALLINT | 16-bit signed |
TypeId::Integer | INTEGER | DUCKDB_TYPE_INTEGER | 32-bit signed |
TypeId::BigInt | BIGINT | DUCKDB_TYPE_BIGINT | 64-bit signed |
TypeId::UTinyInt | UTINYINT | DUCKDB_TYPE_UTINYINT | 8-bit unsigned |
TypeId::USmallInt | USMALLINT | DUCKDB_TYPE_USMALLINT | 16-bit unsigned |
TypeId::UInteger | UINTEGER | DUCKDB_TYPE_UINTEGER | 32-bit unsigned |
TypeId::UBigInt | UBIGINT | DUCKDB_TYPE_UBIGINT | 64-bit unsigned |
TypeId::HugeInt | HUGEINT | DUCKDB_TYPE_HUGEINT | 128-bit signed |
TypeId::Float | FLOAT | DUCKDB_TYPE_FLOAT | 32-bit IEEE 754 |
TypeId::Double | DOUBLE | DUCKDB_TYPE_DOUBLE | 64-bit IEEE 754 |
TypeId::Timestamp | TIMESTAMP | DUCKDB_TYPE_TIMESTAMP | µs since Unix epoch |
TypeId::TimestampTz | TIMESTAMPTZ | DUCKDB_TYPE_TIMESTAMP_TZ | timezone-aware timestamp |
TypeId::Date | DATE | DUCKDB_TYPE_DATE | days since epoch |
TypeId::Time | TIME | DUCKDB_TYPE_TIME | µs since midnight |
TypeId::Interval | INTERVAL | DUCKDB_TYPE_INTERVAL | months + days + µs |
TypeId::Varchar | VARCHAR | DUCKDB_TYPE_VARCHAR | UTF-8 string |
TypeId::Blob | BLOB | DUCKDB_TYPE_BLOB | binary data |
TypeId::Decimal | DECIMAL | DUCKDB_TYPE_DECIMAL | fixed-point decimal |
TypeId::TimestampS | TIMESTAMP_S | DUCKDB_TYPE_TIMESTAMP_S | seconds since epoch |
TypeId::TimestampMs | TIMESTAMP_MS | DUCKDB_TYPE_TIMESTAMP_MS | milliseconds since epoch |
TypeId::TimestampNs | TIMESTAMP_NS | DUCKDB_TYPE_TIMESTAMP_NS | nanoseconds since epoch |
TypeId::Enum | ENUM | DUCKDB_TYPE_ENUM | enumeration type |
TypeId::List | LIST | DUCKDB_TYPE_LIST | variable-length list |
TypeId::Struct | STRUCT | DUCKDB_TYPE_STRUCT | named fields (row type) |
TypeId::Map | MAP | DUCKDB_TYPE_MAP | key-value pairs |
TypeId::Uuid | UUID | DUCKDB_TYPE_UUID | 128-bit UUID |
TypeId::Union | UNION | DUCKDB_TYPE_UNION | tagged union of types |
TypeId::Bit | BIT | DUCKDB_TYPE_BIT | bitstring |
TypeId::TimeTz | TIMETZ | DUCKDB_TYPE_TIME_TZ | timezone-aware time |
TypeId::UHugeInt | UHUGEINT | DUCKDB_TYPE_UHUGEINT | 128-bit unsigned |
TypeId::Array | ARRAY | DUCKDB_TYPE_ARRAY | fixed-length array |
TypeId::TimeNs | TIME_NS | DUCKDB_TYPE_TIME_NS | nanosecond-precision time |
TypeId::Any | ANY | DUCKDB_TYPE_ANY | wildcard for function signatures |
TypeId::Varint | BIGNUM | DUCKDB_TYPE_BIGNUM | arbitrary-precision integer (VARINT before DuckDB 1.4) |
TypeId::SqlNull | SQLNULL | DUCKDB_TYPE_SQLNULL | explicit SQL NULL type |
TypeId::IntegerLiteral | INTEGER_LITERAL | DUCKDB_TYPE_INTEGER_LITERAL | unresolved integer literal |
TypeId::StringLiteral | STRING_LITERAL | DUCKDB_TYPE_STRING_LITERAL | unresolved string literal |
TypeId::Geometry | GEOMETRY | DUCKDB_TYPE_GEOMETRY | spatial geometry value (duckdb-1-5-3) |
TypeId::Variant | VARIANT | DUCKDB_TYPE_VARIANT | self-describing nested value, e.g. Iceberg v3 (duckdb-1-5-3) |
Feature gate for
Geometry/Variant:DUCKDB_TYPE_GEOMETRY(40) andDUCKDB_TYPE_VARIANT(41) require theduckdb-1-5-3feature, which layers on top ofduckdb-1-5and needslibduckdb-sys >= 1.10503.0(DuckDB 1.5.3).VARIANTentered the C type enum in DuckDB 1.5.3, after theduckdb-1-5feature's 1.5.0 floor, andGEOMETRYis gated with it so that one feature covers both. Gating them separately avoids breaking consumers pinned to libduckdb-sys 1.10500–1.10502 (DuckDB 1.5.0–1.5.2). See Known Limitations.
Methods
to_duckdb_type() → DUCKDB_TYPE
Converts to the raw C API integer constant. Used internally by the builder APIs.
#![allow(unused)] fn main() { use quack_rs::types::TypeId; let raw: libduckdb_sys::DUCKDB_TYPE = TypeId::BigInt.to_duckdb_type(); }
from_duckdb_type(raw) → TypeId
Converts a raw DUCKDB_TYPE constant back into a TypeId. Recognizes every
variant available in the active feature set, including TIME_NS, ANY,
BIGNUM, SQLNULL, INTEGER_LITERAL and STRING_LITERAL (no feature needed:
all six exist in every DuckDB this crate supports) and the duckdb-1-5-3
values (GEOMETRY, VARIANT) when that feature is enabled.
Panics if the value does not correspond to any variant available in the current
feature configuration; try_from_duckdb_type returns None instead.
#![allow(unused)] fn main() { use quack_rs::types::TypeId; let type_id = TypeId::from_duckdb_type(libduckdb_sys::DUCKDB_TYPE_DUCKDB_TYPE_BIGINT); assert_eq!(type_id, TypeId::BigInt); assert_eq!(TypeId::try_from_duckdb_type(9_999), None); }
is_composite() and composite_constructor_hint()
DECIMAL, ENUM, LIST, STRUCT, MAP, ARRAY and UNION carry parameters
that a bare type id cannot express, so duckdb_create_logical_type cannot build
them. is_composite() returns true for these seven, and
composite_constructor_hint() names the LogicalType constructor to use
instead:
#![allow(unused)] fn main() { use quack_rs::types::TypeId; assert!(TypeId::List.is_composite()); assert_eq!( TypeId::List.composite_constructor_hint(), Some("LogicalType::list(element_type)") ); assert_eq!(TypeId::BigInt.composite_constructor_hint(), None); }
sql_name() → &'static str
Returns the SQL type name as a static string.
#![allow(unused)] fn main() { use quack_rs::types::TypeId; assert_eq!(TypeId::BigInt.sql_name(), "BIGINT"); assert_eq!(TypeId::Varchar.sql_name(), "VARCHAR"); assert_eq!(TypeId::TimestampTz.sql_name(), "TIMESTAMPTZ"); }
Display
TypeId implements Display, which outputs the SQL name:
#![allow(unused)] fn main() { use quack_rs::types::TypeId; println!("{}", TypeId::Interval); // prints: INTERVAL let s = format!("{}", TypeId::UBigInt); // "UBIGINT" assert_eq!(s, "UBIGINT"); assert_eq!(TypeId::Interval.to_string(), "INTERVAL"); }
VectorReader/VectorWriter mapping
The read and write methods on VectorReader/VectorWriter map to TypeId
variants as follows:
| TypeId | Read method | Write method | Rust type |
|---|---|---|---|
Boolean | read_bool | write_bool | bool |
TinyInt | read_i8 | write_i8 | i8 |
SmallInt | read_i16 | write_i16 | i16 |
Integer | read_i32 | write_i32 | i32 |
BigInt | read_i64 | write_i64 | i64 |
UTinyInt | read_u8 | write_u8 | u8 |
USmallInt | read_u16 | write_u16 | u16 |
UInteger | read_u32 | write_u32 | u32 |
UBigInt | read_u64 | write_u64 | u64 |
Float | read_f32 | write_f32 | f32 |
Double | read_f64 | write_f64 | f64 |
Varchar | read_str | write_varchar | &str |
Interval | read_interval | write_interval | DuckInterval |
HugeInt | read_i128 | write_i128 | i128 |
UHugeInt | read_u128 | write_u128 | u128 |
Blob | read_blob | write_blob | &[u8] |
Uuid | read_uuid | write_uuid | u128 (textual bits) |
Date | read_date | write_date | i32 (days since epoch) |
Time | read_time | write_time | i64 (µs since midnight) |
TimeTz | read_time_tz | write_time_tz | u64 (packed) |
Timestamp | read_timestamp | write_timestamp | i64 (µs since epoch) |
TimestampTz | read_timestamp_tz | write_timestamp_tz | i64 (µs since epoch) |
TimestampS | read_timestamp_s | write_timestamp_s | i64 (s since epoch) |
TimestampMs | read_timestamp_ms | write_timestamp_ms | i64 (ms since epoch) |
TimestampNs | read_timestamp_ns | write_timestamp_ns | i64 (ns since epoch) |
Decimal | read_decimal(row, width) | write_decimal(row, width, v) | i128 (unscaled) |
List, Map, Struct and Array are nested vectors: use the helpers in
Complex Types (ListVector, MapVector,
StructVector, ArrayVector, StructReader / StructWriter, ListBuilder).
Enum, Union, Bit, TimeNs, Any, Varint, SqlNull, IntegerLiteral,
StringLiteral, Geometry and Variant do not yet have dedicated read/write
helpers. Access these via the raw data pointer from duckdb_vector_get_data.
Properties
TypeId implements Debug, Clone, Copy, PartialEq, Eq, and Hash, so
it can be used as a map key or set element and compared in match expressions:
#![allow(unused)] fn main() { use std::collections::HashMap; use quack_rs::types::TypeId; let mut type_names: HashMap<TypeId, &str> = HashMap::new(); type_names.insert(TypeId::BigInt, "count"); type_names.insert(TypeId::Varchar, "label"); }
#[non_exhaustive]
TypeId is marked #[non_exhaustive], so quack-rs can add variants for new
DuckDB types without a breaking change. A match on TypeId outside quack-rs
needs a wildcard arm:
#![allow(unused)] fn main() { use quack_rs::types::TypeId; fn demo(type_id: TypeId) { match type_id { TypeId::BigInt => { /* ... */ } TypeId::Varchar => { /* ... */ } _ => { /* handle future types */ } } } }
LogicalType
For types that require parameters (such as DECIMAL(p, s) or LIST(INTEGER)),
use quack_rs::types::LogicalType:
#![allow(unused)] fn main() { use quack_rs::types::{LogicalType, TypeId}; let lt = LogicalType::new(TypeId::BigInt); // or use the From impl: let lt: LogicalType = TypeId::BigInt.into(); // LogicalType implements Drop → calls duckdb_destroy_logical_type automatically }
LogicalType wraps duckdb_logical_type with RAII cleanup, preventing the
memory leak described in Pitfall L7.
Constructors
| Constructor | Creates |
|---|---|
new(type_id) | Simple type from a TypeId |
from_raw(ptr) | Takes ownership of a raw handle (unsafe) |
decimal(width, scale) | DECIMAL(width, scale) |
list(element_type) | LIST<T> from a TypeId |
list_from_logical(element) | LIST<T> from an existing LogicalType |
map(key, value) | MAP<K, V> from TypeIds |
map_from_logical(key, value) | MAP<K, V> from existing LogicalTypes |
struct_type(fields) | STRUCT from &[(&str, TypeId)] |
struct_type_from_logical(fields) | STRUCT from &[(&str, LogicalType)] |
union_type(members) | UNION from &[(&str, TypeId)] |
union_type_from_logical(members) | UNION from &[(&str, LogicalType)] |
enum_type(members) | ENUM from &[&str] |
array(element_type, size) | ARRAY<T>[size] from a TypeId |
array_from_logical(element, size) | ARRAY<T>[size] from an existing LogicalType |
Each constructor except from_raw has a try_ form (try_new,
try_decimal, try_list, …) that returns an error instead of panicking on
invalid input, such as a composite TypeId passed to new, a DECIMAL width
above 38, or a UNION with more than MAX_UNION_MEMBERS (255) members.
Introspection methods
The introspection methods are all unsafe, because they call into a loaded
DuckDB:
get_type_id, get_alias, set_alias, decimal_width, decimal_scale,
decimal_internal_type, enum_internal_type, enum_dictionary_size,
enum_dictionary_value, list_child_type, map_key_type, map_value_type,
struct_child_count, struct_child_name, struct_child_type,
union_member_count, union_member_name, union_member_type,
array_size, array_child_type.
See Type System for the full introspection table.
Known Limitations
What quack-rs cannot do because the DuckDB C extension API does not allow it, DuckDB behaviour an extension should plan around, and former limitations that have since been resolved.
Window functions are not available
DuckDB window functions (OVER (...) clauses) are implemented entirely in
DuckDB's C++ layer and have no counterpart in the public C extension API.
This is not a gap in quack-rs or in libduckdb-sys — the relevant symbol
(duckdb_create_window_function) simply does not exist in the C API:
| Symbol | C API (1.4.x)? | C API (1.5.0+)? | C++ API? |
|---|---|---|---|
duckdb_create_window_function | No | No | Yes |
duckdb_create_copy_function | No | Yes | Yes |
duckdb_create_scalar_function | Yes | Yes | Yes |
duckdb_create_aggregate_function | Yes | Yes | Yes |
duckdb_create_table_function | Yes | Yes | Yes |
duckdb_create_cast_function | Yes | Yes | Yes |
What this means for your extension:
A custom window operator requires a C++ extension. An ordinary aggregate can be
used as a window (agg(x) OVER (...)), but see the next section before
recommending that to your users: two of those shapes crash every aggregate
registered through the C API.
If DuckDB exposes window registration in a future C API version, quack-rs
will add wrappers in the corresponding release.
Aggregates crash under OVER () and ORDER BY (DuckDB defect, Pitfall L11)
Every aggregate registered through the C API — quack-rs's or anyone else's —
reads out of bounds when DuckDB runs it as a whole-partition window
(agg(x) OVER (), agg(x) OVER (PARTITION BY p)) or as an ordered aggregate
(agg(x ORDER BY y)). The usual result is a segmentation fault that takes the
host process down. The book's own example aggregate does it:
-- with hello-ext loaded: segfaults on DuckDB 1.4.4 and 1.5.5
SELECT max(w) FROM (SELECT word_count(s) OVER () AS w
FROM (SELECT 'a b c' AS s FROM range(5000)));
The cause is in DuckDB (CAPIAggregateUpdate hands the callback a constant
state vector), there is no way for an extension to detect or prevent it, and it
is reported upstream as
duckdb/duckdb#26109. Until it
is fixed, tell your users not to use your aggregates in those two shapes.
Frames that are not whole-partition (ROWS BETWEEN 5 PRECEDING AND CURRENT ROW)
and DISTINCT windows work. See
Pitfall L11.
Aggregate states leak when finalize reports an error (DuckDB behaviour)
When an aggregate's finalize callback reports an error
(AggregateFunctionInfo::set_error), DuckDB 1.5.5 does not call the destructor
for every state the query created: an ungrouped query initialised 2 states and
destroyed 1, a grouped one 4 and 2. When finalize succeeds, every state is
destroyed. The extension cannot tell which states were abandoned, so whatever
they own is leaked: with FfiState<T>, whatever T owns on the heap, and the
box of a T too large to store inline (see the next section). The query still
fails with your message. If that leak matters (a long-lived process whose
queries often fail this way), keep what a state owns small. Only finalize was
measured; errors reported from other callbacks were not. The behaviour is
pinned by
aggregate_states_are_not_all_destroyed_when_finalize_fails in
tests/ffi_roundtrip/lifecycle.rs.
Grouped-aggregate states the scan never reaches are never destroyed (DuckDB defect)
DuckDB 1.4.4 to 1.5.5 destroys a grouped aggregate's states as its result
scan passes them. When the scan stops early, the states it has not reached are
never destroyed — not after the query, not when the connection or the database
closes. That happens under a LIMIT above the aggregate, an error raised
above it, or an interrupt (InterruptHandle::cancel). Measured on one thread
with 300,000 groups: under LIMIT 10, 2,048 of 300,000 states were destroyed
(4,096 on 1.4.x); with an error raised half-way through the result, 151,552.
DuckDB's own aggregates are affected the same way: mode() under LIMIT 10
leaked about 100 MB per query. See docs/upstream-duckdb-reports.md, item 20.
The query's answer is right; the cost is memory. FfiState<T> stores a T
of at most 256 bytes, aligned no more strictly than usize, inside DuckDB's
own state bytes, which DuckDB frees with the hash table, so such a state leaks
nothing unless T itself owns heap memory (a Vec, String or HashMap).
A larger T is boxed, and the box leaks with it. Until DuckDB fixes this,
keep aggregate states small and free of heap allocations where you can.
Pinned by states_a_grouped_scan_never_reaches_leak_no_rust_heap in
tests/aggregate_leaks.rs.
Window frames with EXCLUDE never destroy one state per row (DuckDB defect)
A C API aggregate in a window whose frame has EXCLUDE CURRENT ROW, GROUP
or TIES is evaluated by DuckDB's segment tree in two parts, and the second
part initialises one state per row that is never destroyed: over a 5000-row
window, 5000 states on every release from 1.4.4 to 1.5.5 (none without
EXCLUDE). The answer is right; as above, a small FfiState<T> leaks only
what T owns on the heap. See docs/upstream-duckdb-reports.md, item 35;
pinned by a_window_frame_with_exclude_leaves_states_undestroyed in
tests/ffi_roundtrip/agg_states.rs.
An abandoned stream keeps its table-function state (DuckDB behaviour)
Dropping a streaming QueryResult part-way through, and then its
PreparedStatement, does not free the query's operator states: DuckDB keeps
the active query on the connection until the next statement runs there (or
the connection closes). A table function's state — for a typed table
function, the with_state value and its per-scan clone — therefore lives
until then. Nothing leaks, but a state that holds a file, a lock or a large
buffer holds it for that long; run any statement (SELECT 1) on the
connection to release it. Pinned by
an_abandoned_stream_keeps_its_table_state_until_the_next_statement in
tests/ffi_roundtrip/query_stream.rs.
Running out of memory inside a callback aborts the process
An allocation failure is not an error quack-rs can report. On the Rust side,
the default allocation-error handler aborts. On the DuckDB side,
duckdb_list_vector_reserve, duckdb_vector_copy_sel and
duckdb_vector_assign_string_element_len allocate without catching, so their
std::bad_alloc crosses the Rust callback frame and aborts the process
("Rust cannot catch foreign exceptions"). quack-rs allocates inside
callbacks only on error paths and in data_chunk_to_arrow's pre-export
check, which copies each column it checks. Bound what your callbacks
allocate, and set DuckDB's memory_limit so its own operators fail cleanly
before the process runs out.
One allocation failure is worse than an abort. When duckdb_prepare fails to
allocate while it records a statement's parameter names, it frees the
statement it has already handed back and reports an error; prepare (and
everything built on it) then reads and frees that statement again. This is
undefined behaviour inside DuckDB's C API that no caller can detect; see
docs/upstream-duckdb-reports.md, item 36.
COPY functions (resolved in DuckDB 1.5.0; both directions since)
DuckDB 1.5.0 added duckdb_create_copy_function and related symbols to the public
C extension API. quack-rs wraps these in the copy_function module behind the
duckdb-1-5 feature flag. See CopyFunctionBuilder for usage.
This was previously listed as a known limitation (no C API counterpart prior to 1.5.0).
COPY … FROM was a second, narrower gap: quack-rs wrapped the writing half only.
CopyFunctionBuilder::copy_from now attaches a quack-rs table function as a
format's reader, and a copy function may implement either direction or both — a
read-only format leaves the writing callbacks unset entirely. See the
Copy Functions chapter.
Arrow interop (resolved behind duckdb-1-5-4)
DuckDB's C API has a family of conversion functions (present since 1.4.4) that
move data directly between a duckdb_data_chunk and the Arrow C Data
Interface. quack-rs wraps all eight non-deprecated entries in the arrow module, with no arrow crate
dependency — see the Arrow Interop chapter.
The remaining fourteen Arrow entries in the C API struct are the older
duckdb_query_arrow result API, which lives inside
#ifndef DUCKDB_API_NO_DEPRECATED; they are deliberately not wrapped.
The feature is duckdb-1-5-4 rather than duckdb-1-5 because libduckdb-sys
declared the two Arrow ABI records as opaque zero-sized placeholders until
1.10504.0. The DuckDB functions themselves are present in every release
quack-rs supports, but because the feature implies duckdb-1-5, an extension
built with it needs a DuckDB 1.5.0+ engine.
Callback accessor wrappers (resolved)
quack-rs wraps the callback accessor functions — the C API functions used inside your callbacks to retrieve arguments, set errors, access bind data, and so on:
| Category | Wrapper type | Available |
|---|---|---|
| Scalar function execution | ScalarFunctionInfo | Always |
| Scalar function bind | ScalarBindInfo | duckdb-1-5 |
| Scalar function init | ScalarInitInfo | duckdb-1-5 |
| Aggregate function callbacks | AggregateFunctionInfo | Always |
| Table function bind | BindInfo | Always |
| Table function init | InitInfo | Always |
| Table function scan | FunctionInfo | Always |
| Cast function callbacks | CastFunctionInfo | Always |
| Copy function bind | CopyBindInfo | duckdb-1-5 |
| Copy function global init | CopyGlobalInitInfo | duckdb-1-5 |
| Copy function sink | CopySinkInfo | duckdb-1-5 |
| Copy function finalize | CopyFinalizeInfo | duckdb-1-5 |
Where the C API provides a client context for a callback — scalar bind and
init, table function bind, and the four copy-function callbacks — the wrapper
exposes it as get_client_context, which returns a ClientContext (see the
client_context module).
Complex type creation (resolved)
LogicalType provides constructors for all complex parameterized types:
| Method | Type created |
|---|---|
LogicalType::decimal(width, scale) | DECIMAL(p, s) |
LogicalType::enum_type(members) | ENUM('a', 'b', ...) |
LogicalType::array(child, size) | type[N] |
LogicalType::union_type(members) | UNION(a INT, b VARCHAR) |
LogicalType::list(child) | LIST(type) |
LogicalType::struct_type(fields) | STRUCT(...) |
LogicalType::map(key, value) | MAP(K, V) |
The constructors that take child types (list, array, map, struct_type,
union_type) have _from_logical variants for nested complex types, and each
constructor has a try_ form that returns an error instead of panicking.
Introspection methods (get_type_id, list_child_type, struct_child_count,
decimal_width, etc.) are also available.
VARIANT and GEOMETRY types (resolved — exposed behind duckdb-1-5-3)
The VARIANT type (a self-describing nested value, used for example by
Iceberg v3) entered the C type enum as DUCKDB_TYPE_VARIANT (41) in
DuckDB 1.5.3. GEOMETRY (DUCKDB_TYPE_GEOMETRY, 40) was already present
earlier in the 1.5.x line.
quack-rs exposes these as TypeId::Variant and TypeId::Geometry, gated behind
the duckdb-1-5-3 feature. That feature layers on top of duckdb-1-5 and
requires libduckdb-sys >= 1.10503.0 (DuckDB 1.5.3). The separate gate exists
because VARIANT postdates the duckdb-1-5 feature's 1.5.0 floor; GEOMETRY
is gated with it so that one feature covers both values. Keeping them out of
duckdb-1-5 preserves compatibility for consumers pinned to libduckdb-sys
1.10500–1.10502 (DuckDB 1.5.0–1.5.2).
[dependencies]
quack-rs = { version = "0.18", features = ["duckdb-1-5-3"] }
Neither type yet has dedicated VectorReader/VectorWriter helpers; access
their data via the raw pointer from duckdb_vector_get_data when needed.
Changelog
All notable changes to quack-rs, mirrored from
CHANGELOG.md.
The format follows Keep a Changelog. quack-rs adheres to Semantic Versioning.
Unreleased
0.18.0 — 2026-09-30
This release comes out of a second production-readiness audit (AUDIT.md
section 7) and the three passes that followed it (sections 8 to 10). Each
code defect below was either reproduced against a real DuckDB before it was
fixed or, where nothing could trigger it, derived from DuckDB's source;
AUDIT.md records which (VALIDATED or PROVEN). Each code fix has a
regression test where one could be written, and the trait-bound fixes are
pinned by compile_fail doctests. The exceptions: the fixes AUDIT.md
marks PROVEN only, among them the entry points' NULL duckdb_database*,
AbiPolicy::Warn's eprintln! and catalog lookups in a catalog that is not
"duckdb". The FileHandle drop fix is tested by swapping failing C API
stubs into the dispatch table, since no file system DuckDB ships throws
from Close(). The 32-bit size fixes are unit-tested on wasm32 itself (CI's
wasm job runs the unit tests under node); what DuckDB does with an
oversized allocation there is derived from its source, as no wasm32 build
of DuckDB is available to test against.
Several fixes close holes in the safe API — places where safe code could
cause undefined behaviour, a data race or a process abort — and those needed
signature or trait-bound changes, so this is a breaking release; it follows
0.16.0. Each such entry is marked Breaking:.
0.17.0 was prepared but never published; its changes are included here, and the entries below describe the change from 0.16.0.
A third pass followed before anything was published (AUDIT.md section 8).
It fixed further defects, reproducing them against a real DuckDB wherever that
was possible, added CI gates that compile the Rust examples in the book and the
README and check the book's links, and corrected the documentation's claim that
an aggregate's update never sees NULL rows (see Fixed).
A fourth pass followed that (AUDIT.md section 9). It fixed process aborts,
out-of-bounds reads and writes, and wrong answers. Each was reproduced against
a real DuckDB before it was fixed, and its regression test was shown failing
without the fix; AUDIT.md names the DuckDB versions per finding. The pass
also documented thirteen DuckDB defects in docs/upstream-duckdb-reports.md, each
with a plain-C reproducer.
A fifth pass followed (AUDIT.md section 10). Each code fix has a regression
test shown failing without the fix, except as noted above, and each defect
that involves DuckDB was reproduced against a real DuckDB first, except the
catalog lookup and the 32-bit size fixes, which are PROVEN. It documented
eighteen DuckDB defects (items 20 to 37 of docs/upstream-duckdb-reports.md),
each with a plain-C reproducer run on the releases it names.
Within Added, Changed, Fixed and Security, entries are grouped by the pass that produced them: Fifth audit, Fourth audit, and Earlier passes (the 0.17.0 work, the second audit and the third). Security has no Fifth audit group; that pass's memory-safety fixes are under Fixed. Dependencies is not grouped.
Added
Fifth audit
tests/aggregate_leaks.rs, a test binary with a counting global allocator, andtests/handle_leaks.rs, which bounds the C heap (glibcmallinfo2) across many create-and-drop rounds of 15 handles whoseDropno functional test observed (DbConfig,Value,Appender,TableDescription,OwnedDataChunkandPreparedStatement::parameter_name; withduckdb-1-5alsoExpression,ClientContext,FileOpenOptions,FileSystem,FileHandle,SelectionVector,InstanceCache,ErrorDataandCatalog). It runs only on Linux with glibc.vector::max_child_capacity: the most elementsDuckDBcan hold in a list vector's child buffer, whichListBuildernow respects.- A
value_renderfuzz target (fuzz/, featurelive) that renders arbitrary temporal payloads, alone and nested in lists, through a realDuckDB. - CI:
test-older-enginesalso runs the suite against 1.4.5 (the last 1.4 release) and against 1.5.3 and 1.5.4 with theduckdb-1-5-3andduckdb-1-5-4features: the first releases whose bindings those features compile against (the job pins the bindings to the engine's release, which the test shim requires; aduckdb-1-5-4build needs only a 1.5.0+ engine at run time).tests/append_metadata_cli.rsruns theappend_metadatabinary end to end. - CI: the
mirijob also runs the library tests on a big-endian target (s390x-unknown-linux-gnu, interpreted), which is where the string decoder's byte-order defect shows (see Fixed).
Fourth audit
scalar_bind_callback!/scalar_init_callback!(duckdb-1-5).validate_parameter_name,DUCKDB_UNCALLABLE_KEYWORDS,DUCKDB_UNREFERENCEABLE_PARAMETER_KEYWORDS.value::UNRENDERABLE,callback::EMPTY_PANIC_PLACEHOLDER,callback::panic_c_message,CopyGlobalInitInfo::get_file_path_bytes,scalar::info::EMPTY_ERROR_PLACEHOLDER,aggregate::info::EMPTY_ERROR_PLACEHOLDER.- CI:
test-older-enginesruns the suite against DuckDB 1.4.4 (default features) and 1.5.0 (duckdb-1-5), the oldest release each can load into; every engine-specific defect below went unnoticed without it. A weekly scheduled run; the scaffold job runs the generated project'smake configure release test;check-abi-table.pyfingerprints whole signatures, not names.
Earlier passes (0.17.0, second and third audits)
-
Scalar functions as safe Rust closures.
ScalarFunctionBuilder::map1/map2/map1_str/map2_str/map1_opt/map2_opttake an ordinary closure; parameter and return types come from its signature, NULLs propagate correctly, and a panic becomes a SQL error.VARCHARgets its own constructors so the closure can borrow a&strstraight out of the vector. One indirect call per chunk, not per row. Each returnsResult<TypedScalarFunctionBuilder, _>, which offersname(),volatile()andregister(con)but cannot change the signature (code that passes it to aRegistrarcallsRegistrar::register_typed_scalar), and the trampoline re-checks each chunk's vector types, so the declared result width always matches the closure's. Building one does not call DuckDB, so it works in a unit test withMockRegistrar;registerchecks the types. -
Valuegained constructors for the remaining integer and float widths, the rest of the temporal family,INTERVAL,BLOB,DECIMAL, and the composites (STRUCT,LIST,ARRAY,ENUM;MAPandUNIONbehindduckdb-1-5); there is still none forBITorBIGNUM. Alsois_sql_null(distinct fromis_null, which asks about the handle) andas_enum_index.struct_valuechecks the field count first, becauseduckdb_create_struct_valuetakes no count and reads one value per field of the type.list_value/array_valuetake the element type:duckdb.hcontradicts itself here, and the implementation settles it. The element type may itself be aLISTorARRAY(lists of lists, arrays of arrays); the error explaining the element-type rule appears only when DuckDB itself refuses the value.decimalchecks width (1..=38), scale and the unscaled value's digit count before calling DuckDB, which would otherwise abort the process or store a different number. The temporal constructors other thandateandintervalreturnResultand refuse a payload DuckDB cannot render (see theValue::time_ns/Value::timestampentry under Changed). -
PreparedStatementgained 16 more typed binds andbind_value, the escape hatch for every composite type.bind_decimalvalidates likeValue::decimal:duckdb_bind_decimalchecks nothing, and forwidth <= 18keeps only the low 64 bits of the unscaled value. -
Cancellation and progress.
OwnedConnection::interrupt_handlereturns aSend + SyncInterruptHandle, lifetime-tied to the connection, whosecancela watchdog thread can use to stop a running query and whoseprogressreads itsQueryProgress.OwnedConnection::interrupt/progressdo the same through the connection itself, and theunsafequery::interrupt/query::query_progresstake a raw connection. -
Streaming results.
PreparedStatement::execute_streamingandQueryResult::is_streaming(duckdb-1-5). -
QueryResult::column_logical_typekeeps the nested structure thatcolumn_typecollapses, andresult_kindseparates rows from row counts (the newquery::ResultKind). -
LogicalType::register—CREATE TYPEfrom the C API, so an extension can ship a namedENUMorSTRUCT. Stable-prefix; no feature needed. -
ScalarBindData<T>/ScalarLocalState<T>(duckdb-1-5) — typed bind data and per-thread local state for scalar functions, with theduckdb_delete_callback_tgenerated and panic-safe. The rawset_bind_data/set_stateroute makes the extension author write their ownunsafe extern "C" fnaroundBox::from_raw, which is the abort hazard under Security relocated into user code. DuckDB reads scalar bind data from every executing thread at once, soScalarBindData<T>requiresT: Send + Sync + 'staticandScalarLocalState<T>T: Send + 'static.ScalarBindData::setalso requiresT: Clone(wrap other data inArc<T>): it registers a generated, panic-safe copy callback, without which the bind data is lost whenever the optimizer copies the bound expression — a wrong answer, not an error (Pitfall L10; seeScalarBindInfo::set_bind_data_copyunder Fixed). A secondsetleaks, rather than drops, the first value. -
vector::ops(duckdb-1-5) makesSelectionVectorusable:copy_selected,slice,reference_value,reference_vectorandOwnedVector. Documents thatsliceproduces a dictionary vector, after which every reader in this crate reads the wrong rows.OwnedVector::newwalks the type, child vectors included, with checked arithmetic and refuses a capacity abovevector::ops::MAX_CAPACITYbefore DuckDB allocates: DuckDB computes the buffer size with an unchecked multiply, so(HUGEINT, 2^60 + 2)would get a 32-byte buffer. -
Arrow C Data Interface bridge — the new
arrowmodule behind a newduckdb-1-5-4feature, wrapping the eight-function conversion family already in DuckDB 1.4.4's C API:ArrowOptions<'conn>, which borrows its connection so that a conversion cannot read freed memory after the connection closes (the safefrom_connection(&'conn OwnedConnection), and theunsafefrom_raw_connection,from_resultandQueryResult::arrow_options), owningArrowSchema/ArrowArray/ArrowConvertedSchemaRAII types, andto_arrow_schema/data_chunk_to_arrow/schema_from_arrow/data_chunk_from_arrow. Noarrowcrate dependency: the module works on the ABI recordslibduckdb-sysdefines, which have arrow-rs'sFFI_ArrowSchema/FFI_ArrowArraylayout, so bridging is a pointer cast —ArrowArray::take_frommoves a record out of a foreign wrapper and leaves a released placeholder, so only one side ever callsrelease.The ownership rules were read out of
arrow-c.cppandarrow_converter.cpprather than inferred:duckdb_data_chunk_from_arrowsetsarrow_array->release = nullptrbefore the conversion loop body, so it claims the array on the error path too — hencedata_chunk_from_arrowtakes the array by value, and the by-value binding still releases it in the one case (a zero-column schema) where the loop never runs.to_arrow_schemaanddata_chunk_to_arrowinstallreleaselast, after everything that can throw, so a failed conversion leaves nothing to free.Two crashes DuckDB does not guard are refused here instead:
duckdb_data_chunk_from_arrowindexesarrow_array->children[i]once per schema column with no bounds check and dereferences an already-released array, soArrowConvertedSchemaremembers its column count and both are checked first.data_chunk_from_arrowalso returnsInvalidInputfor a negative length, a nonzero top-level offset (DuckDB ignores it and imports the wrong rows), a nullchildrenpointer, a null child, and a child shorter than the array's length, each of which DuckDB would dereference, read out of bounds or misread, and for a zero-row array, which DuckDB passes on as a zero-byte allocation that a debug build asserts against. It refuses a length, or a nested row count, abovevector::ops::MAX_CAPACITY, and the valid Arrow layouts DuckDB imports from the wrong rows or out of bounds (seesrc/arrow/import_layout.rsand the Fifth audit entries under Fixed). Its Safety section requires that the array conform to the schema — nothing in an Arrow array records its type, so a child whose buffers do not match the type its schema declares cannot be checked — and that the length be the array's true row count: DuckDB allocates the chunk before itstryblock, so a length below that ceiling that the system cannot allocate still aborts the process. Dictionary-encoded and null-type columns are copied into flat vectors before it returns; run-end-encoded children are expanded by DuckDB. TIMETZ comes back as TIME without its offset and BIT as BLOB.The
duckdb-1-5-4feature's floor is set by the bindings, not by DuckDB: the eight functions predate 1.5, and the feature needs only the 1.5.0+ engine thatduckdb-1-5does at run time, butlibduckdb-sysdeclaredArrowSchema/ArrowArrayas opaque zero-sized bindgen placeholders until 1.10504.0.src/arrow.rscarries aconstassertion that says so. -
COPY … FROM—CopyFunctionBuilder::copy_fromattaches a quack-rs table function as a format's reader, so an extension can implement loading as well as writing. Supporting pieces:TableFunctionBuilder::build_handlereturns a configured, unregisteredTableFunctionHandle—registeris now that plusduckdb_register_table_function— because aCOPY … FROMreader is attached to a copy function rather than registered on its own.BindInfo::result_column_count/result_column_name/result_column_type(duckdb-1-5) read the target table's schema, whichCOPY … FROMfixes before the bind callback runs.duckdb.his explicit that such a bind "should not define its own result columns".CopyBindInfo::optionsexposes theCOPY … TOoptions as the STRUCT value DuckDB builds, andValue::struct_field_nameswalks the field names off the value's borrowed logical type without exposing the handle (pitfall P11); for a UNION it returns an empty name for the tag, then the member names.CopyFunctionBuilder::extra_info, with the same ownership-until-transfer guarantee as the other builders.
duckdb_copy_function_set_copy_from_functionreports every rejection by doing nothing at all, socopy_fromchecksduckdb.h's stated precondition — "the table function must take a single VARCHAR parameter (the file path)" — which DuckDB never enforces:CCopyFromBindbuilds the argument list itself and never consultstf.arguments, so a mismatch surfaces much later inside the reader's own bind callback. -
TableFunctionBuilder::with_bind_init(bind, init): immutable bind dataB: Send + Syncand a fresh scan stateS: Sendper execution, noCloneneeded. -
TypedScalarFunctionBuilder;Registrar::register_typed_scalar(with a default implementation, so existingRegistrars compile). -
StructWriter::set_row_null. -
callback::drop_panic_payload,callback::take_panic_messageandcallback::MAX_NESTED_PAYLOAD_DROPS. -
selection_vector::MAX_LEN;datetime::is_valid_date,MICROS_PER_DAY,TIME_TZ_MAX_OFFSET_SECONDS,DECIMAL_MAX_WIDTH. -
CatalogEntryType::is_lookup_supported;ReplacementScanInfo::EMPTY_ERROR_PLACEHOLDER. -
Aggregate function sets support a different return type per overload (#121).
DuckDBresolves an aggregate overload from its parameter types and arity only — the return type takes no part in resolution — so members of one set are free to return different types, which is howDuckDB's ownarg_max(ANY, ANY) -> ANYandarg_max(ANY, ANY, ANY) -> ANY[]coexist. quack-rs had the return type onAggregateFunctionSetBuilderalone, which made that impossible to express.AggregateOverloadBuildernow carriesreturns/returns_logical, andAggregateFunctionSetBuilder::overload(..)takes one fully-configured overload — mirroringScalarFunctionSetBuilder::overload/ScalarOverloadBuilder, which already worked this way:#![allow(unused)] fn main() { AggregateFunctionSetBuilder::new("my_agg") .overload( AggregateOverloadBuilder::new() .param(TypeId::Integer) .returns(TypeId::Integer) // ... callbacks ) .overload( AggregateOverloadBuilder::new() .param(TypeId::Varchar) .returns(TypeId::Varchar) // ... callbacks ) .register(con)?; }This is additive.
returns/returns_logicalon the set now act as a default for every overload that does not set its own, so existingreturns(..).overloads(range, ..)code keeps working unchanged. As for return types, registration fails, naming the overload index, only when an overload has neither its own nor a set-level default; a missing callback, a duplicate signature or a compositeTypeIdfails it too. -
Unit tests asserting that every
AggregateOverloadBuildercallback setter stores into its own field. These setters areconst fn, which makes their cargo-mutants mutants unviable rather than caught —Default::default()cannot be called in a const context, so the replacement fails to compile and the mutation gate is structurally silent about them. That is a property of the gate, not evidence the setters work. -
An AddressSanitizer job (
ci.yml; blocking — it was informational until its first green run).leak-checkanswers "did we forget a destructor"; ASAN answers "did we write outside an allocation, or use one after free" — the class behind the two heap-corruption defects fixed in v0.16.0, and the one a crate doing raw pointer arithmetic into DuckDB's memory is most exposed to. Miri cannot reach those paths: they call foreign functions. -
scripts/duckdb-version-from-lock.sh— derives the DuckDB release tag from thelibduckdb-syspin inCargo.lock(1.10505.0 → v1.5.5). Two CI jobs hard-codedv1.5.4next to a comment asking the next person to keep it in sync with the dependency; bumping the lockfile would have left both linking a libduckdb one release older than the bindings being generated against it.test-older-engineschecks its output against the engine each of its five matrix entries names, two of them in the pre-1.5 scheme. -
scripts/sync-book-changelog.py— the book's changelog page is now generated fromCHANGELOG.md, with a--checkmode wired into thedocjob. Kept by hand, the mirror had fallen ~19 KB behind and the published 0.16.0 entry was missing its entire "Portability and feature-combination breakage" subsection. The deliberate differences — the book page's own preamble, an em dash in release headings, and links to repository files rewritten as GitHub URLs — are applied by the script. -
timeout-minuteson everyci.ymljob (36 at this release). The default is 360 per job, so a hang intest-bundled(which compiles DuckDB from C++ source) orleak-check(-Zbuild-std) burned six hours of runner time. -
First end-to-end coverage of the aggregate function-set registration path (
tests/ffi_roundtrip.rs). Neither the aggregate nor the scalar set builder had an E2E test, despite Pitfall L6 — a set member whose name is unset is dropped silently. Four new tests register a real three-overload set (BIGINT -> BIGINT,VARCHAR -> VARCHAR,(BIGINT, BIGINT) -> DECIMAL(18,2)viareturns_logical), asserttypeof(..)per overload, check the computed values across multiple chunks and underGROUP BY, and cover the set-level default and both rejection paths. -
LogicalType::try_get_type_id, which returnsNonefor a type id this crate does not know;get_type_idpanics there, and now documents that it does. -
A fallible form of every
LogicalTypeconstructor: the newtry_decimal,try_array,try_array_from_logical,try_list_from_logicalandtry_map_from_logicaljoin the existingtry_*functions. Alsotypes::logical_type::MAX_UNION_MEMBERS(255: DuckDB 1.5.6 lowered its limit from 256, and asserts it when building the type). -
VectorWriter::try_write_varchar/try_write_blob, which return an error and write nothing for a value longer than the newvector::string::MAX_STRING_LEN(u32::MAXbytes, DuckDB's string length limit); alsovector::string::check_string_len. -
ListBuilder::with_element_limit, to cap list lengths that come from input at what fits in memory. Staying under the builder's own ceiling is not enough —MAX_LIST_CHILD_CAPACITYbytes per child buffer, whichmax_child_capacityturns into an element limit for the child's type (see the Fifth audit entries under Fixed):duckdb_list_vector_reservehas notry/catch, so a failed allocation below that ceiling also aborts the process.ListVector::reservenow documents this. -
vector::ops::MAX_CAPACITY, the largest capacityOwnedVector::newaccepts (2^37 elements, DuckDB'sMAX_VECTOR_SIZE, on a 64-bit target; 2^28 - 1 on a 32-bit one). -
MockVectorWriter::set_validandMockVectorWriter::is_written(see theMockVectorWriterentry under Changed). -
CatalogEntryType::may_autoload_extension, the name check behind the new catalog-lookup refusal (see Changed). -
An
EMPTY_ERROR_PLACEHOLDERfor table functions (table::info::EMPTY_ERROR_PLACEHOLDER), copy functions (copy_function::info::EMPTY_ERROR_PLACEHOLDER) and casts (CastFunctionInfo::EMPTY_ERROR_PLACEHOLDER), andcallback::CAST_FAILED_WITHOUT_MESSAGE(see Fixed). -
InstanceCacheisSend + Sync. DuckDB's instance cache guards its own state with a mutex, so the wrapper was!Sendonly because it holds a raw pointer. -
WarningSeverityimplementsPartialOrdandOrdin the orderInfo < Low < Medium < High < Critical, soseverity >= WarningSeverity::Highselects the warnings that need attention. -
ScalarOverloadBuildergainsvolatile,varargs,varargs_logicaland, withduckdb-1-5,bind/init;AggregateOverloadBuildergainsextra_info. Each behaves as it does on the single-function builder, including who freesextra_infowhen registration fails. -
The prelude exports
ScalarBindDataandScalarLocalState(withduckdb-1-5, besideScalarBindInfo/ScalarInitInfo). Its documentation now lists every re-export, and a unit test fails when one is missing. -
validate_spdx_licenseaccepts<license> WITH <exception>(for exampleApache-2.0 WITH LLVM-exception), which it used to reject. The exception must be on the SPDX exception list, now public asvalidate::spdx::SPDX_LICENSE_EXCEPTIONS, or anAdditionRef-id. -
classify_extension_versionaccepts one leadingvon the semantic-version forms (v1.0.0): the spelling DuckDB's versioning documentation uses, and what extension-ci-tools stamps whenHEADcarries avX.Y.Ztag. -
Pitfall L12 (
LESSONS.md,book/src/reference/pitfalls.md): an aggregate'supdatereceives NULL rows under the default NULL handling too. See Fixed. -
Projects generated by
generate_scaffoldtest something. The generated project'scargo testran zero tests and itstest/sql/<name>.testheld onlyrequireand commented-out examples, so both CI steps passed whatever the extension did. The project now has a unit test and a SQLLogicTest that queries<name>_helloand checks its output. The generatedMakefilechecks thatextension-ci-toolsis checked out and names both commands:git submodule addin a new repository (where the documentedgit submodule update --initclones nothing) andupdate --initin a clone (Pitfall P4). -
The book and the README are compiled as doctests.
book/doctestis a standalone crate that turns every page underbook/src(except the changelog) andREADME.mdinto rustdoc input, so every Rust block not markedignorecompiles against the working copy, and runs unless it is markedno_run;test-bundled-prebuiltruns it. Each remainingignoreblock says why.mdbook testcannot do this: it passes no--extern quack_rs, so no block that uses the crate resolves it. The blocks that failed, and the content errors that turned up, are listed under Fixed.release.yml's gate also runs the crate's own doctests now (cargo test --doc --features duckdb-1-5-4). -
New CI jobs and checks (
ci.yml):bookbuilds the book with mdBook on every pull request, not only after a merge tomain, then runsscripts/check-book-links.py, which checks every relative link and anchor in the book and README and every docs.rs link against this checkout's rustdoc. mdBook checks neither.autoload-entriesrunsscripts/check-autoload-entries.py, which fails when a DuckDB release from v1.5.0 on autoloads an extension for a type or collation name missing from the lists behind the catalog-lookup refusal.abi-guard-layoutexercises the realLayoutMismatchpath: aduckdb-1-5build stampedC_STRUCTmust be refused by DuckDB v1.5.0 and v1.4.4 and must load into the release its bindings come from. The existingabi-guardjob only ever reached the declared-version check.dependency-floorresolvesduckdb/libduckdb-systo the declared 1.4.4 floor and runscargo test --lib, on stable: at that floor the dependency tree needs a newer rustc than the 1.86.0 MSRV.extension-loadrunsscripts/check-hello-ext.py, which executes every statement documented inexamples/hello-ext/README.mdagainst the loaded extension and compares the output withexamples/hello-ext/sql_checks.txt.semverfails when cargo-semver-checks ran 0 checks against the published baseline for a bump that is not breaking; before, 0 checks passed. For a breaking bump it runs none by design, so a new informational step compares againstorigin/mainwith--release-type patchand lists every breaking change on the branch in the log and the job summary.- Integration tests fail when a documented "N pitfalls" count differs from
LESSONS.md, when the source trees inCONTRIBUTING.mdand the book miss a file undersrc/ortests/or list one that does not exist, and when a documentation paragraph mentionspanic = "abort"without warning against it.
Changed
Fifth audit
- Breaking:
FfiState<T>stores a smallTinDuckDB's state bytes. ATof at most 256 bytes, aligned no more strictly thanusize, is kept inline; a larger one is still boxed. In 0.16.0FfiState<T>was a one-word#[repr(C)]struct with a publicinner: *mut Tfield that always boxedT; its layout is now private — a tag derived fromT'sTypeId(see Security), thenTor the box. Use its callbacks andwith_state/with_state_mutas before. See Fixed. - Breaking:
AggregateStaterequiresSync. A window's segment tree shares its states between threads ascombinesources. A state type holding aCellorRefCellmust switch to atomics or aMutex, or keep that data outside the state. - Documentation: bind-time arguments are seen before the cast to the
parameter type (
ScalarBindInfo::argument); a Safety clause ondata_chunk_from_arrow(validity bitmaps are read one byte past their rows); Known Limitations entries for aggregate statesDuckDBnever destroys, abandoned streams and out-of-memory aborts. Value::as_str,display_stringandDebugrender only types whose every payloadDuckDBcan render; everything else (VARIANT,GEOMETRY, a type quack-rs does not know, andARRAY/UNIONvalues that could hold one) getsUNRENDERABLE. See Fixed.- Breaking: scalar function bind and init callbacks take their own
argument types,
RawScalarBindInfoandRawScalarInitInfo(#[repr(transparent)]overduckdb_bind_info/duckdb_init_info), inScalarBindFn,ScalarInitFn,scalar_bind_callback!,scalar_init_callback!,ScalarBindInfo::newandScalarInitInfo::new. A table function's bind and init callbacks receive the same C types, butDuckDBcasts them to a different, larger struct, sotable_bind_callback!output registered on a scalar function (which safe code could do) wrote past the scalar bind info when it reported a panic. The two kinds no longer type-check in each other's slots (compile_faildoctests). A hand-written raw scalar callback changes its parameter type only. AggregateFunctionBuilder::ffi_state::<T>()andAggregateOverloadBuilder::ffi_state::<T>()installFfiState<T>'sstate_size,initanddestructorcallbacks together. Set one by one, a size callback for oneTwith an init callback for a larger one wrote pastDuckDB's allocation; the setters remain, and the aggregateregistermethods andRegistrarnow state that pairing as a Safety obligation (theRegistrardocs said a call on the entry point's connection was "always sound"). The README and prelude examples use the new method; the prelude's had no destructor, so it leaked every state.- Breaking:
ArrowConvertedSchema::from_rawtakes the Arrow schema the handle was built from (&ArrowSchema) instead of a column count, and is no longerconst:data_chunk_from_arrowchecks each array against that schema's shape (see Fixed). data_chunk_to_arrowrefuses a chunk holding a valueDuckDBwould export as a different value (see Fixed).- Breaking, not named until now (found by
cargo semver-checksagainst 0.16.0, run as a patch release so that it reports every break):AppenderandConnectionare no longerRefUnwindSafe(they gained aCelland aRefCellin earlier passes: the appender's row bookkeeping and the collision check's catalog snapshot), andvalidate::description_yml::DescriptionYmlis#[non_exhaustive], so it cannot be built with a struct literal. Forcatch_unwind, wrap a closure that captures anAppenderor aConnectioninAssertUnwindSafe. - Breaking: a callback macro's body must have type
()(exceptcast_callback!'s, which returns the cast'sbool). The body runs insidecatch_unwind, and its value was discarded: a body that used?compiled, and the error it returned was dropped without being reported. It is now a type error; report the error through the callback'sset_error. - Breaking: a callback macro's body is no longer an
unsafecontext. The body was a closure inside the generatedunsafe extern "C" fn, and a closure inherits its function's unsafe context, so a raw-pointer dereference or anunsafe fncall in a "safe" callback compiled with nounsafekeyword and, on edition 2021 (which the scaffold generates), no warning. The body is now expanded as a nested ordinaryfn; wrap unsafe operations inunsafeblocks, as the documented examples already do.compile_faildoctests pin it. - Breaking: catalog lookups are refused in a catalog
DuckDBdoes not implement itself. For a catalog a storage extension attaches,duckdb_catalog_get_entrystarts that extension's transaction and runs its schema lookup with notry, so an exception there would abort the process.CatalogEntry::lookupandCatalog::get_entrynow return an error for any catalog whose type is not"duckdb".DuckDB's own catalogs (the database's,tempandsystem) all have that type and are unaffected. Look up entries in another catalog with SQL instead, as the error says. CombineFndocuments thatcombinemust leave its source states unchanged (see Fixed), and no longer advises moving out of them.Catalog::type_nameno longer gives"system"as an example type: thesystemandtempcatalogs are of type"duckdb".data_chunk_from_arrow's Safety section requires a fixed-width dictionary's values buffer to be readable one element past its length when the indices can be NULL:DuckDBpoints NULL indices at a sentinel entry there, and the flattening copy reads it (found with an AddressSanitizer-builtlibduckdb; item 29). Its Errors section now lists the layout refusals.- Documentation corrections from a mechanical check of the docs' universal
and numeric claims: the Arrow conversion functions are already in
DuckDB1.4.4's C API (not added in 1.5.0); the unstable API region did not change "in every recent release" (1.4.4 and 1.4.5, 1.5.0 and 1.5.1, and 1.5.2 to 1.5.5 are byte-identical); it was not changed by middle insertions "in four of the last four" versions (two). LESSONS.mdand the book's pitfall catalogue gained L15 (combinemust leave its source states unchanged), L16 (a valid Arrow array is not always oneDuckDBimports correctly), L17 (aCOPY … FROMreader must not declare result columns) and L18 (aLISTreserve moves every buffer below its child): 30 documented pitfalls. The README's and the book's summary tables, which stopped at L14, list all 30.DestroyFn,FfiStateand the book's known limitations document thatDuckDBnever destroys one aggregate state per row of a window frame withEXCLUDE(item 35: 5000 of a 5000-row window, every release from 1.4.4).DestroyFn's docs said it was called for every stateDuckDBcreated.query::prepareand the book's known limitations document that an allocation failure insideduckdb_preparecan hand back a statementDuckDBhas already freed, which the error path then reads and frees (item 36: the last 3 of its allocations, every release from 1.4.4, found by failing each allocation in turn).ClientContext::config_optiondocuments thatDuckDB's function has notry, and why no built-in setting's getter throws through it in practice.FfiInitData::setandFfiLocalInitData::setname the callback each must be called from: both store throughduckdb_init_set_init_data, so the wrong one sets the other kind of init data, which the othergetthen reads as the wrong type.FfiBindData::setsays a typed table function's bind sets the bind data itself.ReplacementScanBuilder::registersaysDuckDBcallsdelete_callbackeven whenextra_datais null, unlike its other destructor slots (a new test observes the call).data_chunk_from_arrowsays which column holds its claim on the Arrow array:DuckDBgivesreleaseto column 0 alone, so a vector made to reference another column (reference_vector) does not keep the producer's buffers alive (a new test observes both cases).StructWriter'swrite_*,set_nullandset_validstate in their Safety sections that no field writer was replaced throughfield_mut: swapping two writers is safe code, and the next write would go to the wrong child vector.OwnedConnection'sSendjustification saidDuckDBforbids concurrent use of one connection; it serialises it. The comment now gives the real reasons, and a test queries and drops a connection on another thread.docs/upstream-duckdb-reports.mdgained items 20 to 37, and item 16 gained aHUGEINTreproducer.- Test gaps the mutation sweeps exposed. The full sweep left 22 mutants
alive, and the end-to-end run over the files
mutants.tomlexcludes left more. Each is now killed by a test, excluded with the reason inmutants.toml, or recorded inAUDIT.mdsection 10 as equivalent with the argument. The new tests coverFfiState's tag structure, a release profile withoutpanic = "unwind", an overload's own return type, every typedValuegetter against a value of its own type, theTIMETZandTIME_NSrange guards, theAppendermethods a no-op replacement survived (another schema,column_type,clear_columns,append_default_to_chunk),StructWriter's child vectors,InMemoryDb::execute's row count,QueryResult::result_kindfor a statement that returns nothing, a createdErrorData,MockVectorWriter::len, andappend_metadata's handling of=in a path, a lone-, signed version numbers, commit hashes of the wrong length, case or alphabet, and a footer whose magic field is wrong.tests/handle_leaks.rskills theDropmutants that survived every functional test. Two comparisons moved into smallconst fns so that a unit test can reach their boundary: the aggregate-state salt and the appender'su32::MAX-inclusive length limit. - The ABI layout table covers
DuckDBv1.5.6, which shares the 546-slot layout of v1.5.2 – v1.5.5. Without it aduckdb-1-5extension was refused by v1.5.6 asUnknownEngineVersionunder the defaultAbiPolicy::Strict. v1.5.6 declares every slot stable for extensions targeting C API v1.5.6; quack-rs still targets v1.2.0, which v1.5.6 loads. scripts/check-abi-table.pyreads the v1.5.6 header correctly: it evaluates the newDUCKDB_API_VERSION_AT_LEAST(...)bands, no longer folds a preprocessor continuation line into the first declaration's fingerprint, and checks that the firstSTABLE_API_SLOT_COUNTdeclarations are identical in every release (the v1.4.0varint→bignumrename aside) instead of requiring an unchanged stable count, which v1.5.6 raised from 357 to 404.- CI runs what the release workflow runs before a tag is pushed. The first
v0.18.0tag failed on Windows whilemainwas green, because no PR job ran theduckdb-1-5*tests on macOS or Windows:test-bundlednow also runscargo test --all-targets --all-featureson all three, and on Linuxcargo doc --all-features, the only rustdoc pass over the test-only items. Thetestjob runs the unit tests with Cargo's default release codegen (16 codegen units, no fat LTO), the build that exposed theFfiStatesalt defect (see Fixed); every other build is debug or has one codegen unit and fat LTO. The release workflow's test steps andtest-bundled's pass--no-fail-fast, so one failing test binary no longer hides failures in the ones after it. release.yml:gh release create --verify-tag, so a missing tag fails instead of being created at the default branch's HEAD; a re-run ofpublishrecognises current cargo's "already exists on crates.io index" (it matched only the older "already uploaded"); every job has atimeout-minutes.- The unit tests run on wasm32, the one 32-bit target quack-rs supports: the
wasmjob installs emsdk 6.0.10 and runscargo test --libunder node, with and withoutduckdb-1-5-4. Before, it only compiled for wasm32, so the 32-bit size limits were only ever tested with 64-bit values. Two tests assumed a 64-bit target and were corrected.criterionis now a non-wasm dev-dependency (its rayon dependency does not build for wasm32). tests/file_handle_close.rstestsFileHandle'sDropwhen the close fails, by swapping stubs into the C API dispatch table; it fails on the oldDrop, which destroyed without closing first.tests/ffi_roundtrip/file_errors.rsruns on every platform: a failed write through a read-only handle, and a seek pasti64::MAX, which is asserted to be refused withInvalidInput. Only the/dev/fullsync failure, which has no portable trigger, stays Linux-only. The over-4 GiB string test runs on macOS as well as Linux..gitattributes: text files are LF in every checkout, so Windows CI tests the same bytes as Linux; fuzz seeds and images are binary.- The book's aggregate examples and
hello-extregister state withffi_state::<T>()instead of wiringFfiState's three callbacks by hand. - The book was reviewed page by page against the source. Corrected, among
others: pages that said every extension binary is tied to one DuckDB
release and should pin
libduckdb-sysexactly (true only for builds that use the unstable API), an out-of-boundsChunkWriterexample, the instance cache's lifetime (it holds weak references),AbiPolicy::Warn(it prints to stderr), and links that rendered as literal brackets. - The site's custom
<head>was never rendered (book.tomldid not point mdBook atbook/theme), so every page shipped an empty description and no canonical or Open Graph tags.scripts/seo-postbuild.py, run bydocs.ymland by CI'sbookjob, now writes a canonical URL and a specific description to each page and fails on a missing or duplicate one; the preview image is a PNG; headings are no longer rendered at half opacity; diagrams follow the theme; mermaid is pinned to 11.17.2. - The crate, README, book, logo and social-preview image no longer call
quack-rs "production-grade" or the logo "FAST"; neither was backed by
evidence, and the audits behind this release found serious defects. The
crates.io description is now "Rust SDK for building DuckDB loadable
extensions on the DuckDB C Extension API". A new README Status section
and the FAQ state what the record shows instead: pre-1.0 API churn, the
defects each audit found (
AUDIT.md), DuckDB's own limits, and what CI checks.
Fourth audit
- CI's "the refused extension registered nothing" checks could not fail
(
duckdb -cstops at the failedLOAD); they feed the statements on stdin and require a marker row. AbiPolicyhandling is a purepolicy_verdict, tested against every check result. The docs ofenforce_abi_policyno longer sayWarngoes throughset_error.- Arrow export documents the INTERVAL and UHUGEINT values
DuckDBcorrupts. docs/upstream-duckdb-reports.mdgained items 7 to 19.
Earlier passes (0.17.0, second and third audits)
-
Breaking:
CopyFunctionBuilderis no longerSendorSync. SupportingCOPY … FROMgave it two new fields that each carry a raw pointer —copy_from: Option<TableFunctionHandle>(aduckdb_table_function) andextra_info: Option<ExtraInfo>(a*mut c_void) — and a raw pointer is neitherSendnorSync.cargo-semver-checksclassifies this asauto_trait_impl_removed, a major break, one of the reasons this release bumps the minor version: for a pre-1.0 crate Cargo treats the leftmost non-zero component as the major, so the minor position is where a break goes (RELEASING.md, "Semantic versioning policy").The change aligns the type with its three siblings rather than making it an outlier —
ScalarFunctionBuilder,TableFunctionBuilderandAggregateFunctionBuilderwere already!Send + !Syncin 0.16.0, each because it holds a rawDuckDBhandle. A builder is a short-lived object constructed and registered insideduckdb_init_c_api, and noDuckDBC API accepts one from another thread, so the traits were never usable for anything real. Code that moved aCopyFunctionBuilderbetween threads must now construct it on the thread that registers it. -
Flaky tests: a shared counter raced across parallel tests. The
extra_infotests and the newarrowones each reset onestatic AtomicUsizeand then asserted on it, butcargo testruns tests in parallel — so one test's reset could land between another's drop and its assertion.dropping_an_untransferred_extra_info_frees_itfailed in CI while the identical job on the identical commit passed. Each test now owns its counter:extra_infocarries it in the allocation, andarrowcarries it in the record's ownprivate_data, which is what that field is for. Verified with 100 repeat runs, zero failures. -
128-bit splitting and reassembly in
ValueandPreparedStatementwas open-coded.Value::as_i128/as_u128/as_uuid/as_decimal/uuidandPreparedStatement::bind_i128/bind_u128/bind_decimaleach did their own<< 64/>> 64word arithmetic, where being silently wrong is easy — a shift in the wrong direction still compiles, still round-trips zero, and still round-trips anything that fits in 64 bits. Those call sites now route through fourpub(crate)helpers (hugeint_from_i128/hugeint_to_i128/uhugeint_from_u128/uhugeint_to_u128) that live next to each other and are unit-tested in both directions, at the extremes and against hand-built records. Mutation testing confirms all seven shift mutants across the three modules are now killed. The appender,datetimeand the vector reader and writer still split and reassemble 128-bit values themselves. -
Test gaps the mutation sweep exposed, once it could see the files. Chief among them the shift direction in
hugeint_from_i128/uhugeint_from_u128: swapping>>for<<still compiles, still round-trips zero and still round-trips anything that fits in 64 bits, so nothing in the suite noticed. AlsoTypeId::composite_constructor_hint's per-variant arms,LogicalType::check_slot's rejection path,composite_message,LogicalTypeError::api_func, thearrowaccessors against a populated record rather than only an empty one, andmap2_str's NULL propagation when just one argument is NULL.With the three gate defects fixed and the two FFI-wrapper modules excluded, the incremental sweep over this branch's 40 changed source files reports 404 mutants — 250 caught, 154 unviable, none missed. What survives after that is annotated in the source with
#[mutants::skip]and a reason, rather than filtered out of sight: the bare FFI reads (DataChunk::size/column_count,ScalarBindData::set,ScalarLocalState::set), twoDropimpls whose effect is only visible in freed memory or to a leak checker (OwnedVector,SecretEntry), the deprecatedFfiBindData::get_from_bindwhose mutant is the function (it returnsNoneunconditionally, becauseDuckDBhas noduckdb_bind_get_bind_data),map2/map2_strwhose per-row NULL check only runs insideDuckDB's expression executor, andValue::as_str_or_default, whose null-handle answer is exactly the mutant'sString::new().Two of those turned into real work rather than an annotation.
SecretEntry's zeroize-on-drop is a security property the crate advertises and nothing asserted — its body is nowSecretEntry::zeroize_in_place, tested directly, withDropleft as a one-line delegation. Andhugeint_to_i128combined its halves with|; because the halves occupy disjoint bits,|→^cannot change the result for any input, so that mutant was unkillable by construction. The halves are added instead: identical here, incapable of overflowing (i64::MIN << 64is exactlyi128::MIN, and the round trips at the extremes would panic in a debug build if that were wrong), and-or*in its place dies at once. -
The mutation-testing gate always reported 100% and always passed.
cargo mutants --output DIRwrites its results toDIR/mutants.out/, so--output mutants.output them inmutants.out/mutants.out/. The report step countedmutants.out/caught.txtand friends, found nothing, computedSCORE=100%fromTOTAL=0, and skipped itsexit 1becauseMISSEDwas also 0. The run on this PR printedMUTATION SCORE: 100%and passed while cargo-mutants' own summary line in the same log read229 missed, 238 caught, 218 unviable. Both jobs now pass--output .. -
Incremental mutation testing skipped every top-level
src/*.rs. The job selected changed files with the pathspecsrc/**/*.rs, but git's default wildmatch lets*cross/, so that pattern requires at least one directory component aftersrc/and matches no top-level file at all.src/value.rs,src/query.rs,src/appender.rsand every other module directly undersrc/were silently excluded while the job reported success. It now filters to.rsin the shell. -
The mutation gate's Display/Debug filter matched nothing.
mutants.tomlcarriedexclude_re = ["^fmt::"], commented "Display/Debug impls". cargo-mutants matches that regex against the whole line it prints for a mutant —src/x.rs:1: replace <impl core::fmt::Debug for T>::fmt -> … with …— which begins with the file path, so a pattern anchored atfmt::can never match and seventeen unkillableDebug::fmt -> Ok(Default::default())mutants survived every sweep. cargo-mutants' own documented form,impl Debug, does not match this crate either: it writesimpl core::fmt::Debug for T, and the qualified path lands in the mutant name. The pattern is nowimpl [a-z:]*Debug for.Displayis deliberately not excluded — those impls render error text that tests assert on, so their mutants die and belong in the gate.The incremental job now also re-applies
exclude_refrommutants.tomlas CLI flags, the way it already did forexclude_globs: cargo-mutants only combines a CLI--exclude-rewith the config file from 27.0.0 onwards, and that job passes one. -
src/value.rsandsrc/query.rsare excluded from mutation testing, and their pure logic moved out so that it is not. Every function left in those two files wraps aDuckDBC call, which the gate's--cargo-arg=--librun — no live engine — structurally cannot reach; that is the same rationalemutants.tomlalready carried for nine sibling modules, and between them the two files accounted for 206 of the 220 survivors. Excluding them wholesale would have swallowed the pure code too, so it moved to siblings that stay in the gate:src/value/hugeint.rs(the four 128-bit word helpers),src/value/defaults.rs(the fourteenas_*_oraccessors, as a second inherentimpl Value) andsrc/query/cstr.rs(to_c_sql,c_str_to_owned). Seven of theas_*_oraccessors had no unit test at all, and nothing exercisedc_str_to_owned's non-null path — all of them were among the 220 — so nine new tests turn those survivors into kills rather than hiding them. No public item moved: the re-export keeps every existing path. -
Four CI jobs never actually ran.
rust-toolchain.tomlpinschannel = "stable", and a rustup toolchain file overrides the defaultdtolnay/rust-toolchainsets — so themiri,leak-check,fuzzandnightlyjobs all resolved a barecargoto stable. Miri and LeakSanitizer failed loudly (the 'miri' component ... is not available for the 'stable-...' toolchain;the -Z flag is only accepted on the nightly channel); the informational nightly job failed silently, re-testing stable. All four now invokecargo +nightlyexplicitly, with a comment saying why.The
test-bundled-prebuiltjob's clippy step was also missing theenv:block its sibling test step has, sobuild.rspanicked looking forduckdb.hppbefore clippy ran. -
cargo docwith-D warningsfailed. Six intra-doc links were broken or redundant —QueryResult::column_logical_type,scalar::typed,vector::ops(twice),TableFunctionBuilder::build_handle— andvector::opscarried a doc comment on both thepub moddeclaration and the file's own//!header, so its module docs were resolved in the parent's scope and every link into the module failed with no source location. Three more links inappenderandtable_descriptionpointed atduckdb-1-5-gated methods and so broke the default-feature doc build.cargo docis now clean under-D warningson all four feature sets. -
Pitfall L9 —
duckdb_data_chunk_from_arrowtakes the array even when it fails.duckdb.hreads like a success-path statement;arrow-c.cppnullsarrow_array->releaseinside the per-column loop, before the work that can throw. Guessing either way gives you a bug — double release, or a leaked Arrow buffer tree when a zero-column schema means the loop never runs. Documented inLESSONS.mdand the book, and made impossible by the by-value signature.The pitfall count in the README, the crate docs and the FAQ was stale at 17, and the book's catalogue was missing L8 and L9. All four now agree with
LESSONS.md, and an integration test fails when a documented count differs from it (see Added). -
A copy function may now implement
COPY … FROMalone.CopyFunctionBuilder::registerused to requirebind,sinkandfinalize, which made a read-only format impossible even thoughduckdb_register_copy_functionaccepts one: it decides what a copy function supports frominfo.sink != nullptrandcopy_from_bind != nullptrindependently. Leaving all three unset is now valid whencopy_fromis set; setting only some of them is an error that says so. ExistingCOPY … TOfunctions are unaffected. -
CI gains four jobs:
miri(546 unit tests under the interpreter),leak-check(LeakSanitizer over the end-to-end suite against a reallibduckdb, now leak-clean),fuzz(cargo-fuzzover the description.yml parser, theduckdb_string_tdecoder and the validators) andsemver(cargo-semver-checks).tests/ffi_roundtrip.rsis now linted — it is feature-gated, so the plain clippy job had been compiling it away to nothing. -
extension-loadis now a matrix overDuckDBv1.4.4, v1.5.0, v1.5.5 andlatest, rather than onereleases/latestrun that silently retargeted wheneverDuckDBshipped. That is the README's compatibility claim, proven rather than sampled. -
New end-to-end coverage for nested types as scalar-function input —
LISToffsets that are cumulative rather than uniform, NULL elements inside a list, aMAPkey miss, and anARRAY's fixed stride across rows. Writing them was already covered; reading them does raw offset arithmetic against a layout onlyDuckDBdefines, which a mock cannot check. -
AUDIT.mdrecords the full review: what was read, what was probed, what was verified correct, and what is still open. -
CI: the AddressSanitizer job is now blocking, as planned when it was added. It passed on
mainand over the merged 125-test end-to-end suite with no reports and no suppressions. -
CI: the informational beta clippy job now also lints the end-to-end tests, with
bundled-test-prebuilt,duckdb-1-5-4against a pre-built libduckdb. Those tests compile only withbundled-test-prebuilt, so beta's newassert_is_emptylint fired on four of them while the job stayed green. -
Mutation testing: the configuration moved to
.cargo/mutants.toml. At the repository root cargo-mutants never read it — so the full sweep ran without its exclusions or features (2,164 mutants listed instead of 1,348) — and it carried two keys cargo-mutants 27.1.0 rejects (cap_timeout,jobs). Itsexamine_globsis gone too: once the file was read, that key overrode the incremental job's--fileflags instead of being narrowed by them. -
Breaking:
TableFunctionBuilder::with_staterequiresS: Clone + Send: the statebindreturns is a template and every execution scans a fresh clone. Use the newwith_bind_initfor state that cannot be cloned. -
Breaking:
TypedTableFunctionBuilder::projection_pushdownis removed — the typed scan closure cannot learn the projection, so enabling it returned the wrong columns. Use the rawTableFunctionBuilderfor pushdown. -
Breaking:
FfiBindData::set/FfiInitData::setrequireT: Send + Sync,FfiLocalInitData::setT: Send;ReplacementScanBuilder::register_with_dataandConnection::register_replacement_scan_with_datarequireT: Send + Sync. -
Breaking: every scalar
Valuegetter returnsOption<T>(as_i8…as_u128,as_f32,as_f64,as_bool, the date/time/timestamp family,as_interval,as_uuid,as_decimal; the newas_enum_indexdoes too):Nonefor a null handle, SQLNULL, a non-scalar value or a failed cast. A failed cast used to return a sentinel (T::MIN,NaN) indistinguishable from a real value. Theas_*_or(default)forms keep their signatures and now cover all four cases. -
Breaking:
Value::as_blobaccepts only aBLOB(DuckDB's cast of anything else toBLOBcould throw) and errors on SQLNULL;Value::as_strerrors on SQLNULL. -
Breaking:
FileSystem<'ctx>borrows itsClientContext. Code that stores aFileSystemneeds a lifetime parameter, and theClientContextmust outlive it. -
Breaking:
DuckDbErrorTypegainsAutoload,SequenceandInvalidConfiguration(40–42), which used to map toInvalid. -
Breaking:
SelectionVector::newreturnsResult<Self, ExtensionError>. -
Breaking:
datetime::date_to_days,time_from_micros,time_tz_from_bits,timestamp_from_micros,timestamp_to_micros,time_tz_bitsanddecimal_to_f64returnOption. -
Breaking:
VectorWriter::set_null/set_null_range(and soDataChunk::propagate_nulls) on aSTRUCTorARRAYvector also null the row's fields / elements, recursively, as DuckDB'sFlatVector::SetNulldoes. To reuse such a row, callset_validon it first, which restores whatset_nullnulled below it, then write any field or element NULLs. -
Breaking:
SqlMacro::to_sqlemits double-quoted identifiers (CREATE OR REPLACE MACRO "add"("a", "b") AS (a + b)). Calling the macro is unchanged: DuckDB resolves quoted identifiers case-insensitively. -
ScalarFunctionBuilder::varargs,varargs_logicalandvolatileno longer requireduckdb-1-5: both C functions are in the stable v1.2.0 API. -
Appenderis re-exported from the prelude without a feature, matching the module, anduse quack_rs::prelude::*now bringsentry_point!/entry_point_v2!into scope, as the prelude's own documentation said. -
AggregateFunctionSetBuilder::overloadsis unchanged, but the builder its closure receives is now namedAggregateOverloadBuilder, for symmetry withScalarOverloadBuilder. The old name remains as a deprecated type alias (quack_rs::aggregate::builder::OverloadBuilder) and still compiles. -
AggregateOverloadBuilderis exported fromquack_rs::aggregateand from the prelude. The oldOverloadBuilderwas reachable only atquack_rs::aggregate::builder::, and had no public constructor, so a caller could not build one outside anoverloadsclosure. -
AggregateOverloadBuildermoved out ofset.rstosrc/aggregate/builder/overload.rs. -
Breaking:
QueryResult::next_chunkreturnsResult<Option<OwnedDataChunk>, ExtensionError>.duckdb_fetch_chunkreturns null both at the end of the rows and when the fetch fails, andnext_chunkreturnedNonefor both, so a streaming result that failed part way — a runtime error in the query, or another statement run on the same connection — read as a complete, shorter result. The error DuckDB recorded is now returned asErr, and keeps being returned on later calls; a clean end staysOk(None). -
Breaking:
Value::time_nsandValue::timestampreturnResult<Value, ExtensionError>and refuse a payload outside the range DuckDB's own SQL produces; so do the newValue::time,time_tz,timestamp_tz,timestamp_s,timestamp_msandtimestamp_ns(see Added). DuckDB stores any 64-bit payload unchecked, and rendering an out-of-range one crashed or aborted the process (Value::time_ns(i64::MAX),Value::timestamp(i64::MIN)) or printed garbage.Value::dateand the newValue::intervalare infallible: DuckDB renders every value of those types. -
Breaking:
CatalogEntry::lookupandCatalog::get_entryreturnResult<Option<_>, ExtensionError>:Ok(None)is "not found",Erris "refused". ATypeorCollationlookup of a name that makes DuckDB autoload an extension (inet,json, an ICU collation name) is refused without calling DuckDB while theautoload_known_extensionssetting is on: a failed autoload insideduckdb_catalog_get_entryaborted the process, and a successful one loaded an extension as a side effect of a lookup. The entry types that were already refused (Schema,Database,PreparedStatement,Invalid) are now anErrtoo, instead of aNonethat looked like "not found". -
Breaking:
FileSystem::openreturnsFileHandle<'_>, which borrows theFileSystem. AFileHandlekept after its database was closed read freed memory (valgrind: an invalid read induckdb::FileHandle::Read).FileHandle::from_rawreturns a handle whose lifetime the caller chooses. -
Breaking:
InitInfo::projected_column_indexreturnsOption<usize>,Nonepast the end of the projection, where DuckDB returns 0 — a real column index.CopyBindInfo::column_typereturnsOption<LogicalType>, bounds-checked likeBindInfo::result_column_type; it used to wrap the null handle DuckDB returns for an out-of-range index. -
Breaking:
TypedTableFunctionBuilder::buildreturns an error whenprojection_pushdown(true)was set on theTableFunctionBuilderbeforewith_state/with_bind_init. That sequence bypassed the typed builder's no-pushdown rule, and the scan returned the wrong columns (SELECT bgot columna's values). -
Breaking:
ConfigOptionBuilder::registerrefuses an option with no default value (DuckDB then reports it as an unrecognized configuration parameter), an option of a pseudo-type (ANY,SQLNULL,INTEGER_LITERAL,STRING_LITERAL) or of a type the typedValueconstructors cannot build (BIGNUM,GEOMETRY,VARIANT, …), and a name thatduckdb_settings()already lists, compared case-insensitively and including aliases: an option namedthreadsused to register and shadow the built-in setting. The default is now handed to DuckDB already typed (see Security). -
Breaking:
TableFunctionBuilder::registerrefuses a name that already belongs to a table function or table macro, andCopyFunctionBuilder::registera format name that already exists (csv,parquet, …). DuckDB dropped such a registration while reporting success, so the function was never called. The C API cannot overload table functions. -
Breaking:
ScalarFunctionBuilder::registerandScalarFunctionSetBuilder::register(and somap1,map2and the other typed constructors) refuse a parameter signature thatduckdb_functions()already lists under that name, built-ins included. DuckDB merges a scalar registration into the existing entry with override on, so the new overload silently replaced the old one for every query in the database — after registeringabs(BIGINT),abs(-5::BIGINT)returned 995 — or, with a different return type, made every call ambiguous. Signatures containingSTRUCT,UNION,ENUMor a literal pseudo-type are registered unchecked, as documented on both builders. -
Breaking:
SqlMacro::registerrefuses a body that holds more than one statement, counted with DuckDB's own parser, and executes nothing. The body1); DROP TABLE t; SELECT (1used to register and drop the table. A scalar body that ends in a--comment, which used to comment out the closing parenthesis, now works. -
Breaking:
Appenderno longer loses rows silently. DuckDB's appender cannot take back a value, so arow()closure that failed after appending part of a row left the row half-written, andclose()then returnedOkand wrote none of the buffered rows. Such arow()now poisons the appender: every laterappend_*,row,end_row,flushandclosereturns an error saying how many buffered rows were not written.close()with a row started but not ended is an error; every mutating method is an error after a successfulclose()(DuckDB accepted appends after close and wrote them at the next flush);append_chunkin the middle of a row is refused. Withduckdb-1-5,clear()resets DuckDB's state and clears the poison. A value rejected first in its row loses nothing and does not poison. -
Breaking: the
LogicalTypeSTRUCT and UNION constructors (struct_type,struct_type_from_logical,union_type,union_type_from_logicaland theirtry_forms) apply the rules DuckDB's binder applies to the same types in SQL: names unique ignoring ASCII case (empty names exempt) and at mostMAX_UNION_MEMBERSunion members. The C API checks nothing, so a scalar returning such a type registered and then failed every call with "duplicate name in struct". Thetry_forms return an error; the others panic. -
Breaking:
VectorWriter::write_varchar/write_blobandStructWriter::write_varchar/write_blobpanic for a value longer thanMAX_STRING_LENinstead of storing a truncated one (see Fixed). Insidescalar_callback!and the typed scalar constructors the panic becomes a SQL error;try_write_varchar/try_write_blobreturn it instead. -
Breaking:
MockVectorWriterbehaves like a real output vector, so a test that passed for a callback that is wrong in DuckDB now fails.set_nullfollowed by awrite_*leaves the row NULL (set_validundoes a NULL); a row never written is valid, not NULL (is_writtentells a test whether the loop wrote it); a write past capacity panics instead of growing the mock; a string overMAX_STRING_LENpanics, asVectorWriternow does. The docs no longer claim one function can be called with both the mock and the real writer: the types differ. -
Breaking:
generate_scaffoldvalidates the free text it writes into the generated files:descriptionandmaintainermust be non-empty, must not start or end with whitespace and must not contain control or bidirectional characters (maintainermust also be one line),github_repomust beowner/repo, andgit_refa commit hash or tag. Those fields went intodescription.ymlunquoted, soFast: analytics,Analytics #1or a newline produced a file that parsed to something else or not at all; they are now YAML double-quoted scalars. Each line of a multi-line description gets its own//!prefix; the second and later lines used to fall outside the doc comment and fail to compile. -
Breaking:
validate_spdx_licenserejects nesting deeper than 64 parentheses. It recursed once per(with no limit, so a license field of a million(overflowed the stack and aborted the process, reachable fromparse_description_ymlon untrusted input. -
TypeId::TimeNs,Any,Varint,SqlNull,IntegerLiteralandStringLiteralno longer requireduckdb-1-5. All six exist in DuckDB 1.4.4, this crate's floor, and 1.4.4 producesTIME_NSandBIGNUMcolumns, so with default featuresLogicalType::get_type_idpanicked on those columns. -
Breaking:
TypeId::Varint.sql_name()returns"BIGNUM", DuckDB's name for the type since 1.4, instead of"VARINT". Both names parse as SQL, but code that compares the returned string sees a different value. -
SecretEntry'sDebugoutput shows[REDACTED]for a non-empty scope: the scope (a bucket or URL prefix) is zeroized on drop as sensitive, butDebugprinted it. -
InMemoryDbchecks, before first use, that thebundled-test-prebuiltC++ shim was compiled against headers whoseduckdb_ext_api_v1has the size thelibduckdb-sysbindings expect, and panics naming both slot counts,DUCKDB_LIB_DIRandcargo tree -i libduckdb-sysif not. With a DuckDB 1.5.0 library under 1.10505 bindings, a slot the older headers lack was left as stack garbage and the tests ran on; larger headers would overrun the buffer. -
scripts/check-abi-table.pylists every upstreamvX.Y.Ztag from v1.2.0 on instead of a hard-coded list, which is how v1.4.5 went missing from the layout table (see Security). -
RELEASING.md: a release is done only once the tag is onorigin, crates.io's newest version is theCargo.tomlversion and docs.rs has built it, each with a command to check; version snippets are bumped in the release pull request itself. -
Internal layout, with no change to any public path:
src/value.rsis split intovalue/composite.rs,value/nested.rs,value/scalars.rs,value/temporal.rsandvalue/temporal_checks.rs; theLogicalTypeconstructors moved totypes/logical_type/construct.rs; andScalarOverloadBuildermoved toscalar/builder/overload.rs.src/query.rs,src/appender.rs,src/arrow.rsandsrc/testing/mock_vector.rsare split the same way into private submodules. -
Every
unsafeblock in library code now states, in a// SAFETY:comment, the invariant it relies on and why it holds; 149 did not.clippy::undocumented_unsafe_blocksis enabled so CI keeps it that way (test code is exempt). -
docs/architecture.mdmatches the crate again: its module table listed neitherabi,arrow,callback,chunk_writer,datetime,query,secrets,tlsnorwarning, and saidappender,table_description,ScalarFunctionBuilder::varargsandvolatileneedduckdb-1-5(they do not). An integration test now fails when the table andsrc/lib.rsdisagree.
Fixed
Fifth audit
- On a 32-bit target the Arrow layout walk's refusal of an oversized row
count named
vector::ops::MAX_CAPACITY(2^28 - 1) as the limit while refusing exactly that many rows: a list child's reserve rounds up to a power of two, so the limit is 2^27. The message now states the effective limit. Found by the first run of the unit tests on wasm32. - Rendering a value aborted the process for values ordinary SQL builds:
a
VARIANTholding an out-of-range timestamp, aDECIMAL(38, 0)holdingi128::MIN(fromsumover two in-range values), and aGEOMETRYbuilt from malformed WKB each madeDuckDB's cast to text throw through the C API; aDECIMAL(38, 38)holding 1.2 rendered with an unwritten first byte, and aborted too when that byte was not valid UTF-8. The render guard is now an allow-list, andDECIMALpayloads are checked against their width. ListBuilderaborted the process on a large list of a wide type.DuckDB's ceiling is 2^37 bytes per child buffer, not elements, checked after it rounds the reservation up to a power of two, so aBIGINTlist of 2^34 + 1 elements, or anINTEGER[1000]list of 2^25 + 1, reached a reserve that throws through the C API. The builder respects the rounded byte ceiling; a row past it is NULL.- Aggregate states
DuckDBmoved were never dropped (unreleased fourth-audit code; 0.16.0 had no tag). That pass'sFfiStatetag was derived from the slot's address, and radix repartitioning copies states to new rows, so every moved state was skipped and itsTleaked (8 of them after one ungrouped query on eight threads in the regression test). The tag follows the slot's contents. FfiState's type tag was salted with the address ofT's type-name string (unreleased fifth-audit code), which is not unique: with more than one codegen unit and no fat LTO (Cargo's default release profile) the init, access and destroy callbacks could see different addresses, sowith_statefound no state anddestroyskipped every one: in such a build everyFfiStateaggregate returned NULL or garbage (10 of the 279 end-to-end tests fail that way). The salt is now a hash ofTypeId::of::<T>().- Aggregate states
DuckDBnever destroys leaked a box each. A grouped aggregate's states that a stopped scan never reached (aLIMITabove it, an error, an interrupt) are never destroyed byDuckDB1.4.4 to 1.5.5; underLIMIT 10over 300,000 groups, 297,952 boxedTs leaked. A smallTis now stored inDuckDB's own state bytes, so it leaks nothing unless it owns heap memory itself. ListBuilder::with_element_limitcalled after the first row could raise the limit pastDuckDB's ceiling for the child type, which the builder applies once, at the first row; a later row past the ceiling then reachedduckdb_list_vector_reserveand aborted the process (aBIGINTrow of 2^34 + 1 intests/ffi_roundtrip/list_limits.rs). The ceiling is kept apart and always applies. On a 32-bit target a row of more than 2^31 elements under a larger limit made the reservation 0 (next_power_of_twooverflowing) while the closure still wrote the row, past the child; the reservation is now the limit there.TableDescription::column_name,column_typeandcolumn_has_defaultaborted the process for indexu64::MAXonDuckDB1.5.0 to 1.5.5: the C API converts the index to anoptional_idx, whose constructor throws for that value outside anytry(upstream item 30). They returnNonefor it without callingDuckDB, as for any other index past the last column.data_chunk_to_arrowexported different values without an error. AnINTERVALof more thani64::MAX / 1000microseconds wrapped whenDuckDBconverted it to nanoseconds; aUHUGEINTof 2^127 or more came out negative; aHUGEINTorUHUGEINTof 39 digits was exported as adecimal128(38, 0)it does not fit. Each is now refused, at any nesting depth; aHUGEINTis accepted whenarrow_lossless_conversionexports it as a 16-byte binary.data_chunk_to_arrowcould return an array that contradicts its schema. Before 1.5.5,BIGNUM(and from 1.5.0GEOMETRY) exported underarrow_output_version = '1.4'are written as binary views while the schema declares plain binary; a consumer reads the views as offsets, and appendingDuckDB's own re-import of the batch crashed on 1.4.4 to 1.5.4 (item 34). The export is now checked against the declared schema and refused on a mismatch.data_chunk_from_arrowimported valid Arrow arrays from the wrong rows, or read and wrote out of bounds.DuckDBmishandles offsets below the top level: a struct inside an offset struct or list, a union's members, a run-end-encoded array's value validity, and a dictionary's validity under a list. It also mishandles overlapping or gapped list views, dictionaries whose values are dictionary-encoded, sparse unions whose type codes are not0, 1, …, and a run-end-encoded array where it reads a plain one. A dictionary-encoded array of more than 2048 rows under a struct with NULL rows had its validity copied past a 2048-row heap mask; the regression test aborted with glibc'scorrupted size vs. prev_size. Each layout is refused, naming the node, and the neighbouring layoutsDuckDBdoes import correctly are still accepted (tests/ffi_roundtrip/arrow_layout.rs).data_chunk_from_arrowaccepted three more layoutsDuckDBimports wrongly, found by the fifth audit's review of the new layout walker: a fixed-size list with NULLs whose child is aSTRUCTwith a dictionary field of more than 2048 rows (the list's NULLs are broadcast into the struct and reach the 2048-row mask of item 25; valgrind reports 156 invalid accesses inSetInvalidpast the 256-byte mask on 1.5.5); a dictionary withnull_count = -1, whose NULL rows came back as values (item 31); and a sparse union with a nonzeronull_count, whose type ids were read as validity, so every row came back NULL (item 32). Each is refused, as is such a layout below a nodeDuckDBconverts as zero rows (the walk used to stop there, thoughDuckDBstill expands a run-end-encoded descendant; valgrind showed its values' validity read 11 bytes past a 1-byte bitmap on 1.5.5, item 24). So is ageoarrow.wkbcolumn read as more than 2048 rows:DuckDB1.5 copies its storage into a 2048-row vector before the cast toGEOMETRY(SIGSEGV on 4096 rows, item 33).- A null-typed field below the top level of an imported Arrow column read
as valid after its first row:
DuckDBimports it as a constant vector, which the readers index as flat. The column is now flattened whenever its type holdsNULLat any depth, not only at the top. - The
CombineFndocumentation advised acombinethat gave wrong window results. It said to move out of the source states. A window's segment tree combines one state into every frame that covers it, so acombinethat consumed its source gave 4985 of 5000 rows wrong overROWS BETWEEN 100 PRECEDING AND CURRENT ROW(pinned bytests/ffi_roundtrip/agg_window.rs). It now says to leave the source unchanged. - A typed table function answered with the wrong column under projection
pushdown switched on after
build().buildrefused pushdown switched on beforewith_state, but not on the raw builder it returns, andSELECT bthen returned columna's value. Registering such a builder (andMockRegistrar::register_table) now fails. - A typed table function's callbacks could be replaced after
build()with the raw builder's safebind,init,local_initandscansetters (orextra_info), after which the typed trampolines read one type's data as another's: an init callback setting au64as init data made the scan lock it as aMutex<S>. Registering such a builder (andMockRegistrar::register_table) now fails; bothextra_infoSafety sections now say the pointee must be what the installed callbacks read. - A
rowclosure that panicked after its first value left the appender unpoisoned, so finishing the row by hand committed a row half written by the closure that panicked. It is poisoned, as an error there poisons it. - Dropping a
FileHandlecould abort the process when the close threw:duckdb_destroy_file_handlecallsClose()with notry. Only an extension's file system can throw there;DuckDB's local close cannot. The drop now closes throughduckdb_file_handle_close, which catches, and destroys the handle only if that succeeded; after a failed close it leaks the handle rather than retry the close unguarded. FileHandle::seekclamped a position pasti64::MAXtoi64::MAX, which a file system that accepts that offset (tmpfs) took, so the call returnedOkat a position the caller never asked for. It is now anInvalidInputerror.append_metadatanamed the wrong default platform on OpenHarmony (*-linux-ohos):DuckDBappends_muslthere, as for any musl-based Linux, and the tool did not.MockRegistrarstill accepted builders whose types the real registration refuses: a composite or literalTypeIdin any parameter, varargs, return, named-parameter, cast or config-option slot, and anANYreturn type. Its module doc said these checks needDuckDB; they do not. Each builder now runs one sequence of type checks, withDuckDBwhen registering and without it in the mock, so the messages are the same.- Four writer contracts named only a
LIST/MAPvector's direct child as moved by areserveon thatLIST/MAP.VectorWriter::from_vector,StructWriter::new,StructVector::field_writerandValidityBitmap::ensure_writablenow say the same about every STRUCT field and ARRAY element vector below that child, down to the nextLISTorMAP:DuckDB1.5.5 reallocates all of their data and validity buffers (tests/ffi_roundtrip/nested_reserve.rsmeasures which move). - Arrow import: a fixed-size list format DuckDB accepts could bypass the
layout checks.
DuckDBreads the size in+w:Nwithstd::stoi, so+w:2x,+w: 2and+w:+2are fixed-size lists to it; quack-rs parsed the size strictly, took them for leaves and skipped every check below them. A fixed-size list → struct → dictionary layout the checks refuse under+w:2crashed the process (SIGSEGV) under+w:2x. The size is now parsed asstoiparses it. - Arrow import: offsets below the top level were not checked for being
negative, and their sums could overflow (a panic in a debug build).
Every node's length and offset must be non-negative, and a sum past
i64::MAXis refused. ArrowArray::releaseandArrowSchema::releasecalled a producer's callback again at drop when the callback did not null itself, as the Arrow specification requires it to; they now null it themselves.entry_point!andentry_point_v2!aborted the process when an argument expression panicked.$policyand$registerwere evaluated in the generatedextern "C"function before the panic guard ("panic in a function that cannot unwind"); they are now evaluated under it, and the load fails with the panic's message.- A typed table function used as a
COPY … FROMreader could invalidate the database. Its bind declares columns, and underCOPY … FROMDuckDBappends each to theINSERT's own expected types, so every chunk reaching the table is too wide: aDuckDBbuilt with assertions failschunk.ColumnCount() == types.size()and invalidates the database; a release build drops the column. The typed bind now fails with a message when a column is declared there (upstream item 37). VARCHARandBLOBvalues were decoded as little-endian.duckdb_string_tholds its length and pointer in the target's own byte order;DuckStringView,read_duck_stringandread_duck_blobread both as little-endian, so on a big-endian target an inline string's length read aslength << 24and its inlined bytes were followed as a pointer (reproduced under Miri with--target s390x-unknown-linux-gnu). Both are now read natively, and the pointer is read as a pointer at the target's width, which also keeps its provenance. NoDuckDBextension platform is big-endian, and nothing changes on a little-endian one; bytes passed toDuckStringView::inline_from_bytesare now read in native order too.- The Arrow export check read
HUGEINT/UHUGEINTrows with their halves swapped on a big-endian target.DuckDBstores them as{lower, upper}; the check read the 16 bytes as a nativei128, so on s390x it passed10^38(read as about1.27 * 10^37) into a lossy export and refused2^63. It now readsDuckDB's struct (reproduced under Miri on s390x). - On a 32-bit target (wasm32),
ListBuilderandOwnedVector::newcould askDuckDBfor a buffer whose size wraps.DuckDBcomputes a buffer's size as a 64-bitidx_tand passes it tomalloc, which narrows it to a 32-bitsize_tunchecked. The limits bounded the element count byusize::MAX, not the bytes, so aBIGINTorVARCHARlist child could reserve 2^32 elements andmallocreceive 0 bytes, andOwnedVector::new(HUGEINT, 2^28)succeeded with the same wrap. The list limit now also fits one allocation, andvector::ops::MAX_CAPACITYis 2^28 - 1 on a 32-bit target. 64-bit targets are unchanged. data_chunk_from_arrowpassed on lengths no vector can hold.DuckDBsizes the chunk from the array's length before its error handling starts, and sizes each list child it reserves from the list's element count; a run-end-encoded column declares any length with a few bytes of buffers. On a 32-bit target a length of 2^28 (VARCHAR) or 2^29 (BIGINT) had its byte size narrowed bymalloc, and the import then wrote past the buffer; on any target a length past 2^37 threw through the C API. The layout walk now refuses, at any node, a row count whose next power of two exceedsvector::ops::MAX_CAPACITY(the rounding a list child's reserve adds).- A run-end-encoded child of a zero-row list crashed the Arrow import
when the list's offset was not 0.
DuckDBtreats a list it converts as zero rows as empty whatever its offsets say, and reads an empty list's child as a plain array (upstream item 24): a run-end-encoded child there is read from buffers it does not have (SIGSEGV on every release from 1.4.4). The layout walk refused this only when the list's offset was 0, and a valid array (an inner list under an empty outer row, whose offsets need not start at 0) got through. It now decides asDuckDBdoes, by the row count. - A dictionary-encoded child of a fixed-size list crashed the Arrow import
when the fixed-size list was a
MAPvalue.DuckDBverifies a map by flattening its entries, and flattening anARRAYflattens its child over the child vector's capacity, which a map allocates at the chunk's row count rather than the entries it holds; the dictionary child's selection vector was built only for the entries converted, so the flatten read it out of bounds (SIGSEGVon 1.5.0-1.5.2, an AddressSanitizerheap-buffer-overflowon 1.5.5; upstream item 24). The layout walk refuses a dictionary-encoded fixed-size-list child under a map; the same shape at the top level or under a plain list, which is not over-flattened, still imports.
Fourth audit
- A C API aggregate in a running window returned the wrong answer. Without
a destructor
DuckDBstreams it and re-reads the first row (sum-like:1 2 3 4 5for1 3 6 10 15). Every aggregate now registers one (a no-op when none is given). - Arrow import with a nonzero parent offset imported the wrong rows
(
DuckDBignores the offset for values); it is refused. ListBuilderoverwrote rows already in the list vector when it started on a non-empty one; it appends after them.VectorWriter::set_validleft a nested row's children NULL; it restores them when the row was NULL.- Scalar collision check. It missed signatures containing a type alias,
could be shadowed by a user macro named like the catalog functions it
queries (it now qualifies them with
system.main.), and accepted overloads thatDuckDB's binder finds ambiguous with varargs. Breaking: such overlapping overloads are refused at registration.Connectionkeeps a snapshot of existing scalars, so 300 registrations throughRegistrartake 18.6–35.6 ms instead of 5.2–6.3 s (release build, three runs each). - Keyword names. Breaking:
validate_function_namerefuses the 53 keywords (DUCKDB_UNCALLABLE_KEYWORDS) thatDuckDBcannot call unquoted —coalesce(x)silently ran the built-in — andSqlMacroparameters use the newvalidate_parameter_name(79 keywords). - Catalog entries outlived their handle. Breaking:
CatalogEntrycopies the name and type at lookup and owns nothing afterwards. - Secret scopes were parsed from a string; they are read as a list,
NULL-safely, from
system.main.duckdb_secrets(). CopyGlobalInitInfo::get_file_pathtruncated at a NUL. Breaking: it returnsResult<String>;get_file_path_bytesreturns the raw bytes.interval_to_microsreported overflow for totals that fit (an intermediate sum overflowed); it computes the exact total ini128.Appender: a failed automatic flush (every 204,800 rows) poisoned the appender with a false "half-written row" message; the row counts as ended and the constraint error is reported as it is.- On DuckDB 1.4.x, registering a scalar under any existing name (a new
overload of
abs, say) failed with no reason, because the C API registers withCREATEbefore 1.5.0; quack-rs refuses it first and says why. AndLogicalType::try_new(TypeId::TimeNs)returned anINVALIDtype there (the 1.4.x C API does not knowTIME_NS); it is an error now, as is any type id the running engine hands back changed. LogicalType::registerblamed a taken name when the type containedANY; it names the cause.try_decimalvalidates width and scale itself (DuckDBdoes only from 1.5.4). Breaking:try_array(_, 0), an emptyunion_typeand an emptyValue::array_valueare refused, as in SQL.MockRegistraraccepted builders the real registration refuses. Breaking: it runs the same checks (missing callback or return type, empty function set, incomplete copy function, config option without type or default) and records nothing on failure.description.yml: a quoted, flow or block value on the line after its key kept its quotes; values thatPyYAML(YAML 1.1) reads as a boolean, number, date or null draw a warning; the scaffold quotesname,githubandref. The "text after a comment" error named the line the value started on rather than the line of the comment that ended it.- Scaffold: the generated
lib.rsfailedcargo fmt --checkfor names of 4 characters or fewer or 40 or more; the generated CI's Linux SQLLogicTest step was skipped by extension-ci-tools. append_metadatarefuses a platform group (linux) and awasm_*platform without--wasm, and recognises a footer with an empty ABI field.validate_semverrefuses numeric pre-release identifiers with leading zeros;validate_spdx_licensenames the canonical spelling for a case-only difference.- ABI refusal for a development engine no longer suggests a declaration
the build ignores;
build.rswarns about a malformedQUACK_RS_TARGET_DUCKDB_VERSION.
Earlier passes (0.17.0, second and third audits)
-
Pitfall L8 —
DEFAULT_NULL_HANDLINGdoes not propagate NULLs for scalar functions. quack-rs documented thatDuckDB"automatically returns NULL if any argument is NULL, without your function callback being called". For a scalar function registered through the C API that is false at run time:CAPIScalarFunctioncalls the callback for every row including NULL ones and never inspects the result's validity, and the only NULL check inExpressionExecutor::ExecuteisVerifyNullHandling, whose entire body is inside#ifdef DEBUG. A callback that ignores validity therefore returns a non-NULL answer for a NULL input, silently, in every release build.SELECT f(NULL)still returns NULL — a literal NULL is constant-folded before the function is reached — which is why the bug survives review. From a column it does not. NewDataChunk::propagate_nulls/any_nullrestore SQL semantics in one line; the new typed constructors (see Added) get it right by construction; the docs, the book chapter andLESSONS.mdnow state whatDuckDBdoes, with the source quoted. A regression test pins the behaviour. Aggregates are no different: theirupdatereceives NULL rows under either setting too (Pitfall L12; see Added and the aggregate entry below). -
Composite
TypeIds silently produced an invalid type.duckdb_create_logical_type"returns an invalid logical type" forDECIMAL,ENUM,LIST,STRUCT,MAP,ARRAYandUNION— a non-null handle wrappingLogicalTypeId::INVALID, so the existing null check never fired..param(TypeId::Struct)failed much later with a message that named neither the parameter nor the fix, andget_type_id()on one panicked. NewTypeId::is_composite/composite_constructor_hint;LogicalType::newasserts,try_newerrors, and every builder validates before allocating anyDuckDBhandle. -
extra_infoleaked when a builder was not registered.DuckDBonly takes ownership atduckdb_*_set_extra_info; a dropped builder dropped the pointer. This reached users through APIs that never mention a pointer —TableFunctionBuilder::with_stateboxes two closures. Found by Miri. -
Two stale-borrow bugs in
src/secrets.rs's own tests, which took a pointer into aString, called a&mutmethod, then read through the stale pointer. The library'szeroize_stringwas correct throughout. -
The reference example disabled every panic guard in the crate.
examples/hello-ext/Cargo.tomlshippedpanic = "abort"— the settingvalidate_release_profilerejects outright and the scaffold refuses to generate, because it makes everycatch_unwindin quack-rs inert. CI built that example, loaded it into a realDuckDB, and held it up as the way to do this. Two new tests hold the example and the scaffold's generated profile to quack-rs's own validator, so the generator and the validator cannot drift. -
The
ScaffoldConfigexample in the README and two book pages did not compile. The struct gained three fields and the exhaustive literals were never updated; rustdoc examples are compiled bycargo testbut Markdown code fences are not, so the copy a new user reaches for was the broken one. All now use..ScaffoldConfig::default(). -
Registration failures now name what to check.
duckdb_register_*_functionreports failure as a bareDuckDBErrorwith no message. There are exactly three causes, and a name collision with aDuckDBbuilt-in (list_sum,array_sum, …) looks identical to a type error. The message names all three and points atSELECT * FROM duckdb_functions() WHERE function_name = '<name>'. -
Wrong answers, no error:
- A NULL row of a
STRUCTresult kept its fields valid, so(f(x)).areturned the stale field value instead of NULL. - A typed table function failed the second time a plan ran
(
PREPARE … ; EXECUTE p; EXECUTE p;, or a recursive CTE):initmoved the state out of bind data DuckDB reuses for every execution. Valuegetters cast the value in place:Value::double(1.5).as_i32()turned the value intoDOUBLE 2.0. Getters now read a private copy.BindInfo::add_result_columnwith a type containingANY/INVALIDwas dropped by DuckDB, shifting every later column; it is now a bind error.datetime::time_tz_bitssilently corrupted an out-of-range offset.
- A NULL row of a
-
Duplicate overload signatures in a scalar or aggregate function set are rejected at
register, naming both overloads, instead of registering and then failing every call with "Could not choose a best candidate function". -
Expression::foldreturnsErrfor a non-foldable expression instead ofOkwith a null-handleValue. -
CastFunctionBuilder::registerleakedextra_infowhen DuckDB rejected anANY/INVALIDtype; those are now rejected before anything is handed over. -
ReplacementScanInfo::set_error("")was ignored by DuckDB, so the query fell through to "table does not exist". -
SqlMacrowith a SQL keyword as a name or parameter produced a parser error. -
Documentation that was false against the code or DuckDB:
set_max_threads(it does not needlocal_init);DbConfig::setaccepts unknown option names; Pitfall L4 (a skippedensure_validity_writablesilently drops the NULL rather than segfaulting); the book'spanic = "abort"advice (it must be"unwind", which the crate itself enforces); several README and book examples that did not compile; and the stable ABI prefix, which is ABI-identical since v1.2.0 but not byte-identical (two slots were renamedvarint→bignumin v1.4.0). -
ClientContext's constructors now state that the context must not outlive its connection: DuckDB's wrapper holds a reference, not an owner. -
Pitfall L10 (
LESSONS.md,book/src/reference/pitfalls.md) — scalar bind data is dropped whenDuckDBcopies a bound expression. The book's pitfall summary table was also missing L8 and L9; all three rows are now there. -
ScalarBindInfo::set_bind_data_copy— scalar bind data was silently lost wheneverDuckDBcopied a bound expression.CScalarFunctionBindData::Copy()populates the copy's bind data only if a copy callback is registered, and quack-rs never exposed the setter, soget_bind_datacould return null on a copied expression: a wrong answer, not a crash. -
An invalid
mutants.ymlthat silently disabled the mutation gate. A shell comment added earlier in this branch wrote out an empty workflow expression while explaining not to pass values that way. GitHub evaluates expressions anywhere in the file, including inside shell comments, so the empty one invalidated the whole workflow:Invalid workflow file: .github/workflows/mutants.yml (Line: 203, Col: 14): An expression was expectedThis fails silently by design: GitHub records a run with zero jobs and the workflow stops running.
mutants-incrementaltherefore stopped executing on pull requests while every other check stayed green — the same "a gate that is not actually running" failure this release is otherwise about, introduced by this branch rather than found in it.scripts/check-workflow-expressions.pynow runs in thedocjob and rejects this class. It fails on the exact commit that broke and passes on the fix. Neitheryaml.safe_loadnor a JSON-Schema check catches it, because both treat therun:block as an opaque string. -
22
assert!(x.is_empty())/assert!(!x.is_empty())assertions that beta clippy's newassert_is_empty/assert_is_not_emptylints reject, across 10 files. These are not style noise:clippy-betawas running stable clippy before this release, so it had never reported them, and when beta promotes to stable the blockingclippyjob inherits every one. Each now usesassert_eq!/assert_ne!against an empty value, which is what the lint asks for and prints the actual value on failure. -
Four CI quality gates were testing nothing, each verified against the files rather than inferred:
- The MSRV job (
ci.yml) and the release gate's MSRV entry ran a barecargo checkafter selecting 1.86.0.rust-toolchain.tomlpinschannel = "stable"and a toolchain file overrides the rustup default, so both ran stable. Nowcargo +1.86.0 check. clippy-betaran stable clippy for the same reason.- Miri ran with default features,
cfg-ing out everyduckdb-1-5*module — includingsrc/arrow.rs, the largest block of pure-Rustunsafehere. - Doctests were never compiled: every invocation used
--all-targets, which excludes them. All 183 passed when the gate was added.
- The MSRV job (
-
release.ymlstill usedfail-fast: true, the settingLESSONS.mdblames for a release that shipped with two platforms broken. -
MUTANTS_EXITcapturedtee's status rather than cargo-mutants'. -
SECURITY.mdrecommendedpanic = "abort"; that makescatch_unwindinert and disables the crate's entire panic-containment mechanism.Cargo.tomlhas always setunwind, andvalidate_release_profilerejectsabort. -
RELEASING.mdStep 1 listed 9 check names against 30 CI jobs, and.github/workflows/README.mdlisted 14. A maintainer following either could tag with Miri, LeakSanitizer,osv-scanorsemverred. The workflow README now carries a table generated fromci.yml, with the generator inline. -
An aggregate's
updatereceives NULL rows; the documentation said it did not.NullHandling, thenull_handlingsetter on both aggregate builders, the null-handling book page and the first-extension tutorial said that DuckDB skips NULL rows before an aggregate'supdateunlessSpecialNullHandlingis set. DuckDB'sCAPIAggregateUpdatepasses every row of the chunk toupdate, NULL rows included, underDefaultNullHandlingandSpecialNullHandlingalike; for an aggregate, DuckDB reads the setting only when decorrelating a correlated subquery. Anupdatethat reads a row without checking its validity reads a meaningless value for every NULL row. The docs now say so, the new Pitfall L12 gives the symptom and the fix (skip rows whoseis_validis false), and the end-to-end testaggregate_update_receives_null_rows_under_either_null_handlingpins the behaviour. L12 also records that, under either setting, a count-like aggregate in a correlated subquery returns NULL, not 0, for an outer row with no match; DuckDB rewrites that NULL to 0 only for its owncount. -
Wrong answers, no error, third pass:
- Once a row exceeded
ListBuilder's child-capacity ceiling,push_row/push_map_rowwrote no list entry for it or for any later row of the chunk. DuckDB reuses output vectors, so those rows returned the previous chunk's lists as valid values. A refused row is now NULL. VectorWriter::write_varchar/write_blobstored a value longer thanu32::MAXbytes as a truncated string, forVARCHARpossibly cut inside a UTF-8 sequence: a 4 GiB + 1 byte value came back one byte long.- A panic in a
cast_callback!body underTRY_CASTleft the output vector as it was, soTRY_CASTreturned zeros or a previous query's values rather than NULL: DuckDB ignores a cast's return value inTRYmode and nulls only the rows passed toset_row_error. The macro now marks every row of the chunk as an error. - Two overloads of a scalar set that differ only in their varargs type, such
as
f(BIGINT)andf(BIGINT, BIGINT...), were rejected as duplicates. The varargs type is now part of the signature, as it is in DuckDB.
- Once a row exceeded
-
A
cast_callback!body that returnedfalsewithout callingset_errorfailed a regularCASTwithConversion Error:and no text. The macro now setscallback::CAST_FAILED_WITHOUT_MESSAGEbefore running the body; the body's ownset_errorreplaces it. Hand-written cast callbacks are unchanged. -
set_error("")on a table function, cast or copy function reached the user as an error with no text after the prefix (Binder Error:,Conversion Error:); the newEMPTY_ERROR_PLACEHOLDERis reported instead. -
An error message containing a NUL byte lost everything after it on some paths and kept it on others. Every path now replaces the NUL with
?:set_erroron scalar, aggregate, table, cast, copy and replacement-scan functions,ExtensionError::to_c_string,ErrorData::new, and the entry point's report of a failed registration. -
Expression::folderrors carried DuckDB's exception serialized as JSON ({"exception_type":"Conversion","exception_message":…}) with the typeInvalidInput. They now carry the plain message and the matchingDuckDbErrorType(Conversion,OutOfRange, …). -
With
duckdb-1-5, a typed table function whose bind declares no column fails with an ordinary bind error; DuckDB raised anINTERNAL Errorwith a C++ stack trace. -
FileFlag::CreateNewdid not create a file: it set only DuckDB's exclusive flag, which is ignored without the create flag, so an existing file opened and a missing one failed.set_flag(FileFlag::CreateNew, true)now also setsCreate. On Windows an existing file is still opened rather than refused: DuckDB's Windows file system ignores the exclusive flag, which is now documented onFileFlag::CreateNew. -
WarningCollectordropped every warning, silently, once a panic had poisoned its lock. It now recovers the lock and keeps working. -
ScalarFunctionBuilder::varargs/ScalarOverloadBuilder::varargswith a compositeTypeId(TypeId::List, …) panicked inside the setter.registernow refuses it with an error naming the varargs slot. -
ScalarFunctionSetBuilderandAggregateFunctionSetBuildercheck every overload for a return type and its required callbacks before creating any DuckDB handle, and the error names the overload index; the scalar set used to report "overload missing function callback" with no index. -
AbiPolicy::Strict's refusal of an unknown engine recommendedAbiPolicy::AllowUnknownEngine; it now recommends a build that uses only the stable C API. The layout-mismatch message no longer advises rebuilding aduckdb-1-5extension against a pre-1.5 engine. -
append_metadatarejected a 32-byte footer value, which DuckDB reads in full without a terminating NUL, and it now also accepts--option=value. -
examples/parse_descriptions.rs, pointed at a community-extensions checkout (extensions/<name>/description.yml), found no files and reported "0 parsed, 0 rejected" with exit status 0. It now reads that layout, fails when it finds no files, and exits non-zero when any file is rejected. -
Documentation that was false against the code or DuckDB, third pass:
- Aggregates: the
state_sizecallback runs whenever an operator sizes a state buffer, not once at registration;finalizeruns once per result batch;destroyalso runs on the source states ofcombine;combinetargets are initialised by theinitcallback (StateInitFn;T::default()forFfiState<T>), not zeroed;updatereceives one state pointer per input row, not per group. - Scalars: the
null_handlingsetters repeated the claim Pitfall L8 disproved (that DuckDB skips the callback for NULL arguments). Identical calls share bind data: forSELECT f(i), f(i)the bind callback runs twice, but both columns use the data from its first run, so a bind that reads a counter, clock or RNG needsvolatile. Theregistermethods now state DuckDB's collision rules: an aggregate cannot take a name already used by a scalar function, aggregate or macro, while a scalar merges into an existing function of the same name (an identical signature is now refused, see Changed). The entry-point docs say that registration is not transactional: functions registered before the registration closure fails stay registered. - Table and copy functions:
set_cardinality'sis_exactworks the other way round fromduckdb.h;local_initdoes not enable parallelism, onlyset_max_threadsdoes; tableextra_infomust point toSend + Syncdata;CopyBindInfo::options' shape is now described. - Casts:
set_row_errorrequiresrow < count, because DuckDB checks that bound only in debug builds; the cast docs saidfalsebecomes NULL underTRY_CAST, which holds only for rows passed toset_row_error. - Queries and values: a multi-statement string passed to
query()/execute()runs every statement and returns the first row-producing statement's result; every prepared parameter has a name (a positional?is named by its position);DbConfig::get_flagreturns the extension's name, not a description, for an extension setting;check_valid_utf8agrees withstd::str::from_utf8rather than being stricter. - Dates and intervals:
date_from_daysofinfinitygives 5881580-07-11; DuckDB's 30-day month applies to interval comparison andepoch_us, not to interval arithmetic orepoch;DuckInterval'sEqcompares fields, so1 monthdiffers from30 dayshere although SQL calls them equal. InstanceCache::get_or_createreturns an error for a different config on an already-open database (the docs said the config was ignored), and never caches in-memory paths.FileSystem'sset_flag(flag, false)does not clear a flag.SqlMacrodocuments where the macro is created, that it persists in a database file and can replace a user's macro or shadow a built-in; its parameter errors say "parameter name".- Table macros read a table parameter through
query_table(tbl): theSqlMacroexamples wroteFROM tbl, which DuckDB binds at creation time to a table literally namedtbl, so the module example failed to register. - Loading: the getting-started pages and Pitfall P3 said to
LOADa bare.so, which every supported DuckDB refuses; they now show theappend_metadatastep andduckdb -unsigned, and no recipe setsallow_extensions_metadata_mismatchany more, since a correctly stampedC_STRUCTbuild loads without it. Pitfall P2's symptom was wrong: a bad-dvis stamped without complaint andLOADrefuses the file. - The book's Known Limitations page suggested approximating window semantics
with aggregate functions, the exact shape that crashes every C API
aggregate (Pitfall L11); that page, the README and the hello-ext README now
warn about it. The hello-ext README also stopped telling readers to verify
panic = "abort"and to addduckdbwithbundledas a dev-dependency (Pitfall P9). - The installation page's MSRV rationale,
SECURITY.md's claim thatTlsConfigProviderenforces TLS 1.2+ by default (min_tls_versionis a required method), the entry-point page's macro expansion, the crate's and FAQ's pitfall counts, and the README's validated-fields table (an unlisted SPDX id is a warning, not a rejection). - Book examples that did not compile or did not do what they said: among
them a
DuckStringViewexample using the deprecatedfrom_bytes(it returnsNonefor any string over 12 bytes), amy_bindthat leaked theduckdb_valuefromduckdb_bind_get_parameter, a READMEdescription.ymlexample that did not compile, and scaffold examples that, run as written, overwrote the current package'sCargo.tomlandsrc/lib.rs. Broken links in the book are fixed, and hand-kept test counts, several of them wrong, are removed from the repository trees.
- Aggregates: the
-
# Safetycontracts that allowed undefined behaviour (found while documenting everyunsafeblock; contract text only, no signature or behaviour change):FfiLocalInitData::get/get_mut,FfiInitData::get_mutandFfiBindData::get_from_init/get_from_functiondid not requireTto be the type passed toset;set::<u8>thenget_mut::<[u64; 64]>met every stated clause and wrote 512 bytes through a 1-byte allocation. The init-data getters also did not bound the returned lifetime by the scan call. Both requirements are now stated.MapVector::set_sizedid not require the size to equal the entries written, so a larger size made DuckDB read past the child vectors.read_duck_blobdid not require the vector to outlive the returned slice, which borrows the vector's own buffer for a blob of 12 bytes or fewer.
-
Documented: when an aggregate's
finalizereports an error, DuckDB 1.5.5 does not destroy every state the query created (ungrouped: 2 initialised, 1 destroyed; grouped: 4 and 2), so whatever those states own leaks, and for anFfiState<T>whoseTis boxed (see Changed) the box too. Nothing in an extension can detect it; Known Limitations andAggregateFunctionInfo::set_errornow say so, and an end-to-end test pins it. Two of this release's own tests leaked on these paths and made the LeakSanitizer job fail; both now use states that own nothing. -
cargo docfailed with default features on a link toPreparedStatement::execute_streaming, which needsduckdb-1-5. CI now also builds the docs with default features, the set a dependent crate documents. -
Cargo.toml'sduckdb-1-5description gave the requirement aslibduckdb-sys >= 1.5.0(the crate is versioned 1.10500.0) and claimed the feature has no effect against 1.4.x, which was never checked; the hello-ext README saidvarargsandvolatileneedduckdb-1-5.
Security
Fourth audit
- An Arrow import with a large dictionary wrote past a heap buffer.
data_chunk_from_arrowon a dictionary-encoded array with NULLs and more than 2048 entries — aLISTof 1025+ two-element dictionary-encoded lists is enough — madeDuckDBoverflow a validity mask (valgrind: invalid write inGetValidityMask; SIGSEGV or SIGABRT on 1.4.4, 1.5.0 and 1.5.5). Such arrays are refused beforeDuckDBis called. - An Arrow import read out of bounds. A dictionary-encoded or null-type
column came back as a dictionary or constant vector, which every quack-rs
reader reads as flat: wrong values, then reads past the buffer. Those
columns are copied into flat vectors before
data_chunk_from_arrowreturns. - Rendering a timestamp from SQL aborted the process.
make_timestamp(-9223372036854775808)passed to a table function, thenValue::as_str,display_stringor{:?}:DuckDB's rendering threw through the C API (exit 134, 1.4.4 to 1.5.5), also inside a LIST, STRUCT or MAP. Every temporal payload is checked first;as_strreturnsErr(UNRENDERABLE).as_timeand the other converting getters no longer hand such a payload toDuckDB's cast, which overflowed a signed multiply. - A scalar bind callback inspecting a subquery argument aborted the process
on DuckDB 1.5.0 to 1.5.4. Those releases copy the argument outside any
try;SELECT f((SELECT 1))threw a C++ exception through the callback. Breaking:ScalarBindInfo::argumentnow asks for nothing on those releases and fails the bind with an explanation;get_argument's Safety section states the requirement. - An out-of-range TIME or timestamp crashed
DuckDBlater.PreparedStatement::bind_time/bind_timestamp/bind_timestamp_tzandAppender::append_time/append_timestampaccepted any payload;i64::MINas a TIME segfaulted when rendered, and at the append itself into a VARCHAR column. They are refused (Breaking).bind_str,bind_blobandappend_bytesrefuse more than 4 GiB, whichDuckDBstored modulo 2^32 (a 4 GiB + 3 byte blob became 3 bytes). - A panic in a scalar function's bind or init callback aborted the
process. New
scalar_bind_callback!/scalar_init_callback!macros (duckdb-1-5) catch it and fail the query with its message; a panic with an empty message reportsEMPTY_PANIC_PLACEHOLDER. - A null
get_apioraccessfromDuckDBpanicked or crashed the entry point.init_extensionchecks both beforelibduckdb-sysunwraps them. - A literal type in a registration invalidated the database.
TypeId::StringLiteral/IntegerLiteralas a parameter, return or registered type: the first query that used it raised an internal error and every later query failed.LogicalType::try_newrefuses both. FfiStatedestroyed statesinitnever ran on. After astate_initerror,DuckDBpasses never-initialised states todestroy; each state now carries a tag derived fromT'sTypeId(and, for a boxedT, the box's address) thatdestroychecks (Breaking: see theFfiState<T>entry under Changed).SecretEntryleft secret bytes in spare capacity after truncation; zeroisation now covers the whole allocation.
Earlier passes (0.17.0, second and third audits)
-
A panicking
Dropin extension state aborted the process. Every FFI destructor quack-rs generates —FfiState<T>::destroy_callback,FfiBindData/FfiInitData/FfiLocalInitData::destroy,replacement_scan::drop_box,TypedCallbacks::destroy_extra— dropped aBox<T>of arbitrary user data directly inside anextern "C" fn. Since Rust 1.81 an unwind across that boundary is a guaranteed process abort. Reproduced againstDuckDB1.5.4: an aggregate whose state type has a panickingDropkilled the process withSIGABRTfrom insideduckdb::RowOperations::DestroyStates, on a task-scheduler thread. All of them now run under the newcallback::catch_ffi_panic, which is public so extensions writing their ownextern "C"destructors get the same containment.Where
DuckDBoffers an error channel the panic is now reported instead of swallowed:CAPIAggregateStateInitchecks the error flag and throws, so a panickingDefault::default()becomes an ordinary SQL error rather than a silent NULL. The state destructor has none (CAPIAggregateDestructortakes no info and returns nothing), so there the message is discarded.FfiState::init_callbackalso no longer forms a&mut Selfover the possibly-uninitialised allocationDuckDBhands it. -
Soundness: safe code could corrupt memory, race, or read freed memory.
FileSystemheld a rawClientContextpointer with no lifetime, so it could be used after its connection closed and read freed memory (confirmed under valgrind). Breaking:FileSystem<'ctx>.SelectionVector::newexposed uninitialised memory through the safeas_slice()(stale0xDEADBEEFobserved), and a large length made DuckDB compute a wrapped allocation size, so safe indexing segfaulted. Breaking: it returnsResult, rejects lengths aboveMAX_LENbefore DuckDB is called, and zeroes the buffer.datetime::decimal_to_f64read past DuckDB's powers-of-ten tables forscale > 38.FfiBindData,FfiInitDataand replacement-scan data are shared across threads by DuckDB. Breaking:FfiBindData::set/FfiInitData::set/register_with_datarequireT: Send + Sync,FfiLocalInitData::setrequiresT: Send.
-
Process aborts from ordinary input. Each of these let a DuckDB C++ exception unwind into Rust ("Rust cannot catch foreign exceptions"), killing the host process; each is now validated in Rust first and reported as an error or
None:- every
Value::as_*getter (and the_orforms) on a SQLNULL— e.g. a table function called withf(n := NULL); a null handle was dereferenced; datetime::date_to_dayson an invalid date,timestamp_from_micros/timestamp_to_microson infinities and the far-negative range;datetime::time_from_micros/time_tz_from_bitson a time outside00:00:00–24:00:00, in a DuckDB built with assertions (a debug build, as thebundled-testfeature compiles), which failsTime::Convert'sD_ASSERT; a release build returned out-of-range fields;SelectionVector::newabove DuckDB's allocation limit;- a config option whose default does not cast to its type;
- catalog lookups for
Schema,Database,PreparedStatementandInvalidentry types; - a panic whose payload's own
Droppanics (panic_any(value)), in every callback macro,catch_ffi_panic, the typed scalar and typed table trampolines and the entry point's registration guard; - the entry points dereferenced a NULL
duckdb_database*when DuckDB'sget_databasefailed.
- every
-
Documented, not fixable here: C API aggregates crash under
agg(x) OVER ()andagg(x ORDER BY y). DuckDB'sCAPIAggregateUpdatedoes not flatten the state vector, and the window-constant and sorted-aggregate executors pass a one-element state array withcount > 1, so every aggregate registered through the C API — not only quack-rs's — reads out of bounds. Reproduced in plain C against DuckDB 1.4.4, 1.5.0 and 1.5.5; reported upstream as duckdb/duckdb#26109. Documented onAggregateFunctionBuilder,AggregateFunctionSetBuilder,FfiState, the aggregate book pages and as Pitfall L11. -
The one active advisory suppression is gone, because the crate behind it is. RUSTSEC-2026-0235 (
rkyv0.7.46) was suppressed inosv-scanner.toml, reachable only as quack-rs →duckdb→rust_decimal→rkyv. Induckdb1.10505.0rust_decimalbecame an optional dependency (it was required in 1.10504.0), and quack-rs does not enable it — so neither crate is inCargo.lockany more. Both the suppression and the CI step that re-proved it have been removed. -
cargo denywas scanning the wrong dependency graph.deny.tomlhadgraph.all-features = false; sinceduckdbis optional and the crate declares nodefaultfeature, neither the advisory scan nor the license scan ever evaluatedduckdb,arrow,chronoorrust_decimal. Nowall-features = true. -
ci.ymlandmutants.yml— the two workflows that build and execute pull-request code — had nopermissions:block, so the token inherited the repository default. Both now takecontents: read. -
mutants.ymlinterpolated a PR-derived file list straight into arun:block, where$(...)expands before bash parses the script. Moved toenv:. -
persist-credentials: falseon all 48actions/checkoutsteps. -
Soundness, third pass: more ways safe code, or a malformed input, could corrupt or read freed memory.
- A
FileHandlecould outlive its database and read freed memory. Breaking:FileHandle<'fs>(see Changed). - DuckDB v1.4.5 was missing from the ABI layout table, so a
duckdb-1-5build loaded into it was treated as an unknown engine instead of a layout mismatch, and underAbiPolicy::AllowUnknownEngineit loaded and segfaulted. v1.4.5 is now in the table, so such a build is reported as a layout mismatch there, as it is on v1.4.4.
- A
-
Process aborts, third pass. Each of these killed the host process:
ClientContext::config_optionon an option whose value is NULL:enable_profilingbefore it is set, or any option afterSET <option> = NULL. It now returnsNone.- A catalog
TypeorCollationlookup of a name that makes DuckDB autoload an extension, when the autoload fails. Breaking: now refused (see Changed). - A config option default that SQL's
TRY_CASTaccepts but DuckDB's built-in cast does not, such as aTIMESTAMPTZdefault with a time-zone name while ICU is loaded.ConfigOptionBuilder::registernow converts the default through the connection's ownTRY_CASTand hands DuckDB aValueof the option's type, so no cast runs inside the C API. Valuetemporal getters whose cast DuckDB implements by throwing: for exampleas_time()of'infinity'::TIMESTAMP(reachable from a table function's named parameter) oras_timestamp_ns()of anyTIMESTAMPafter 2262. They now returnNonewhere DuckDB would throw, and also for a result outside the target type's range, which DuckDB could produce but not render.- Rendering a temporal
Valuebuilt from an out-of-range payload. Breaking: the constructors now validate (see Changed). - The
AbiPolicy::Warndiagnostic usedeprintln!, which panics when writing to stderr fails, outside the entry point's panic guard. A failed write now loses the warning instead. validate_spdx_licenseon deeply nested parentheses (stack overflow). Breaking: nesting is capped at 64 (see Changed).
-
SqlMacro::registerran every statement in a body. A body of1); DROP TABLE t; SELECT (1registered and dropped the table. Breaking: a body with more than one statement is now refused (see Changed). -
SecretEntry::with_fieldon an existing key,with_providerandwith_scopefreed the value they replaced without zeroizing it (andwith_fieldalso the duplicate key), althoughDropzeroizes those same values. Each now zeroizes before replacing. A test that inspects every buffer as it is freed covers the replaced value, provider and scope; the duplicate-key case is covered only by reading the code.
Dependencies
libduckdb-sys/duckdb1.10504.0 → 1.10505.0 (DuckDB 1.5.4 → 1.5.5),cc1.2.64 → 1.4.7,arrow58.1.0 → 58.4.0. The relock removed 100 packages net:libduckdb-sys1.10505.0 swapped itsreqwestbuild-dependency forureq, taking the hyper/tokio/quinn/rustls trees with it. MSRV is unchanged at 1.86.0.- No
duckdb-1-5-5feature was added, deliberately. DuckDB 1.5.5 adds no C Extension API surface:extension_api.hppis byte-identical between v1.5.4 and v1.5.5 (sha2560232a22a…3017031, 89456 bytes, 546 function pointers in each). There would be nothing to gate. See the note inCargo.toml. - GitHub Actions, each SHA resolved against the upstream tag:
actions/checkoutv7.0.0 → v7.0.1,Swatinem/rust-cachev2.9.1 → v2.9.2,codecov/codecov-actionv7.0.0 → v7.1.1,actions/attest-build-provenancev4.1.1 → v4.2.2,actions/deploy-pagesv5.0.0 → v5.0.1. Theactions/configure-pagespin was already v6.0.0; only its comment said v5.0.0. - Pinned the four CI tools installed unpinned (
cargo-mutants27.1.0,cargo-semver-checks0.50.0,cargo-llvm-cov0.9.1,cargo-fuzz0.13.2).mutants.ymlreasons about behaviour introduced in cargo-mutants 27.0.0, which held only by luck of whatevercargo installfetched. Everycargo installstep now also clears the workflow-wideRUSTFLAGS: "-D warnings", which was compiling third-party trees with warnings-as-errors. dtolnay/rust-toolchainis pinned to a commit on the action'smasterbranch in every workflow and in the workflowgenerate_scaffoldwrites. The old pin was a commit of its regeneratedstablebranch that no ref reaches any more, so GitHub may garbage-collect it. Every use now passestoolchain:explicitly, asmaster'saction.ymlrequires; the pin does not pin the Rust version.
0.16.0 — 2026-08-19
Security
-
New
abimodule:duckdb_ext_api_v1layout verification.DuckDBhands a loadable extension a struct of function pointers. Its first 357 slots — the "stable prefix" — have been byte-for-byte identical in every release from v1.2.0 through v1.5.5, but everything past that is the unstable region, andDuckDBinserts new entries in the middle of it between releases (duckdb_appender_clearat slot 410 in v1.5.0,duckdb_geometry_type_get_crsat slot 493 in v1.5.2). Every quack-rs wrapper behind theduckdb-1-5/duckdb-1-5-3features — 105 C API functions covering scalar bind/init, copy functions, catalog access,ErrorData,FileSystem,Expression,SelectionVector, config options, table descriptions and the client context — lives in that region.DuckDBdoes not catch this: an extension stampedC_STRUCT+v1.2.0(the default) is accepted by anyDuckDBwhose C API version is at least v1.2.0 and then handed the whole struct, unstable region included. Loading such a build into aDuckDBwith a different layout silently dispatches to the wrong function pointers. Verified end-to-end: an extension built againstDuckDB1.5.0's headers, stampedC_STRUCT/v1.2.0, loaded intoDuckDB1.5.5 aborts the process withdouble free or corruption.abi::checkcompares the slot count of the compiled-in layout against the layout the running engine uses (resolved fromduckdb_library_version(), which sits at stable slot 7 and is therefore always dispatched correctly).init_extension/init_extension_v2and theentry_point!/entry_point_v2!macros now run that check under the newAbiPolicy::Strictdefault wheneverduckdb-1-5is enabled, turning the memory corruption above into aLOADerror that names the mismatch and the remedy.AbiPolicy::WarnandAbiPolicy::Trustopt out;Trustis the right choice for binaries stampedC_STRUCT_UNSTABLE, whereDuckDBalready pins the release. Extensions that stay on the stable prefix are unaffected and keep their forward compatibility.scripts/check-abi-table.pyre-derives the layout table from every upstream release header and runs in CI, so the table cannot drift asDuckDBreleases. -
DuckStringView::from_byteswas unsound. It was safe to call yet dereferenced the heap pointer embedded in bytes 8–15 of a pointer-formatduckdb_string_t, so safe code holding attacker-influenced bytes could read arbitrary memory. Replaced by two honest constructors:from_raw(unsafe, honours pointer format — what callbacks want) andinline_from_bytes(safe, returnsNonefor pointer-format values).from_bytesis deprecated and no longer dereferences. -
CopyGlobalInitInfo::get_file_pathcorrupted the heap. It calledduckdb_freeon the pointer fromduckdb_copy_function_global_init_get_file_path, which returnsinfo_ref.file_path.c_str()— the interior pointer of a C++std::stringDuckDBstill owns and destroys itself. EveryCOPY ... TOthrough a quack-rs copy function handed the allocator a pointer it never issued; the first live test of the path aborted withcorrupted size vs. prev_size in fastbins.Every other
duckdb_freecall site in the crate was then audited againstDuckDB's implementation, and all twelve are correct. The signature is not sufficient to decide:char *returns are owned andconst char *returns are usually borrowed, butduckdb_parameter_nameis declaredconst char *and returnsstrdup(...), so it is owned. Recorded asLESSONS.mdP11 with the full table.
Portability and feature-combination breakage
-
MAX_LIST_CHILD_CAPACITYwas typedusize, making the crate fail to compile forwasm32.DuckDB'sDConstants::MAX_VECTOR_SIZEis1ULL << 37ULL— anidx_t, not a pointer-sized value. As ausizeconst,1 << 37is a const-eval overflow wherever pointers are 32 bits, which is everywasm32target — andDuckDB's own extension CI builds three of them. Now typedu64, with a separateusize-clamped constant for the capacity arithmetic; on a 32-bit target the ceiling is larger than any allocationusizecan describe, sousize::MAXis the real limit. -
Three unit tests called
duckdb-1-5-gated methods without acfggate, breaking--features bundled-teston its own — a combination that builds the live-DuckDB tests but not the 1.5 wrappers.Value::display_stringandTableDescription::column_count/column_typeare the gated methods; the assertions around them are now gated too.Both defects compiled cleanly under every other feature combination.
scripts/check-matrix.shnow runs the combinations CI runs — includingbundled-testalone and thewasm32legs — in one command.
Security scanning
-
Security (OSV / GHSA)failed onRUSTSEC-2026-0235(rkyv0.7.46), an advisory that no build of this crate can reach.osv-scannerreadsCargo.lock, and Cargo pins optional dependencies there whether or not their feature is enabled, so the job flagged a crate that is never compiled. The chain is quack-rs →duckdb(optional dependency, enabled bybundled-test) →rust_decimal1.40.0 →rkyv(optional feature ofrust_decimal, not enabled).rust_decimalitself is built;rkyvis not. It is also not fixable here —rust_decimal1.40 constrainsrkyvto^0.7, and the fixed version is 0.8.17. The advisory is present onmainas well.Suppressed via a new
osv-scanner.tomlcarrying the full reachability argument. The suppression does not rest on that comment staying true: theosv-scanjob now re-derives it on every run, failing the build if anrkyvnode ever appears incargo tree --all-features --target all. The check tests the forward tree, becausecargo tree -iexits 0 with "nothing to print" for a lockfile-only package and so cannot tell "not built" from "built" by exit code; it also anchors onrust_decimalbeing present, so a truncated or failed tree reports as unverified rather than as clean.
CI guards
-
scripts/check-abi-table.pytreated an unreachable release header as proof the release did not exist. Itsfetchswallowed every exception — 404, timeout, DNS, 5xx alike — and returnedNone, which the caller printed as "not published (skipped)" and dropped from the derivation. One transient failure onv1.4.4therefore narrowed the derived range fromv1.4.0–v1.4.4tov1.4.0–v1.4.3and failed CI reportingsrc/abi.rsas stale. Following that advice would have shrunk the layout table and made the runtime guard refuseDuckDBversions it should accept — the exact failure the table exists to prevent.A definite 404 is now distinguished from every other failure, transient errors are retried, and a tag that could not be downloaded suspends the staleness comparison (exit 2, "could not check") instead of failing it. The other three guards each fetch a single file, so an empty fetch already meant no data rather than partial data.
-
All four guard jobs treated exit 2 as a failure. The scripts document it as "upstream unreachable, could not check", but the workflow ran them bare, so any non-zero failed the job — making every guard a network-flake away from a red build. They now surface exit 2 as a warning and fail only on exit 1.
Generated CI
- The generated CI workflow left one action unpinned. Three of its four
actions were SHA-pinned;
dtolnay/rust-toolchain@stablewas not, justified by a comment claiming its SHA "changes with each Rust release". That is not how the action works — it reads the toolchain fromrust-toolchain.tomlor itstoolchain:input at run time, so pinning the action's SHA does not pin the Rust version. quack-rs's own CI SHA-pins the same action and gets current stable. A branch is a moving target its owner can repoint, and a workflow step runs arbitrary code in the user's CI. All four are now pinned to the same SHAs quack-rs itself uses, and a test asserts everyuses:in the generated workflow carries a 40-character hex ref.
Fixed
The release-profile validator required the setting that breaks panic safety
-
validate_release_profilerequiredpanic = "abort", which makes every one of quack-rs's panic guards inert. quack-rs wraps everyextern "C"entry point — the extension entry point and every scalar/table/aggregate/cast/copy callback macro — incatch_unwind, so a panic in an extension's code becomes aDuckDBerror instead of a crash.catch_unwindcatches nothing underpanic = "abort": the runtime aborts before unwinding starts. Demonstrated directly rather than assumed —rustc -O panic_probe.rs → caught, process survived, exit 0 rustc -O -C panic=abort … → Aborted, exit 134— so the validator was telling extension authors to configure the one setting that turns a recoverable SQL error into a
SIGABRTthat kills the user's wholeDuckDBsession.The crate already disagreed with itself: the scaffold has generated
panic = "unwind"since the panic-safety work in this release, with a comment explaining why.validate_release_profilenow requires"unwind"and rejects"abort"with that explanation;ReleaseProfileCheck::panic_abortis renamedpanic_unwind. A new test asserts the scaffold and the validator agree, so they cannot drift apart again.The original justification — "panics across FFI boundaries are undefined behavior" — is also out of date: Rust defines an unwind escaping
extern "C"as an abort, and quack-rs catches panics before the boundary regardless.quack-rs's own
[profile.release]also saidpanic = "abort". Cargo ignores a dependency's profile so it changed nothing downstream, but it contradicted the crate's own advice; it now says"unwind".
A validator made legal function names unregisterable
-
validate_function_namerejected mixed-case names, and it gatestry_new— soScalarFunctionBuilder::try_new("myFunc")returnedErrand the function could not be registered through quack-rs at all.DuckDBitself shipsformatReadableSizeandformatReadableDecimalSize, and registering a camelCase name through the C API succeeds: verified againstDuckDB1.5.5, where the function is then callable asformatReadableThing,formatreadablethingandFORMATREADABLETHING, becauseDuckDBidentifiers are case-insensitive.The rule was justified as avoiding "catalog issues"; that test disproves it. Letters of either case are now accepted. Everything that would genuinely break is still rejected — a name needing quotes in SQL (
my-func,my func,my.func), one starting with a digit, one over 256 characters, one with an interior NUL.snake_caseremains the right convention and is documented as one, rather than enforced as a rule that blocks a legal name.The same relaxation applies to
AggregateFunctionBuilder,TableFunctionBuilderandSqlMacroparameter names, which share the validator.A regression test now runs
validate_function_nameover every function induckdb_functions()(746 of them) andvalidate_extension_nameover every entry induckdb_extensions(), asserting that everything identifier-shaped is accepted and every operator is not. That is how the defect was found.
A documented convention that was not being followed
-
"Every
unsafeblock inside this crate has a// SAFETY:comment" was not true.clippy::undocumented_unsafe_blocksreports 180 blocks in the library. Most are inside anunsafe fnand merely forward that function's own documented contract —unsafe_op_in_unsafe_fnis denied crate-wide, so those blocks are required syntax rather than new assertions — but around forty were in safe functions, where the crate rather than the caller is asserting the invariant, and those had nothing.The claim is replaced with the convention actually worth following, and that convention is now met: every
unsafeblock in a safe function carries a// SAFETY:comment. Auditing them also turned up three comments that described the wrong thing — twoduckdb_freecalls and aduckdb_destroy_valueannotated as if they were uses of the enclosing handle; those now say which allocation they own and why, cross-referencingLESSONS.mdP11.
The scaffold generated a description.yml that would be rejected
-
repo.refwas generated asmain.DuckDB's community-extension documentation is explicit: "Provide the hash of the latest commit on the branch targeting stable asref". The repository builds exactly that revision and signs the result, so a branch makes the build unreproducible. Of the 43 published extensions sampled, 41 pin a full 40-character hash and two pin a tag; none uses a branch.ScaffoldConfiggainsgit_ref, defaulting toREF_PLACEHOLDER("REPLACE_WITH_COMMIT_HASH") — deliberately not a valid revision, so it cannot be submitted by accident the waymainsilently could. The generated file carries a comment saying why, and a commented-outref_next. -
DescriptionYmlsilently droppedrepo.ref_next. It is a documented field: while a newDuckDBrelease is being prepared, the community repository tests an extension against both the latest stable release andmain, andref_nextnames the revision compatible withmain. Now parsed intogit_ref_next, empty when absent. -
The generated
description.ymlhad nodocs:section. All 43 published extensions have one — it is what renders on the community-extensions documentation site. The scaffold now emitshello_worldandextended_descriptionstubs.
Two more documented behaviours that were not the real ones
-
ClientContext::catalogdocumented an empty name as "the default catalog";DuckDBrejects it outright.duckdb_client_context_get_catalogstarts withif (!context || !name || strlen(name) == 0) return nullptr;— an empty string is the one value guaranteed to fail. The catalog of an in-memory database is namedmemory; a file database's is the file's stem. The doc now says so, along with the otherNonecase the C API imposes and quack-rs never mentioned:DuckDBcheckstransaction.HasActiveTransaction(), so this works inside a callback but not on an idle auto-commit connection. Both verified against 1.5.5 by a live test. -
ClientContext::config_optionaborts the process when asked for a setting that does not exist — on aDuckDBbuilt with debug assertions.duckdb_client_context_get_config_optioncallsTryGetCurrentSetting(...).GetScope()without first checking the lookup succeeded, andGetScope()assertsscope != SettingScope::INVALID. A releaseDuckDBcompiles the assertion out and the function's owndefault:arm returnsNULLas documented, so this never reproduces for end users and always reproduces in a test suite linking a debugDuckDB.This is a
DuckDBdefect, not a quack-rs one, but it makes the obvious "does the user have this setting?" probe unsafe. Documented on the method with the source lines, recorded asLESSONS.mdP12, and the abort-free alternative given:SELECT count(*) FROM duckdb_settings() WHERE name = ?.
Documentation claimed a bridge that cannot exist
-
The
secretsmodule described itself as bridging intoDuckDB's secrets system. There is no such bridge, and there cannot be. The extension C API has zero secret functions — not oneduckdb_secret_*among the 546 slots ofduckdb_ext_api_v1inDuckDB1.5.5. An extension cannot askDuckDBfor a credential through the C API at all.The only route is the
duckdb_secrets()table function, andDuckDBredacts sensitive fields there. Verified against 1.5.5:CREATE SECRET s (TYPE s3, KEY_ID 'AKIAEXAMPLE', SECRET 'super-secret-value'); SELECT secret_string FROM duckdb_secrets(); -- ...;key_id=AKIAEXAMPLE;secret=redactedThe module docs now say this plainly, and say what
SecretsManageractually is: a trait over the extension's own credential source, carrying the redactingDebug, zeroize-on-drop and absentPartialEqthat credential handling needs, rather than a route toDuckDB's store.The zeroize claim is also narrowed to what is true: it covers the buffers a
SecretEntryowns, not aStringthe caller still holds or one aStringabandoned when it grew.
The description.yml validator rejected 84% of real extensions
-
parse_description_ymlrejected 36 of the 43 published community extensions it was tested against. Its entire purpose is to tell an author their submission is valid before they open a PR, and it told almost everyone they were invalid. Four independent causes:-
requires_toolchainswas treated as required. It is not — only 14 of the 43 set it, and the community-extensions documentation does not list it as required. This alone rejected half the corpus. It is now optional;validate_rust_extensionstill requiresrustin it when present. -
YAML quotes were not stripped.
parse_kvdeliberately returned quoted values with their quotes and left stripping to each caller, and onlyexcluded_platformsdid. 12 of 43 files writeversion: '2025120401', so the parser saw'2025120401'— quotes included — and every version check failed on it.parse_kvnow unquotes, with a real balanced-quote check rather thantrim_matches, which would also eat""doubled""and a trailinga". -
validate_extension_versionimposed a formatDuckDBdoes not. It accepted only semver or a git hash; 11 of 43 published extensions use a date-based build id (2025120401).DuckDB's community-extension documentation specifies no version format at all — it says the descriptor carries "the version of the extension" and points at existing extensions as examples. The check is now what would actually break something: empty, over 64 characters, or containing anything outside[A-Za-z0-9._+-](whitespace, path separators, control characters).classify_extension_versionis unchanged —DuckDB's three-tier stability scheme is documented and is strict, and that function is where it belongs. -
windows_amd64_rtoolswas rejected. It is the R-tools Windows build (DuckDBPlatform()emits it underDUCKDB_PLATFORM_RTOOLS), it is not in the distribution matrix, and 14 of 43 published extensions exclude it.DUCKDB_PLATFORMSnow also accepts it and the four group names (linux,osx,wasm,windows— the top-level keys ofdistribution_matrix.json), while the newDUCKDB_CI_PLATFORMSkeeps the matrix-derived list the guard script checks. Empty segments from a trailing;— which five real files have — are skipped rather than reported as a platform named"".
All 43 now parse, with every name matching its directory.
-
-
Prose in the
docs:section was parsed as metadata. The scan was flat, so aversion:orlicense:line insidedocs.extended_description— free-form prose in 42 of the 43 files — silently overwrote the extension's real values. Demonstrated: alicense: FAKE-LICENSEline inside a documentation block made a valid file fail validation, and the same mechanism could have made an invalid one pass. The parser is now section-aware (onlyextension:andrepo:are read) and understands block scalars:key: |andkey: >bodies are captured as the field's value — literal blocks keeping line breaks, folded blocks joined — instead of being scanned for mappings. -
Three doc examples showed indented YAML that was not indented. A
\line-continuation in a Rust string literal eats the following line's leading whitespace, sodescription.ymlexamples inparse_description_yml,validate_description_yml_strandvalidate_rust_extensionwere parsing fully-unindented text. They only passed because the parser ignored indentation; making it section-aware exposed them. Rewritten as real multi-line literals.
Validators were giving wrong answers
-
The DuckDB platform list was stale in both directions.
validate::platformrejectedlinux_amd64_muslandlinux_arm64_musl— real, currently-built targets — so an extension that legitimately cannot support musl could not declare it. And it acceptedlinux_amd64_gcc4, whichDuckDBretired:DuckDBPlatform()induckdb/common/platform.hppnow raises a compile error for the legacy CXX ABI rather than emitting a_gcc4suffix, and it is absent from the distribution matrix. Excluding it was a silent no-op.The list is now derived from
config/distribution_matrix.jsoninduckdb/extension-ci-tools— the file the community-extensions build actually reads — andscripts/check-platform-table.pyplus a CI job fail when the two diverge. AddsDUCKDB_OPT_IN_PLATFORMSandis_opt_in_platform, because three of the twelve (linux_amd64_musl,linux_arm64_musl,windows_arm64) are only built on request, so excluding one of those is also a no-op.linux_amd64_gcc4gets a targeted error saying what happened to it, rather than "not a recognized DuckDB build target". -
validate_spdx_licenseclaimed valid licenses did not exist.COMMON_SPDX_LICENSESis a 42-entry shortlist of a 733-entry registry, but the rejection message read "is not a recognized SPDX identifier" — false forCC0-1.0,Python-2.0,BSD-4-Clauseand roughly 690 others. It now says the identifier is not on quack-rs's shortlist and points at the registry.Every entry was checked against
spdx/license-list-data: all 42 are real and none are deprecated.scripts/check-spdx-list.pyand a CI job keep it that way, and flag any newly-added identifier that is not OSI-approved (SSPL-1.0is listed and deliberately is not). The list is now sorted, with a test keeping it so. Also fixes the module doc, which called the fieldextension.licence; realdescription.ymlfiles — and quack-rs's own parser — uselicense.
Silent data corruption
-
The
UUIDaccessors disagreed about which 128 bits they meant, and the documentation said they agreed. AUUIDcolumn is physically aHUGEINT, butDuckDBstores it with the top bit flipped so that signed integer ordering matches UUID string ordering (BaseUUID::FromUHugeintinsrc/common/types/uuid.cppsubtracts 2^63 from the upper half). So:Accessor Returned For '11111111-…'::UUIDVectorReader::read_uuid(old)raw storage 0x9111…Value::as_uuidtextual bits 0x1111…Both were documented as "matching" the other. Handing one to the other — the obvious thing to do when a table function reads a
UUIDand builds aValuefrom it — silently changed the UUID's first hex digit.read_uuid/write_uuid(onVectorReader,VectorWriter,StructReader,StructWriterand both mocks) now apply the flip and take/returnu128textual bits, the same convention asValue::uuid/Value::as_uuidand every RustUuidtype.Value::uuid/as_uuidmove fromi128tou128for the same reason. The type change is deliberate: it turns a silent behaviour change into a compile error at every affected call site.read_i128/write_i128still read and write the raw storage, and the newvector::uuid_from_storage/vector::uuid_to_storageconvert between the two. Pinned by a live test that asserts the raw storage and the textual bits really do differ, so the conversion cannot quietly become a no-op.
Wrong results and unloadable builds
-
ChunkWriterno longer hardcodes a 2048-row capacity.DuckDBcan be built with a differentSTANDARD_VECTOR_SIZE, which is exactly why the C API exposesduckdb_vector_size(); assuming 2048 against a smaller build overruns the output vectors.ChunkWriter::newnow reads the running engine's value.ChunkWriter::newandDataChunk::into_chunk_writerare consequently no longerconst fn. -
The scaffold produced an extension
DuckDBrefuses to load. The generatedMakefilesetDUCKDB_PLATFORM_VERSION, whichextension-ci-toolsdoes not read, alongsideUSE_UNSTABLE_C_API=1.TARGET_DUCKDB_VERSIONtherefore fell back to itsv0.0.1default and the binary was stampedC_STRUCT_UNSTABLE/v0.0.1, whichDuckDBrejects with "The file was built specifically for DuckDB version 'v0.0.1'". The generatedMakefilenow setsEXTENSION_NAME(notEXT_NAME, whichbase.Makefileignores),TARGET_DUCKDB_VERSIONandUSE_UNSTABLE_C_APIfrom the newScaffoldConfigfields, and defines theall/configure/debug/release/test/cleantargets its own README and CI invoke. -
The scaffold generated
panic = "abort", which makes thecatch_unwindinscalar_callback!,table_scan_callback!and the extension entry point inert — so any panic in extension code killed the wholeDuckDBprocess instead of surfacing as a SQL error. Now generatespanic = "unwind". -
The scaffold pinned
quack-rs = "0.13"regardless of the generating crate's version. It now tracks the current major.minor. -
A freshly scaffolded project failed its own generated CI.
cargo clippy --all-targets -- -D warnings(which the generated workflow runs) rejected the generatedsrc/lib.rsforclippy::redundant_closureandsrc/wasm_lib.rsforspecial_module_name. Both are fixed; a newscaffold-e2eCI job builds the generated project, stamps its metadata footer, loads it into a realDuckDB, asserts the query result, and runs the generated lint gate. -
The generated CI referenced a nonexistent action (
duckdb/duckdb-build@v1) and ranmake testwithoutmake configure/make release, so it could not have passed. Replaced with a workflow that configures, builds and tests throughextension-ci-tools. -
The extension entry point ran user registration code without
catch_unwind. A panic in a registration closure unwound to theextern "C"entry point, aborting the process; it now becomes aLOADerror. Anapi_versioncontaining an interior NUL is also rejected up front instead of panicking insidelibduckdb-sys.
Behaviour documented after verification
Value::display_stringrenders a SQL literal, not display text:Value::varchar("hello")gives'hello'andValue::date(0)gives'1970-01-01'::DATE. Now documented with a table, since silently getting quotes and a cast suffix in a diagnostic is surprising.Value::as_strtruncates at an interior NUL, becauseduckdb_get_varcharreturns a NUL-terminatedchar *.DuckDBstores the full bytes; only this read path is limited. Documented on bothas_strandValue::varchar, and pinned by a test.
Documentation
-
The crate documented an "architectural limitation" that does not exist.
Cargo.toml,testing::in_memory_dband the book all stated thatVectorReader,VectorWriterandConnection::register_*"cannot be called incargo test" because they route through the dispatch table. Opening anInMemoryDbpopulates that table for the whole process, after which the entire C API — registration included — works. The newtests/ffi_roundtrip.rsregisters real scalar functions and round-trips every vector type through SQL: every integer width at its extremes,HUGEINT/UHUGEINTat theirs, floats and NaN, strings across the 12-byte inline/pointer boundary and multi-byte UTF-8, blobs containing NUL and non-UTF-8 bytes, all temporal types cross-checked againstDuckDB's own rendering,UUID,INTERVAL's three fields,DECIMALat all four physical widths, NULL in and out, multi-chunk scans, and a panicking callback surfacing as a SQL error. -
Documentation examples pinned
quack-rs = "0.13".
Added
Live tests for every previously untested C API path
-
Copy functions and replacement scans had no live tests at all. Between them they had 19 unit tests, none of which registered anything against a running
DuckDB— which is how a heap-corrupting free survived in a shipped API. Both now have end-to-end coverage:- A
COPY ... TO 'f' (FORMAT my_format)over 5000 rows, threading bind data and global state through all four lifecycle phases, asserting the sink saw every row and that both destructors ran exactly once (a leak or a double free is invisible without counting). - A replacement scan rewriting
SELECT * FROM '10.myfmt'into a table function call, plus the decline path — an identifier the callback ignores must still reachDuckDB's own error handling — and a panicking scan surfacing as a SQL error.
- A
-
Six more modules had unit tests but no live registration: scalar bind/init/local state,
Expression::fold, catalog lookup, config options, selection vectors and the instance cache. All now run against a realDuckDB, which turned up two more documentation defects (below) and confirmed the rest. -
copy_bind_callback!,copy_global_init_callback!,copy_sink_callback!andcopy_finalize_callback!. Every other callback kind had a panic-safe macro; the four copy-function phases did not, so a panic in one of them had nothing to catch it. Each routes the message through that phase's ownduckdb_copy_function_*_set_error. -
TypeId::try_from_duckdb_type— returnsOption<TypeId>instead of panicking on a type value this build does not know. Extensions routinely meet these: a column of a type added in a newerDuckDB, or a 1.5.x type reaching a build withoutduckdb-1-5.from_duckdb_typestill panics and now documents that callbacks should not use it. -
Fallible
LogicalTypeconstructors that previously panicked on an interior NUL in a caller-supplied name:try_struct_type_from_logical,try_union_type,try_union_type_from_logical,try_enum_type,try_set_alias. -
entry_point!/entry_point_v2!accept an optionalAbiPolicyas their second argument;init_extension_with_policy/init_extension_v2_with_policyare the function-level equivalents. -
examples/scaffold_to_dir.rs— writes a scaffolded project to disk, used by the newscaffold-e2eCI job.
Panic safety
-
A panic-safe wrapper macro for every callback kind. Only
scalar_callback!andtable_scan_callback!existed, so the other six kinds — table bind, table init, aggregate update/combine/finalize/destroy, cast, and replacement scan — were unguarded, and a panic in any of them aborted theDuckDBprocess. The aggregate ones are the worst case: they run on worker threads, so the abort comes from a thread the user never sees. New macros:table_bind_callback!,table_init_callback!,aggregate_update_callback!,aggregate_combine_callback!,aggregate_finalize_callback!,aggregate_destroy_callback!,cast_callback!,replacement_scan_callback!. Each routes the panic message to that callback kind's ownset_error;cast_callback!also returnsfalsesoTRY_CASTyields NULL. The aggregate destructor has no error channel in the C API, so its panic is caught and dropped — leaking beats aborting during query teardown. Verified end-to-end: a panicking aggregateupdateand a panicking cast both surface as SQL errors and leave the connection usable. -
The two existing macros now share
callback::panic_messageandcallback::message_to_c_stringwith the new ones. The latter replaces an interior NUL rather than dropping the diagnostic, which the oldif let Ok(c_msg) = CString::new(msg)silently did. -
TypedTableFunctionBuilderreported every panic as the same fixed string. It now includes the payload, so the user learns which assertion failed. -
Deprecated
FfiBindData::get_from_bind, which always returnedNoneand always will:DuckDBexposes noduckdb_bind_get_bind_data. Being safe and returningOption, it silently sentif let Some(..)down the wrong branch.
Capabilities
-
ListBuilderforLISTandMAPoutput vectors.duckdb_list_vector_reservetakes a total capacity and reallocates the child vector when it grows, so aVectorWriterobtained beforehand is left dangling. That makes the natural "reserve as you go, keep one writer" loop a use-after-free.ListBuilderre-fetches the child writer after every reserve, tracks the running offset, writes each parent{offset, length}entry, and grows geometrically so building a list is not quadratic.push_map_rowdoes the same forMAP. It also refuses capacities aboveMAX_LIST_CHILD_CAPACITY(duckdb::DConstants::MAX_VECTOR_SIZE), above whichDuckDBthrows a C++ exception that its own C API does not catch — an exception unwinding into Rust would be undefined behaviour. Covered by tests building 2000 lists and 1500 maps of varying length through real SQL. -
Valuegained the extractors and constructors it was missing. A table function declared with aTIMESTAMPorLISTparameter handed the bind callback aduckdb_valuethat could only be read viaas_str()and reparsed. Addsas_date,as_time,as_time_tz,as_timestamp,as_timestamp_tz,as_timestamp_s/ms/ns,as_interval,as_uuid,as_decimal,as_u128,list_len/list_child/list_items,struct_child,map_len/map_key/map_value, and the constructorsboolean,bigint,double,date,timestamp,varchar,uuid,null_value. -
querymodule — running SQL from inside an extension. The C API has everything needed (duckdb_query,duckdb_prepare,duckdb_bind_*,duckdb_fetch_chunk) and it is all in the stable prefix, but each handle has adestroythat must run exactly once, including on error paths.QueryResult,OwnedDataChunk,PreparedStatementandOwnedConnectionare RAII wrappers for those;Connectiongainsquery,execute,prepareandopen_connection.OwnedConnectioncovers the case the borrowed registration connection cannot: aduckdb_connectionholds its own reference to the database instance, so one opened during load stays valid afterwards — for a callback or a background thread. Verified by a test that closes theduckdb_databasehandle and keeps querying. -
datetimemodule — calendar conversions.DATE,TIMEandTIMESTAMPmove through vectors as raw integers; turning those into year/month/day meant reimplementing the proleptic Gregorian calendar andDuckDB's infinity sentinels.DuckDBalready exposes the conversions in the stable API, so this wraps them:date_from_days/date_to_days,time_from_micros/time_to_micros,timestamp_from_micros/timestamp_to_micros,time_tz_bits/time_tz_from_bits, the fouris_finite_*predicates, andHUGEINT/UHUGEINT/DECIMAL↔f64.Also exports the exact sentinel values as constants.
-infinityis-i32::MAX/-i64::MAX, noti32::MIN/i64::MIN—i32::MINis an ordinary finite date, and treating it as infinity would silently drop real rows. -
VectorWritercaches its validity bitmap.set_nullcalledduckdb_vector_ensure_validity_writable+duckdb_vector_get_validityon every row; both are now resolved once per vector (2 FFI calls instead of 4096 for an all-NULL 2048-row vector). Addsset_null_rangefor the batched case. -
Vector accessors for the remaining physical layouts:
write_u128/read_u128(UHUGEINT),write_decimal/read_decimal(which selecti16/i32/i64/i128from the declared width the wayDuckDBdoes),write_time_tz/read_time_tz, andTIMESTAMPTZ/TIMESTAMP_S/TIMESTAMP_MS/TIMESTAMP_NSaccessors.VectorReader::containsbounds-checks an index against the row count. -
Callback signature aliases are re-exported at their module roots:
scalar::ScalarFn(plusScalarBindFn/ScalarInitFnunderduckdb-1-5) andaggregate::{StateSizeFn, StateInitFn, UpdateFn, CombineFn, FinalizeFn, DestroyFn}, matching whattablealready did. -
The prelude re-exports
AbiPolicy, thedatetimetypes and thequerytypes. -
Registrar::register_config_option— the trait already covered scalar, scalar set, aggregate, aggregate set, table, SQL macro, cast and copy functions, but not config options, so an extension registering one could not have its whole registration closure exercised throughMockRegistrar. Added, withconfig_option_names/has_config_optionon the mock. -
secrets::list_duckdb_secrets— reads the secret metadataDuckDBdoes expose, viaduckdb_secrets(): name, type, provider, persistence, storage, scope prefixes and the redactedsecret_string. Enough to pick a scope, warn that a required secret is missing, or choose a provider. It returns aDuckDbSecretInfo, deliberately not aSecretEntry, so nothing suggests it carries credentials. A live test asserts both halves: the metadata comes through, and the credential provably does not. -
The appender is no longer behind
duckdb-1-5, and gained the row-at-a-time API it never had.duckdb_appender_*occupies slots 281–291 and 330–356 — the frozen stable prefix, unchanged since v1.2.0 — yet the whole module was gated onduckdb-1-5, whose wrappers live in the unstable region. Using the appender therefore forced an extension onto the version-pinned unstable ABI, for functionality that has been portable for four minor releases. Only three methods actually need 1.5 and stay gated:error_data,clearandappend_default_to_chunk.The 24 row-at-a-time functions were wrapped for the first time:
append_bool/_i8/_i16/_i32/_i64/_i128/_u8/_u16/_u32/_u64/_u128/_f32/_f64/_str/_bytes/_date/_time/_timestamp/_interval/_value/_null/_default,end_row,column_count,column_type,add_column,clear_columns, and arow(|row| …)helper that callsend_rowfor you. Previously the only way to insert a row was to build a wholeDataChunk.Three details that are easy to get wrong and are handled here:
append_strusesduckdb_append_varchar_length, so interior NUL bytes survive; that function narrows its length touint32_twith an unchecked cast inDuckDB's release builds, so longer strings are refused rather than truncated; andduckdb_append_valuedereferences its argument with no null check, so a nullValuehandle is refused. Covered by live tests that append every scalar type at its extremes, 5000 rows across several vectors, a short row, a constraint violation surfacing atclose, and aDEFAULT-filled column subset.New
appender::AppendErrorisErrorDatawithduckdb-1-5andExtensionErrorwithout, so enabling the feature upgrades the error type in place without changing any method's shape — existingduckdb-1-5code is unaffected. -
table_descriptionis no longer behindduckdb-1-5either. Slots 292–297 are stable; onlycolumn_countandcolumn_typeare 1.5 additions and stay gated. AddsTableDescription::with_catalog(duckdb_table_description_create_ext, for tables in another catalog) andcolumn_has_default(duckdb_column_has_default) — the latter being the only way to know whetherAppender::append_defaultwill succeed. -
FileHandlegained the looping I/O helpers, andsize/tellbecame fallible.duckdb_file_handle_readandduckdb_file_handle_writereturn "the number of bytes actually read/written" — a single call can come up short, which overhttpfsis routine rather than theoretical. Addsread_exact,read_to_endandwrite_all, which loop.size()andtell()changed fromi64toResult<u64, ErrorData>: the C API signals failure with a negative return, and the previous signature madehandle.size().max(0) as usize— silently treating an error as an empty file — the obvious thing to write. It was in this crate's own documentation. -
Value::type_id().Valuehad fortyas_*accessors and no way to ask what the value actually is, so reading aVARCHARwithas_i64()returned garbage rather than an error. Wrapsduckdb_get_value_type(stable prefix, slot 137, unchanged since v1.2.0), returningNonefor a null handle or a type id newer than this build knows. -
Every public type implements
Debug. 58 of them did not, which is Rust API guideline C-DEBUG and not cosmetic:Result::unwrap,Result::expect_err,assert_eq!, and#[derive(Debug)]on any downstream struct storing a quack-rs type all fail to compile without it.LogicalTypeandValueprint decoded state (type id, alias,DECIMALwidth/scale,DuckDB's own rendering) rather than a pointer; builders printset/unsetper callback, which is the question you have whenregisterreports a missing function;WarningCollectorusestry_lockso printing can neither block nor deadlock.missing_debug_implementationsis now enabled crate-wide, and CI's-D warningsmakes it an error.testing::InMemoryDbwas a 59th, only visible once the lint ran withbundled-teston.
Changed
MSRV
-
MSRV lowered 1.87.0 → 1.86.0. DuckDB's reusable
_extension_distribution.yml— the workflow the community-extensions repository builds every extension with — pinsdtolnay/rust-toolchain@… # 1.86.0for the WebAssembly job. quack-rs required 1.87.0, so Cargo refused, and no quack-rs extension could be built forwasm_mvp/wasm_eh/wasm_threadsby the official pipeline — despite the crate advertisingwasm32-unknown-emscriptensupport since 0.14.0.The entire 1.87 requirement was five
const fnaccessors callingVec::len(stabilised as const in 1.87). None can be reached in a const context —MockVectorWriter,StructReaderandStructWriterare all built at runtime — so droppingconstcosts nothing. 1.86.0 is now the floor for the library, its dev-dependencies (criterionneeds 1.86) and thehello-extexample, all verified.New
scripts/check-msrv-vs-duckdb-ci.pyand a CI job re-derive DuckDB's pinned toolchains from that workflow and fail if the MSRV creeps back above them. -
Breaking:
ScaffoldConfiggainstarget_duckdb_versionanduse_unstable_c_api.ScaffoldConfignow implementsDefault, so existing struct literals can add..ScaffoldConfig::default().generate_scaffoldrejects combinations that produce an unloadable binary — aC_STRUCTbuild claiming aDuckDBrelease as its-dv, or aC_STRUCT_UNSTABLEbuild claiming the C API version.
CI / tooling
- New
abi-tablejob:scripts/check-abi-table.pyverifiessrc/abi.rs's layout table against every upstreamDuckDBrelease header. - New
abi-guardjob: builds an extension againstDuckDB1.5.0's header layout, stamps itC_STRUCT, and asserts the load is refused with a layout diagnostic — a regression test for the corruption described above. - New
scaffold-e2ejob (see above). extension-loadnow stamps a real metadata footer and asserts query results rather than grepping the log for the word "error"; loading a bare.sobypassedDuckDB's metadata validation entirely.
0.15.0 — 2026-07-16
Added
Value::as_blob()for copying arbitrary binary data from aduckdb_value. (Thanks @adonm.)
Fixed
VectorReader::read_blob()now preserves non-UTF-8 bytes instead of returning an empty slice. (Thanks @adonm.)
Changed
- Dev/CI DuckDB bumped to 1.5.4 —
libduckdb-sys/duckdb1.10503.1 → 1.10504.0 in the root lockfile, thehello-extexample lockfile, and thebundled-test-prebuiltCI download (v1.5.3→v1.5.4). 1.5.4 is a bugfix release in the 1.5.x line; its C extension API version is unchanged (v1.2.0, verified fromduckdb_extension.h), soDUCKDB_API_VERSIONis unchanged and the publiclibduckdb-sysdependency range (>=1.4.4, <2) is untouched — downstream consumers are unaffected.
Security
crossbeam-epoch0.9.18 → 0.9.20 (root lockfile), resolving RUSTSEC-2026-0204 (invalid pointer dereference in thefmt::Pointerimpl). Reaches the tree only as a dev-dependency viacriterion → rayon → crossbeam-deque.quinn-proto0.11.14 → 0.11.15 (root and example lockfiles), resolving RUSTSEC-2026-0185 (CVSS 7.5). Reaches the lockfiles vialibduckdb-sys → reqwest → quinn(feature-union only; the loadable-extension build never links it).
CI / tooling
- Refreshed SHA-pinned GitHub Actions via Dependabot:
actions/checkoutv6.0.2 → v7.0.0,codecov/codecov-actionv6.0.1 → v7.0.0,actions/cachev5.0.5 → v6.1.0, andactions/attest-build-provenancev4.1.0 → v4.1.1. Also bumped theccbuild-dependency 1.2.63 → 1.2.64.
0.14.0 — 2026-06-07
Added
wasm32-unknown-emscriptensupport (the DuckDB-WASM target). The crate no longer hard-rejects non-64-bit targets with a top-levelcompile_error!, and theduckdb_string_tpointer slot is read as au64then narrowed tousize— lossless on 64-bit, and on wasm32 it yields the low 4 bytes of the 8-byte slot (the upper 4 are zero padding in DuckDB's 16-byte layout). The full public API, including theduckdb-1-5-3surface,cargo checks forwasm32-unknown-emscripten; CI now guards this. (Thanks @killzoner.)bundled-test-prebuiltfeature — links a pre-built libduckdb instead of compiling DuckDB from C++ source, for a much faster test build. Supply the library viaDUCKDB_DOWNLOAD_LIB=1(libduckdb-sysdownloads the upstream release zip) orDUCKDB_LIB_DIR=...(a libduckdb tree you already have).bundled-testcontinues to compile DuckDB from source. (Thanks @killzoner.)InMemoryDb::open_unsigned()opens an in-memory database withallow_unsigned_extensions=true, allowing downstream extension crates toLOADtheir own locally-built (unsigned).duckdb_extensionartifact for integration testing. (Thanks @killzoner.)
Changed
duckdbis now a purely optional dependency, activated only bybundled-test/bundled-test-prebuilt. It is no longer a dev-dependency, and there is no defaultbundledfeature. As a result, a plaincargo test— and every downstream consumer'sCargo.lock— no longer pulls the DuckDB + arrow tree, and the default test build no longer compiles DuckDB.
Security
tar0.4.45 → 0.4.46 in both the root and example lockfiles, resolving GHSA-3pv8-6f4r-ffg2 ("PAX header desynchronization", Moderate).taris alibduckdb-sysbuild-dependency, so it appears in bothCargo.lockfiles and raised one Dependabot alert each — the two moderate alerts reported onmain. This advisory is published in the GitHub Advisory Database (GHSA) but not the RustSec database, socargo denydid not flag it; the new OSV scan below closes that gap.- Bumped
cc1.2.62 → 1.2.63 (which movesshlex1.3.0 → 2.0.1) and refreshed thecodecov/codecov-actionpin to v6.0.1.
CI / tooling
- Added an OSV / GHSA advisory scan to CI (
osv-scanner, pinned to v2.3.8 via a checksum-verified binary) covering bothCargo.lockfiles.cargo denyconsults only the RustSec database; OSV.dev aggregates GHSA and RustSec, so GHSA-only advisories (such as thetarone above) now fail CI alongside the existing cargo-deny gate.
0.13.0 — 2026-05-24
Added
New safe wrappers for the DuckDB 1.5.0+ C extension API, all gated behind the
duckdb-1-5 feature, plus a new duckdb-1-5-3 feature that surfaces the two
DuckDB 1.5.3 type-enum values. DuckDB 1.5.3's C extension function-pointer API
(version v1.2.0) is unchanged from 1.5.2; the one new C addition — the
DUCKDB_TYPE_VARIANT (41) type-enum value — is now exposed as TypeId::Variant
behind the duckdb-1-5-3 feature (see below). So the additions below mostly
expose 1.5.x capabilities the SDK had not previously wrapped rather than anything
new to 1.5.3 specifically.
error_datamodule —ErrorData, an RAII wrapper overduckdb_error_data(the structured error type returned by several 1.5 APIs). Carries aDuckDbErrorTypecategory and a message, and converts intoExtensionError. Adds the free functioncheck_valid_utf8, exposingDuckDB's own UTF-8 validator.expressionmodule —Expression, an RAII wrapper overduckdb_expression, withreturn_type,is_foldable, andfold. This closes a real gap:ScalarBindInfoalready returned a raw, unusableduckdb_expressionfromget_argument; the newScalarBindInfo::argumentreturns a safeExpression, so bind callbacks can inspect argument types and pre-fold constant arguments once at bind time.file_systemmodule —FileSystem,FileHandle,FileOpenOptions, andFileFlag: read and write files throughDuckDB's virtual file system (honouringhttpfs, in-memory files, and other registered file systems) instead of reaching forstd::fs.appendermodule —Appender: bulk row insertion (create, append aDataChunk, flush, close) plus the 1.5 additionsclear(revert buffered rows),error_data(structured errors), andappend_default_to_chunk.selection_vectormodule —SelectionVector: allocate and fill zero-copy row-index selection vectors.instance_cachemodule —InstanceCache: share one underlying database instance across repeated opens of the same path.Valuegainsdisplay_string(canonical string rendering of any value, viaduckdb_value_to_string) andTIME_NSaccessorsValue::time_ns/Value::as_time_ns(pairing with the existingTypeId::TimeNs).Cataloggainstype_name(the catalog's storage type, e.g."duckdb"or a storage extension's name).- All new public types are re-exported from the
preludebehind theduckdb-1-5feature. duckdb-1-5-3feature +TypeId::Variant/TypeId::Geometry— a new feature flag (duckdb-1-5-3, which impliesduckdb-1-5) exposes theDUCKDB_TYPE_VARIANT(41, added in DuckDB 1.5.3) andDUCKDB_TYPE_GEOMETRY(40) type-enum values asTypeId::VariantandTypeId::Geometry, with the matchingto_duckdb_type/from_duckdb_type/sql_name/Displaycoverage. It is a separate gate because these constants postdate theduckdb-1-5feature's 1.5.0 floor and requirelibduckdb-sys >= 1.10503.1; keeping them out ofduckdb-1-5preserves compatibility for consumers pinned to libduckdb-sys 1.5.0–1.5.2.ErrorDatais now a first-class error type — implementsstd::fmt::Displayandstd::error::Error, gains a structuredDebugimpl, and converts intoExtensionErrorviaFrom(alongside the existinginto_extension_error) so it propagates through?.DuckDbErrorTypenow implementsDisplay(backed by a newpub const fn as_str).TableDescription::as_raw()— exposes the raw handle, matching the accessor convention of the other 1.5 wrappers.
Changed
duckdb/libduckdb-sys1.10502.0 → 1.10503.1 (DuckDB 1.5.2 → 1.5.3) in both the workspace andexamples/hello-extCargo.lock. DuckDB 1.5.3 is a bugfix release (announcement); since the>=1.4.4, <2constraint already permitted it, the bundled fixes are picked up purely by the lock-file update with no source changes required for the bump itself.cc→ 1.2.62 in bothCargo.lockfiles — workspace (1.2.61 → 1.2.62, folding in Dependabot PR #89, thepatch-updatesgroup) andexamples/hello-ext(1.2.57 → 1.2.62, re-syncing the example lock's oldercc). Build-dependency; no API impact.- MSRV corrected to 1.87.0. The crate declared
rust-version = "1.84.1", butlibduckdb-sys(1.5.x line, a non-optional dependency) isedition = "2024"/rust-version = "1.85.1"— so quack-rs has in fact required Rust ≥ 1.85.1 since before this release (cargo +1.84.1 checkcannot even parse the manifest). The declared MSRV, the CIMSRVjob (now explicitly pinned withtoolchain: "1.87.0"so it genuinely gates instead of silently falling back to therust-toolchain.tomlstable channel), the release matrix, and all docs/badges are updated to 1.87.0 — a small headroom margin above the 1.85.1 floor.
Fixed
TypeId::from_duckdb_typeno longer panics on theduckdb-1-5type-enum values. It previously recognised only the base (1.4) values andpanic!ed on everything else — including theduckdb-1-5values (TIME_NS,ANY,BIGNUM/VARINT,SQLNULL,INTEGER_LITERAL,STRING_LITERAL). Because the publicLogicalType::get_type_id()calls it, inspecting such a type inside a bind callback could panic across the FFI boundary (Pitfall L3). It now maps every variant available in the active feature set (plus theduckdb-1-5-3GEOMETRY/VARIANTvalues when that feature is enabled).TableDescription'sDropnow null-checks the handle before destroying it, matching every other RAII wrapper in the crate.
Documentation
- New book section "DuckDB 1.5+ APIs" — dedicated guide pages for the
error_data,expression,appender,file_system,selection_vector, andinstance_cachemodules, wired intoSUMMARY.md. - Refreshed the reference docs (
docs/architecture.md,docs/ffi-reference.md, theTypeIdreference,CONTRIBUTING.md/book source trees) to cover the new modules, and updated the VARIANT/GEOMETRY entries inKnown Limitations,concepts/types.md, and theTypeIdreference to document the newduckdb-1-5-3gate (previously tracked as a follow-up). - Added
// SAFETY:comments to previously-undocumentedunsafeblocks in theget_client_contextaccessors (scalar,copy_function) andTableDescription::create, and SPDX headers tobenches/interval_bench.rsand the test submodule files — closing the last gaps against the crate's own "every file / everyunsafeblock" conventions. - Corrected the README install note (it claimed v0.11.0 was the latest published
crate; v0.12.1 was in fact already on crates.io) and bumped install-example
version references throughout the README, book, and scaffold template to
0.13.
CI
- docs.rs now builds with
duckdb-1-5-3([package.metadata.docs.rs]), so the feature-gated modules (appender,error_data,file_system, …) and the newTypeIdvariants render on docs.rs and the README's docs.rs links resolve. Previously docs.rs built the empty default feature set and omitted them. - CI exercises the
duckdb-1-5-3feature — the feature job now runscheck/test/clippyforduckdb-1-5-3alongsideduckdb-1-5, and theClippy (beta)anddocjobs useduckdb-1-5-3. - Fixed the
NightlyCI job silently running stable — the SHA-pinneddtolnay/rust-toolchainstep lackedwith: toolchain: nightly, so it fell back to therust-toolchain.tomlstable channel (the same class of bug previously fixed for the MSRV job). - Mutation testing scoped to testable code — DuckDB FFI-wrapper modules
whose methods require a live runtime (and whose tests are
bundled-test-gated or absent) are excluded fromcargo mutants, since their mutants can't be killed by unit tests. This extends the existing exclusion pattern to the 1.5.x wrappers —expression,file_system,appender,selection_vector,instance_cache,table_description, and the scalar/copy*Infoaccessors. Pure-logic code (e.g.DuckDbErrorType, theTypeIdconversions) stays in scope. The mutants feature set is bumped toduckdb-1-5-3.
0.12.1 — 2026-05-01
Security
Closes nine GitHub Dependabot alerts (two High, seven Low) split across
the workspace Cargo.lock and examples/hello-ext/Cargo.lock.
-
rustls-webpki0.103.10 → 0.103.13 — picks up the fix for three RustSec advisories reachable via thebundledDuckDB build's transitivereqwest→rustlschain:- RUSTSEC-2026-0098 (GHSA-965h-392x-2mh5) —
nameConstraintswith URI name restrictions were silently ignored instead of enforced. Patched in 0.103.12+; the URI-name path is not on the public Web PKI, so impact is limited to private-PKI consumers. - RUSTSEC-2026-0103 (GHSA-xgp8-3hg3-c2mh) — name-constraint enforcement accepted certificates asserting a wildcard subject name. Reachable only after signature verification and requires misissuance. Patched in 0.103.12+.
- RUSTSEC-2026-0104 — reachable panic when parsing certificate
revocation lists with a syntactically valid empty
BIT STRINGin theonlySomeReasonselement of anIssuingDistributionPointCRL extension. Affects only applications that use CRLs. Patched in 0.103.13+.
Neither path is exercised by
quack-rsitself, but the advisories tripcargo denyfor any downstream consumer that has not yet bumped, so shipping a release that resolves them is the path of least friction. - RUSTSEC-2026-0098 (GHSA-965h-392x-2mh5) —
-
rand0.9.2 → 0.9.4 / 0.8.5 → 0.8.6 — picks up the fix for RUSTSEC-2026-0097 (GHSA-cq8v-f236-94qc) —ThreadRngcould produce an aliased&mut BlockRng<ReseedingCore>(Stacked-Borrows UB) when a custom logger reenteredrand::rng()from inside a reseed at trace-level logging. Triggering the unsoundness requires a custom global logger that pulls fromrand::rng()while reseeding, which is not a patternquack-rsuses, but the advisory matches by version range so resolving it removes the alert noise. Patched on every affected line: 0.8.6+, 0.9.3+, 0.10.1+.
Changed
- Workspace
Cargo.lockbumps —cc1.2.59 → 1.2.61 (build-dep; no API impact)duckdb/libduckdb-sys1.10501.0 → 1.10502.0 (latest patch release; no API impact forquack-rs)rand0.8.5 → 0.8.6 (transitive viarust_decimal; security)rand0.9.2 → 0.9.4 (transitive viaproptestdev-dep; security)
examples/hello-extCargo.lockbumps —libduckdb-sys1.10501.0 → 1.10502.0rand0.9.2 → 0.9.4 (security)rustls-webpki0.103.10 → 0.103.13 (security; matches workspace)
CI
-
GitHub Actions pin updates —
actions/cachev5.0.4→v5.0.5actions/upload-artifactv7.0.0→v7.0.1actions/upload-pages-artifactv4.0.0→v5.0.0
All updates retain SHA-pinned references for supply-chain integrity.
-
New informational
Clippy (beta)job — runs the samecargo clippy --all-targets --features duckdb-1-5 -- -D warningsinvocation on thebetaRust toolchain. Markedcontinue-on-errorso a beta-only lint regression does not block the merge queue, but surfaces six weeks before the lint reachesstable. Originally added in response toclippy::map_unwrap_orgraduating tostablein Rust 1.95.0 and bitingsrc/warning.rsafter the toolchain rolled forward.
Fixed
clippy::map_unwrap_oronWarningCollector::len—self.warnings.lock().map(|w| w.len()).unwrap_or(0)rewritten asself.warnings.lock().map_or(0, |w| w.len()). Behaviour-preserving; fixesClippyandTest duckdb-1-5 featurejobs under Rust 1.95.0.clippy::map_unwrap_or_defaultonWarningCollector::snapshot(defensive) — same rewrite for the siblingmap(|w| w.clone()).unwrap_or_default()call. Caught proactively alongside the above; otherwise would have surfaced the next time the lint promotion round-trips throughpedanticornursery.
0.12.0 — 2026-04-09
Added
-
TypedTableFunctionBuilder<S>with closure-basedbind/scan— new high-level layer on top ofTableFunctionBuilderthat lets extensions register table functions via two safe Rust closures instead of hand-rolledunsafe extern "C" fntrampolines. Entry point isTableFunctionBuilder::with_state::<S, _>(|bind| Ok(S { ... })), followed by.scan(|state, chunk| { ... Ok(()) })and.build()?to recover a fully configuredTableFunctionBuilderusable with anyRegistrar. Highlights:- The
bindclosure receives&BindInfo, declares the output schema, reads parameters, and returns the typed scan stateS: Send + 'static. - The
scanclosure receives&mut Sand aDataChunkfor the output chunk. Returning with chunk size zero signals end-of-stream. - Panics in user closures are caught via
std::panic::catch_unwindand surfaced throughduckdb_bind/init/function_set_error; the scan forces chunk size to zero on panic so the query terminates safely. - Scan state is carried from
bindthroughinitintoinit_dataso the scan callback can hold&mut Swithout extra ceremony. - Because
Sis only required to beSend, scans are serialised by callingset_max_threads(1). Extensions that need true multi-worker parallelism should continue to use the rawTableFunctionBuilderwithlocal_init. - Re-exported from the prelude as
TypedTableFunctionBuilder.
This is proposal A from the duck_net "quack-rs enhancements" list and eliminates the raw bind/init/scan trampolines that every FFI-heavy extension would otherwise write by hand.
- The
-
ExtensionError: additionalFromimpls —From<std::io::Error>,From<std::ffi::NulError>, andFrom<std::fmt::Error>allow the?operator to propagate common error types directly inregister_all()without.map_err(). This eliminates the need forpanic!()when operations like tokio runtime allocation fail during extension initialization. -
tlsmodule —TlsConfigProvidertrait for type-erased TLS client configuration injection. HTTP-capable extensions (e.g.,duck_net) implement this trait to supply custom CA bundles, client certificates for mTLS, or restricted cipher suites through a uniform interface. Usesstd::any::Anysoquack-rshas no dependency on any specific TLS library. Security hardened:client_config()returnsResultfor fallible config creation,accepts_invalid_certs()andmin_tls_version()enable security auditing,config_type_name()allows safe pre-downcast verification, andaudit_tls_provider()integrates with thewarningmodule to automatically flag CWE-295 (cert validation bypass) and CWE-327 (deprecated TLS versions). IncludesTlsVersionenum withis_deprecated()andOrdordering. -
warningmodule — structured security warning API withExtensionWarning,WarningSeverity(Info/Low/Medium/High/Critical), andWarningCollector. Extensions that touch external resources emit warnings with machine-readable codes and optional CWE identifiers.WarningCollectoris thread-safe (Mutex-backed) and supportsemit(),snapshot(),drain(), andclear(). -
secretsmodule —SecretsManagertrait andSecretEntrytype for bridging into DuckDB's nativeCREATE SECRETstorage. Extensions implementSecretsManagerto provideget_secret(),list_secrets(), andremove_secret()through a safe Rust interface.SecretEntryuses a builder pattern withwith_provider(),with_scope(), andwith_field(). Security hardened:Debugredacts field values,Dropzeroizes sensitive data viawrite_volatile,PartialEqintentionally omitted to prevent timing side-channels, and fields are private with accessor methods. -
StructWriter::child_list_vector(field_idx)— semantic alias forchild_vector()that makes the intent clear when a struct field has LIST type. Returns the rawduckdb_vectorhandle for use withListVectormethods (reserve,set_entry,set_size,child_writer, etc.). -
Prelude additions —
TlsConfigProvider,ExtensionWarning,WarningSeverity,WarningCollector,SecretEntry,SecretsManagerre-exported fromquack_rs::prelude.
0.11.0 — 2026-03-30
Added
-
StructWriter::child_vector(field_idx)— returns the rawduckdb_vectorhandle for a struct field, enablingListVector/MapVector/ArrayVectoroperations on nested complex types without raw FFI calls. -
StructReader::child_vector(field_idx)— read-side counterpart for accessing nested complex type fields within STRUCT input vectors. -
ChunkWriter::vector(col_idx)— rawduckdb_vectoraccess for complex column types (LIST, MAP, ARRAY) from within aChunkWriter. -
ChunkWriter::column_count()— returns the number of columns in the chunk without needing a separateDataChunk. -
VectorWriter::set_valid(row)— marks a row as non-NULL, undoing a previousset_null()call. Callsensure_validity_writableautomatically. -
StructWriter::set_valid(row, field_idx)— batched version ofVectorWriter::set_valid()for STRUCT fields. -
ReplacementScanInfo::add_parameter_raw(value)— adds anyduckdb_valueas a parameter to a replacement scan redirect, enabling non-VARCHAR parameter types (INTEGER, BIGINT, BOOLEAN, etc.). -
ReplacementScanInfo::add_i64_parameter(value)— convenience method for adding BIGINT parameters to replacement scan redirects. -
ReplacementScanInfo::add_bool_parameter(value)— convenience method for adding BOOLEAN parameters to replacement scan redirects.
Changed
table_scan_callback!error reporting — the macro now extracts the panic message and reports it to DuckDB viaduckdb_function_set_errorbefore setting chunk size to 0. Previously, panics silently ended the stream with no error message visible to the user.
0.10.0 — 2026-03-29
Added
-
StructWriter(vector::struct_writermodule) — batched, typed writer for STRUCT output vectors. Pre-createsVectorWriters for all fields at construction, then exposeswrite_bool,write_varchar,write_i64,write_date,write_timestamp,write_time,write_blob,write_uuid,set_null, etc. Eliminates ~120 rawduckdb_struct_vector_get_childcalls across typical extensions. -
StructReader(vector::struct_readermodule) — batched, typed reader for STRUCT input vectors. Read-side counterpart toStructWriterwithread_bool,read_str,read_i64,read_date,read_timestamp,read_blob,read_uuid,is_valid, etc. -
ChunkWriter(chunk_writermodule) — auto-sizing chunk writer for table function scan callbacks. Tracks rows vianext_row()and automatically callsduckdb_data_chunk_set_sizeonDrop, preventing forgotten-set-size bugs. -
scalar_callback!/table_scan_callback!macros (callbackmodule) — wrapunsafe extern "C"callbacks withstd::panic::catch_unwind, preventing undefined behaviour from panics unwinding across the FFI boundary. Scalar errors are reported viaduckdb_scalar_function_set_error; table scan panics set chunk size to 0 (end of stream). -
Valueextraction methods —as_i8(),as_i16(),as_u8(),as_u16(),as_u32(),as_u64(),as_i128()covering every DuckDB integer type viaduckdb_get_int8/int16/uint8/uint16/uint32/uint64/hugeint. Plusas_str_or(),as_str_or_default(), and_or(default)null-safe variants for all types. -
VectorReader—read_date(),read_timestamp(),read_time(),read_blob(),read_uuid()semantic methods for DATE, TIMESTAMP, TIME, BLOB, and UUID column types. -
VectorWriter—write_date(),write_timestamp(),write_time(),write_blob(),write_uuid()semantic methods matching reader additions. -
DataChunkconvenience methods —struct_writer(col, fields),struct_reader(col, fields),struct_field_reader(col, field),into_chunk_writer()bridging to the newStructWriter,StructReader, andChunkWritertypes. -
ChunkWriter::struct_writer(col, fields)— convenience bridge toStructWriterfrom within aChunkWriter. -
MockVectorWriter—write_blob(),write_date(),write_timestamp(),write_time(),write_uuid()matching realVectorWriteradditions. -
MockVectorWriter/MockVectorReader—try_get_i8(),try_get_i16(),try_get_u8(),try_get_u16(),try_get_u32(),try_get_u64(),try_get_f32(),try_get_i128(),try_get_blob(),try_get_uuid()closing the type coverage asymmetry between mock and real vector types. -
MockVectorReaderconstructors —from_i8s(),from_i16s(),from_u8s(),from_u16s(),from_u32s(),from_u64s(),from_f32s(),from_i128s(),from_intervals(),from_blobs()for everyMockDuckValuevariant. -
MockDuckValue::Blob(Vec<u8>)— new variant for BLOB testing. -
Prelude additions —
StructReader,StructWriter,ChunkWriterre-exported.
Changed
-
TableDescription::column_type()now returnsOption<LogicalType>(RAII) instead of rawduckdb_logical_type, eliminating manual destroy calls by callers. -
Version references updated — all documentation, examples, scaffold templates, and book pages now reference
quack-rs = "0.10"(was"0.9").
Fixed
-
FFI callback panic safety — replaced 13
CString::new(...).expect(...)calls in FFI callback contexts (table/info, scalar/info, cast/builder, aggregate/info, copy_function/info, replacement_scan) with non-panickingstr_to_cstring()that truncates at interior null bytes. Fully honours the "no panics across FFI" design principle (Pitfall L3). -
Non-idiomatic
&mut { expr }syntax — replaced 8 instances in builderregister()methods and 1 in replacement scan with idiomatic&raw mut.
0.9.0 — 2026-03-29
Added
-
ValueRAII wrapper (valuemodule) — owned wrapper aroundduckdb_valuewith automatic cleanup viaDrop. Typed extraction methods:as_str(),as_i64(),as_i32(),as_f64(),as_f32(),as_bool(). Eliminates manualduckdb_destroy_valuecalls and prevents memory leaks in bind parameter extraction. -
DataChunkwrapper (data_chunkmodule) — ergonomic non-owning wrapper aroundduckdb_data_chunkwithreader(col),writer(col),size(),set_size(n),column_count(), andvector(col)methods. Eliminates rawduckdb_data_chunk_get_vector/duckdb_data_chunk_set_sizecalls in scan callbacks. -
VectorWriter::write_str(idx, value)— alias forwrite_varcharfor discoverability. Extension authors searching forwrite_strnow find it immediately. -
BindInfo::get_parameter_value(index)— returns an ownedValueinstead of a rawduckdb_value, preventing memory leaks. -
BindInfo::get_named_parameter_value(name)— same for named parameters. -
MapVector::key_writer(vector)/value_writer(vector)— createVectorWriterinstances for MAP key and value child vectors directly. -
MapVector::key_reader(vector, count)/value_reader(vector, count)— createVectorReaderinstances for MAP key and value child vectors. -
MockVectorWriter::write_str(idx, value)— alias forwrite_varcharmatching theVectorWriterAPI addition. -
Prelude additions —
Value,DataChunk, andValidityBitmapare now re-exported fromquack_rs::prelude.
Changed
- Version references updated — all documentation, examples, scaffold
templates, and book pages now reference
quack-rs = "0.9"(was"0.7").
0.8.0 — 2026-03-28
Added
-
LogicalType::from_raw(ptr)— construct aLogicalTypefrom an existing rawduckdb_logical_typehandle, taking ownership. -
LogicalTypecomplex type constructors —decimal(width, scale),array(element, size),array_from_logical(element, size),union_type(members),union_type_from_logical(members),enum_type(members). -
LogicalType_from_logicalvariants —struct_type_from_logical,list_from_logical,map_from_logicalacceptLogicalTypevalues for nested complex types that cannot be expressed as simpleTypeId. -
LogicalTypeintrospection methods (20 methods) —get_type_id,get_alias,set_alias,decimal_width,decimal_scale,decimal_internal_type,enum_internal_type,enum_dictionary_size,enum_dictionary_value,list_child_type,map_key_type,map_value_type,struct_child_count,struct_child_name,struct_child_type,union_member_count,union_member_name,union_member_type,array_size,array_child_type. -
TypeId::from_duckdb_type(raw)— reverse conversion from rawDUCKDB_TYPEC enum toTypeId. -
ScalarFunctionBuilder::extra_info(data, destroy)— attach arbitrary data to a scalar function, accessible viaduckdb_function_get_extra_infoin callbacks. -
ScalarOverloadBuilder::extra_info(data, destroy)— same for scalar function set overloads. -
AggregateFunctionBuilder::extra_info(data, destroy)— attach arbitrary data to an aggregate function. -
TableFunctionBuilder::param_logical(logical_type)— add a positional parameter with a complexLogicalType. -
TableFunctionBuilder::named_param_logical(name, logical_type)— add a named parameter with a complexLogicalType. -
CastFunctionBuilder::new_logical(source, target)— construct a cast builder usingLogicalTypevalues for complex source/target types. -
ScalarFunctionInfo— callback wrapper withget_extra_info(),set_error(), and (duckdb-1-5)get_bind_data(),get_state(). -
ScalarBindInfo(duckdb-1-5) — scalar bind callback wrapper withargument_count(),get_argument(),get_extra_info(),set_bind_data(),set_error(),get_client_context(). -
ScalarInitInfo(duckdb-1-5) — scalar init callback wrapper withget_extra_info(),get_bind_data(),set_state(),set_error(),get_client_context(). -
AggregateFunctionInfo— aggregate callback wrapper withget_extra_info()andset_error(). -
CopyBindInfo(duckdb-1-5) — copy bind callback wrapper withcolumn_count(),column_type(),get_extra_info(),set_bind_data(),set_error(),get_client_context(). -
CopyGlobalInitInfo(duckdb-1-5) — copy global init callback wrapper withget_bind_data(),get_extra_info(),get_file_path(),set_global_state(),set_error(),get_client_context(). -
CopySinkInfo(duckdb-1-5) — copy sink callback wrapper withget_bind_data(),get_extra_info(),get_global_state(),set_error(),get_client_context(). -
CopyFinalizeInfo(duckdb-1-5) — copy finalize callback wrapper withget_bind_data(),get_extra_info(),get_global_state(),set_error(),get_client_context(). -
BindInfo::get_parameter(index)— retrieve positional parameter value in table function bind callbacks. -
BindInfo::get_named_parameter(name)— retrieve named parameter value in table function bind callbacks. -
BindInfo::get_extra_info(),InitInfo::get_extra_info(),FunctionInfo::get_extra_info()— access extra info from table function callbacks. -
get_client_context()— available onBindInfo(table),ScalarBindInfo,ScalarInitInfo,CopyBindInfo,CopyGlobalInitInfo,CopySinkInfo,CopyFinalizeInfo. Returns aClientContextRAII wrapper. -
ArrayVector— helper for fixed-size array vectors withget_child(). -
vector_size()— returns the default DuckDB vector size (typically 2048). -
vector_get_column_type(vector)— returns theLogicalTypeof a vector. -
Prelude additions —
StructVector,ListVector,MapVector,ArrayVector,ScalarFunctionInfo,AggregateFunctionInfonow re-exported fromquack_rs::prelude.
Changed
-
CastFunctionBuilder::source()/target()now returnOption<TypeId>instead ofTypeId, returningNonewhen the builder was created vianew_logical(). This is a breaking change. -
CastRecord::source/targetfields changed fromTypeIdtoOption<TypeId>to match the builder change.
0.7.1 — 2026-03-27
Added
-
TypeId::Any— wildcard type for function overload resolution. Maps toDUCKDB_TYPE_ANYin the C API. Requiresduckdb-1-5feature. -
TypeId::Varint— variable-length arbitrary-precision integer. Maps toDUCKDB_TYPE_BIGNUMin the C API, exposed asVARINTin SQL. Requiresduckdb-1-5feature. -
TypeId::SqlNull— explicit SQL NULL type representing the type of a bareNULLliteral before type resolution. Maps toDUCKDB_TYPE_SQLNULLin the C API. Requiresduckdb-1-5feature. -
TypeId::IntegerLiteral— internal type for unresolved integer literals during overload resolution. Maps toDUCKDB_TYPE_INTEGER_LITERAL. Requiresduckdb-1-5feature. -
TypeId::StringLiteral— internal type for unresolved string literals during overload resolution. Maps toDUCKDB_TYPE_STRING_LITERAL. Requiresduckdb-1-5feature. -
MockVectorReader/MockVectorWritertests — 12 new tests coveringfrom_i32s,from_f64s,from_boolsconstructors, typed getters (i32,f64,bool),u16/i128/intervalround-trips, wrong-type returns None, andis_empty. -
DuckDB v1.5.1 compatibility evaluation — comprehensive analysis of all 80+ changes in DuckDB v1.5.1 against quack-rs. See
docs/duckdb-v1.5.1-evaluation.md.
Fixed
- ARM64 / aarch64 build — replaced all
.cast::<i8>()and*const i8pointer casts withstd::os::raw::c_char, which resolves toi8on x86-64 andu8on ARM64 (where Ccharis unsigned). EliminatesE0308/E0277mismatched-types errors when cross-compiling or building natively on aarch64. Affected files:replacement_scan/mod.rs,types/logical_type.rs,vector/writer.rs.
Changed
- DuckDB v1.5.1 compatibility — updated
DUCKDB_API_VERSIONdoc comment and version range documentation to explicitly cover v1.5.1. The C API version remains"v1.2.0"(unchanged from v1.5.0). Users are strongly recommended to upgrade their DuckDB runtime to v1.5.1 for critical WAL corruption and ART index correctness fixes.
Internal
-
CI action update —
dtolnay/rust-toolchainpinned to631a55b12751854ce901bb631d5902ceb48146f7(PR #59). -
Mutation testing —
mutants.tomlnow setsfeatures = ["duckdb-1-5"]so thatcargo mutantscompiles and tests feature-gated code paths. Previously, four mutants inMockRegistrar::copy_function_names,has_copy_function, andtotal_registrationswere unreachable because their tests were also feature-gated. Addedmock_registrar_total_registrations_scalar_plus_copy_functionto robustly kill the+ with -mutation intotal_registrationsby using a non-zerobasecount.
0.7.0 — 2026-03-22
Added
-
duckdb-1-5feature modules — theduckdb-1-5feature flag is no longer a placeholder. When enabled, it gates five new modules wrapping DuckDB 1.5.0 C Extension API additions:catalog— catalog entry lookup (CatalogEntry,Catalog,CatalogEntryType)client_context— client context access (ClientContext) for retrieving catalogs, config options, and connection IDs from within registered function callbacksconfig_option— extension-defined configuration options (ConfigOptionBuilder,ConfigOptionScope) registered viaSET/RESET/current_setting()copy_function— customCOPY TOhandlers (CopyFunctionBuilder) with bind → global init → sink → finalize lifecycletable_description— table metadata queries (TableDescription) for column count, names, and logical types
-
TypeId::TimeNs— newTIME_NScolumn type variant for nanosecond- precision time of day (DuckDB 1.5.0+, requiresduckdb-1-5feature) -
ScalarFunctionBuilder::varargs()/varargs_logical()— mark a scalar function as accepting variadic arguments (requiresduckdb-1-5) -
ScalarFunctionBuilder::volatile()— mark a scalar function as volatile (re-evaluated for every row even with constant arguments, requiresduckdb-1-5) -
ScalarFunctionBuilder::bind()— set a bind callback invoked once during query planning for per-query state allocation (requiresduckdb-1-5) -
ScalarFunctionBuilder::init()— set an init callback invoked once per thread for per-thread local state allocation (requiresduckdb-1-5)
Changed
-
DuckDB 1.5.0 support — upgraded default
libduckdb-sysfrom 1.4.4 to 1.10500.0 (DuckDB 1.5.0) andduckdbfrom 1.4.4 to 1.10500.0. The version range">=1.4.4, <2"inCargo.tomlis unchanged, preserving backward compatibility with DuckDB 1.4.x. -
Transitive dependency updates —
cc1.2.56→1.2.57,tar0.4.44→0.4.45,rustls-webpki0.103.9→0.103.10,arrow56.2.0→57.3.0,clap4.5.60→4.6.0,tempfile3.14.0→3.27.0, plus ~30 other minor/patch updates. -
CI action updates —
Swatinem/rust-cachev2.8.2→v2.9.1,actions/download-artifactv8.0.0→v8.0.1,actions/cache5.0.3→5.0.4,codecov/codecov-action5.4.3→5.5.3.
Fixed
- COPY format handlers — previously listed as a known limitation (no C API
counterpart). DuckDB 1.5.0 adds
duckdb_create_copy_functionand related symbols; the newcopy_functionmodule wraps them behindduckdb-1-5.
0.6.0 — 2026-03-12
Added
-
InMemoryDbdispatch table initialisation —InMemoryDb::open()now correctly initialises theloadable-extensiondispatch table from bundled DuckDB symbols before opening a connection, allowing all threeInMemoryDbunit tests to pass undercargo test --features bundled-test. Previously every call toInMemoryDb::open()panicked with"DuckDB API not initialized or DuckDB feature omitted"because theloadable-extensiondispatch table was never populated incargo test. -
src/testing/bundled_api_init.cpp— thin C++ shim that wraps DuckDB's internalCreateAPIv1()function (fromduckdb/main/capi/extension_api.hpp) as a C-linkage symbol (quack_rs_create_api_v1). Called once at test startup to populate all 459AtomicPtrslots in the dispatch table with real bundled DuckDB function pointers. -
build.rs— Cargo build script that, when thebundled-testfeature is active, locates thelibduckdb-sysbuild output directory, finds the bundled DuckDB include path, and compilesbundled_api_init.cppvia thecccrate. -
CI:
test-bundledjob — new CI job runscargo test --all-targets --features bundled-teston all three platforms (Linux, macOS, Windows) on every push and pull request, closing the gap that allowed this failure to reach the release workflow undetected. -
Pitfall P9 documented —
LESSONS.mdnow contains a full analysis of theloadable-extensiondispatch table failure mode: root cause, theCreateAPIv1()solution, ABI compatibility details, risks of relying on DuckDB's internal C++ header, and a mitigation table.
Fixed
InMemoryDb::open()no longer panics when called incargo testwith thebundled-testfeature enabled. This was a regression introduced whenInMemoryDbwas first shipped in 0.5.1 without the dispatch table initialisation step.
Changed
bundled-testfeature documentation updated to accurately describe the dispatch table initialisation behaviour (previously claimed to "bypass" the dispatch mechanism; it now correctly initialises it).
0.5.1 — 2026-03-12
Added
-
Testing primitives (
quack_rs::testing) — new mock types for unit-testing extension logic without a live DuckDB process:MockVectorWriter— in-memory output buffer matching theVectorWriterAPI; use to test scalar/aggregate finalize/scan callbacksMockVectorReader— in-memory input buffer with convenience constructors (from_i64s,from_strs,from_bools,from_f64s,from_i32s)MockDuckValue— typed enum covering all DuckDB scalar typesMockRegistrar— implements theRegistrartrait using interior mutability; records registered functions without any C API callCastRecord— records source/target types for cast registrations
-
bundled-testCargo feature — links the bundled DuckDB static library via theduckdbcrate and enablesInMemoryDb::open()for SQL-level assertions incargo test. Does not initialize theloadable-extensiondispatch table. -
InMemoryDb— wrapsduckdb::Connectionfor SQL-level integration tests; available behind thebundled-testfeature. -
Builder introspection accessors —
pub fn name(&self) -> &stradded toScalarFunctionBuilder,ScalarFunctionSetBuilder,AggregateFunctionBuilder,AggregateFunctionSetBuilder, andTableFunctionBuilder.pub fn source(&self) -> Option<TypeId>andpub fn target(&self) -> Option<TypeId>added toCastFunctionBuilder.
Security
- Bump
quinn-proto0.11.13 → 0.11.14 in root andexamples/hello-extCargo.lockfiles (addresses RUSTSEC advisory).
0.5.0 — 2026-03-10
Added
-
param_logical(LogicalType)on all builders — register parameters with complex parameterized types (LIST(BIGINT),MAP(VARCHAR, INTEGER),STRUCT(...)) thatTypeIdalone cannot express. Available onAggregateFunctionBuilder,AggregateFunctionSetBuilder::OverloadBuilder,ScalarFunctionBuilder, andScalarOverloadBuilder. Parameters added viaparam()andparam_logical()are interleaved by position, so the order you call them is the order DuckDB sees them. -
returns_logical(LogicalType)on all builders — set a complex parameterized return type. When bothreturns(TypeId)andreturns_logical(LogicalType)are called, the logical type takes precedence. Available onAggregateFunctionBuilder,AggregateFunctionSetBuilder,ScalarFunctionBuilder, andScalarOverloadBuilder. This eliminates the need for raw FFI when returningLIST(BOOLEAN),LIST(TIMESTAMP),MAP(K, V), or any other parameterized type. -
null_handling(NullHandling)on set overload builders — per-overload NULL handling configuration forAggregateFunctionSetBuilder::OverloadBuilderandScalarOverloadBuilder. Previously only available on single-function builders.
Notes
- Upstream fix:
duckdb-loadable-macrospanic-at-FFI-boundary — the safe entry-point pattern developed inquack-rs(using?/ok_or_elsethroughout instead of.unwrap()) was contributed upstream as duckdb/duckdb-rs#696 and merged 2026-03-09. All users of theduckdb_entrypoint_c_api!macro fromduckdb-loadable-macroswill receive this fix in the nextduckdb-rsrelease.quack-rsusers have always been protected via the safeentry_point!/entry_point_v2!macros provided by this crate.
0.4.0 — 2026-03-09
Added
-
ConnectionandRegistrartrait — version-agnostic extension registration facade (src/connection.rs).Connectionwraps theduckdb_connectionandduckdb_databasehandles provided at initialization time. TheRegistrartrait provides uniform methods for registering all extension components (scalar, scalar set, aggregate, aggregate set, table, SQL macro, cast), making registration code interchangeable across DuckDB 1.4.x and 1.5.x. Replacement scans are exposed as direct methods onConnectionsince they requireduckdb_database, not the connection handle. -
init_extension_v2— new entry point helper that passes&Connectionto the registration callback instead of a rawduckdb_connection. Prefer this overinit_extensionfor new extensions. -
entry_point_v2!macro — companion macro toentry_point!that generates the#[no_mangle] unsafe extern "C"entry point usinginit_extension_v2. -
duckdb-1-5cargo feature — placeholder feature flag for DuckDB 1.5.0-specific C API wrappers. Currently empty; will be populated whenlibduckdb-sys1.5.0 is published on crates.io.
Changed
- DuckDB version support broadened to 1.4.x and 1.5.x — the
libduckdb-sysdependency requirement was relaxed from an exact pin (=1.4.4) to a range (>=1.4.4, <2). DuckDB v1.5.0 (released 2026-03-09) does not change the C API version string (v1.2.0) used induckdb_rs_extension_api_init; the existingDUCKDB_API_VERSIONconstant remains correct for both releases. Extension authors can now pin their ownlibduckdb-systo either=1.4.4or=1.5.0and resolve cleanly againstquack-rs. The scaffold template and CI workflow template were updated to default to DuckDB v1.5.0.
0.3.0 — 2026-03-08
Added
-
TableFunctionBuilder— type-safe builder for registering DuckDB table functions (theSELECT * FROM my_function(args)pattern). Covers the full bind/init/scan lifecycle with ergonomic callbacks, eliminating ~100 lines of raw FFI boilerplate. Helper typesBindInfo,FfiBindData<T>, andFfiInitData<T>manage parameter extraction and per-scan state with zero raw pointer manipulation. Seetableandexamples/hello-ext(generate_series_ext) for a fully-tested end-to-end example verified against DuckDB 1.4.4. -
ReplacementScanBuilder— builder for registering DuckDB replacement scans (theSELECT * FROM 'file.xyz'pattern where a file path triggers a table-valued scan). The builder handles callback registration, path extraction, and bind-info population through a 4-method chain. Seereplacement_scan. -
StructVector— safe wrapper for reading and writing STRUCT child vectors.get_child(vec, idx),field_reader(vec, idx, row_count), andfield_writer(vec, idx)replace manual offset arithmetic over child vector handles. -
ListVector— safe wrapper for reading and writing LIST child vectors.get_child,get_entry,set_entry,reserve,set_size,child_reader, andchild_writercover the complete LIST read/write workflow without raw pointer casts. -
MapVector— safe wrapper for DuckDB MAP vectors (stored asLIST<STRUCT{key, value}>).keys(vec),values(vec),struct_child(vec),reserve,set_size,set_entry, andget_entryexpose the full MAP interface. -
vector::complexmodule — re-exportsStructVector,ListVector,MapVectoratquack_rs::vector::complexand documents the read-vs-write workflow for nested types with working code examples in the module doc. -
preludeadditions —TableFunctionBuilder,BindInfo,FfiBindData,FfiInitData,ReplacementScanBuilder,StructVector,ListVector,MapVector,CastFunctionBuilder,CastFunctionInfo,CastModeare now all re-exported fromquack_rs::prelude. -
CastFunctionBuilder— type-safe builder for registering custom type cast functions viaduckdb_cast_function_*. Covers both explicitCAST(x AS T)and implicit coercions (with optionalimplicit_cost). The companionCastFunctionInfowrapper exposescast_mode(),set_error(), andset_row_error()inside callbacks, giving correctTRY_CAST/CASTerror handling with zero raw pointer boilerplate. Seecastfor the full API. -
DbConfig— RAII wrapper forduckdb_config(extension configuration parameters). Builder-style.set(name, value)?chain, automaticduckdb_destroy_configon drop, andflag_count()/get_flag(index)for enumerating all available options. Useful when an extension needs to open a secondaryDuckDBdatabase from within its callbacks. Seeconfig. -
ScalarFunctionSetBuilder— builder for registering scalar function sets (multiple overloads under one name), mirroringAggregateFunctionSetBuilder. -
TypeIdvariants —Decimal,Struct,Map,UHugeInt,TimeTz,TimestampS,TimestampMs,TimestampNs,Array,Enum,Union,Bit. -
From<TypeId> for LogicalType— idiomatic conversion fromTypeId. -
#[must_use]on builder structs —ScalarFunctionBuilder,AggregateFunctionBuilder,AggregateFunctionSetBuilder, andOverloadBuildernow warn at compile time if constructed but never consumed. -
NullHandlingenum and.null_handling()builder method — configurable NULL propagation for scalar and aggregate functions viaduckdb_scalar_function_set_special_handling/duckdb_aggregate_function_set_special_handling. -
VectorWriter::write_interval— writes INTERVAL values to output vectors using the correct 16-byte{ months: i32, days: i32, micros: i64 }layout. -
append_metadatabinary — native Rust replacement for the Pythonappend_extension_metadata.pyscript, now shipping with the crate. Install withcargo install quack-rs --bin append_metadata. -
hello-extcast function demo —examples/hello-extnow registers aCAST(VARCHAR AS INTEGER)cast function usingCastFunctionBuilder, demonstrating bothCAST(abort-on-error) andTRY_CAST(NULL-on-error) code paths. Five unit tests coverparse_varchar_to_int, including boundary values and overflow.
Not implemented (upstream C API gap)
-
Window functions —
duckdb_create_window_functionand related symbols do not exist in DuckDB's public C extension API. They are implemented only in the C++ layer and are therefore not wrappable byquack-rsor any other C-API binding. Verified against the DuckDB stable C API reference andlibduckdb-sys1.4.4 bindings. -
COPY format handlers —
duckdb_create_copy_functionand related symbols are similarly absent from the C extension API for the same reason.
Fixed
hello-extgs_bindcallback — replaced incorrectduckdb_value_int64(param)(wrong arity: takes 3 arguments) withduckdb_get_int64(param)(correct 1-argument form). The extension now builds cleanly and all 11 live SQL tests pass against DuckDB 1.4.4.
Changed
- Bump
criteriondev-dependency from0.5to0.8. - Bump
Swatinem/rust-cacheGitHub Action fromv2.7.5tov2.8.2. - Bump
dtolnay/rust-toolchainCI pin fromv2.7.5to latest SHA. - Bump
actions/attest-build-provenancefromv2tov4. - Bump
actions/configure-pagesto latest SHA (d5606572…). - Bump
actions/upload-pages-artifactfromv3.0.1tov4.0.0.
0.2.0 — 2026-03-07
Added
-
validate::description_ymlmodule — parse and validate a completedescription.ymlmetadata file end-to-end. Includes:DescriptionYmlstruct — structured representation of all required and optional fieldsparse_description_yml(content: &str)— parse and validate in one stepvalidate_description_yml_str(content: &str)— pass/fail validationvalidate_rust_extension(desc: &DescriptionYml)— enforce Rust-specific fields (language: Rust,build: cargo,requires_toolchainsincludesrust)- 25+ unit tests covering all required fields, optional fields, error paths, and edge cases
-
preludemodule — ergonomic glob-import for the most commonly used items.use quack_rs::prelude::*;brings in all builder types, state traits, vector helpers, types, error handling, and the API version constant. Reduces boilerplate for extension authors. -
Scaffold:
extension_config.cmakegeneration — the scaffold generator now producesextension_config.cmake, which is referenced by theEXT_CONFIGvariable in the Makefile and required byextension-ci-toolsfor CI integration. -
Scaffold: SQLLogicTest skeleton —
generate_scaffoldnow producestest/sql/{name}.test, a ready-to-fill SQLLogicTest file withrequiredirective, format comments, and example query/result blocks. E2E tests are required for community extension submission (Pitfall P3). -
Scaffold: GitHub Actions CI workflow —
generate_scaffoldnow produces.github/workflows/extension-ci.yml, a complete cross-platform CI workflow that builds and tests the extension on Linux, macOS, and Windows against a real DuckDB binary. -
validate::validate_excluded_platforms_str— validates theexcluded_platformsfield fromdescription.ymlas a semicolon-delimited string (e.g.,"wasm_mvp;wasm_eh;wasm_threads"). Splits on;and validates each token. An empty string is valid (no exclusions). -
validate::validate_excluded_platforms— re-exported at thevalidatemodule level (previously only accessible asvalidate::platform::validate_excluded_platforms). -
validate::semver::classify_extension_version— returnsExtensionStability(Unstable/PreRelease/Stable) classifying the tier a version falls into. -
validate::semver::ExtensionStability— enum for DuckDB extension version stability tiers (Unstable,PreRelease,Stable) withDisplayimplementation. -
scalarmodule —ScalarFunctionBuilderfor registering scalar functions with the DuckDB C Extension API. Includestry_newwith name validation,param,returns,functionsetters, andregister. Full unit tests included. -
entry_point!macro — generates the required#[no_mangle] extern "C"entry point with zero boilerplate from an identifier and registration closure. -
VectorWriter::write_varchar— writes VARCHAR string values to output vectors usingduckdb_vector_assign_string_element_len(handles both inline and pointer formats). -
VectorWriter::write_bool— writes BOOLEAN values as a single byte. -
VectorWriter::write_u16— writes USMALLINT values. -
VectorWriter::write_i16— writes SMALLINT values. -
VectorReader::read_interval— reads INTERVAL values from input vectors via the correct 16-byte layout helper. -
CI: Windows testing — the CI matrix now includes
windows-latestin thetestjob, covering all three major platforms (Linux, macOS, Windows). -
CI:
example-checkjob — CI now checks, lints, and testsexamples/hello-extas part of every PR, ensuring the example extension always compiles and its tests pass. -
validate::validate_release_profile— checks Cargo release profile settings for loadable-extension correctness. Validatespanic,lto,opt-level, andcodegen-units.
Fixed
- MSRV documentation now consistently states 1.84.1 across
README.md,CONTRIBUTING.md, andCargo.toml(previouslyREADME.mdstated 1.80).
0.1.0 — 2025-05-01
Added
- Initial release
entry_pointmodule:init_extensionhelper for correct extension initializationaggregatemodule:AggregateFunctionBuilder,AggregateFunctionSetBuilderaggregate::statemodule:AggregateStatetrait,FfiState<T>wrapperaggregate::callbacksmodule: type aliases for all 6 callback signaturesvectormodule:VectorReader,VectorWriter,ValidityBitmap,DuckStringViewtypesmodule:TypeIdenum,LogicalTypeRAII wrapperintervalmodule:DuckInterval,interval_to_micros,read_interval_aterrormodule:ExtensionError,ExtResult<T>testingmodule:AggregateTestHarness<S>for pure-Rust aggregate testingvalidatemodule:validate_extension_name,validate_function_name,validate_semver,validate_extension_version,validate_spdx_license,validate_platform,validate_release_profilescaffoldmodule:generate_scaffoldfor generating complete extension projectssql_macromodule:SqlMacrofor registering SQL macros without FFI callbacks- Complete
hello-extexample extension - Documentation of all 15 DuckDB Rust FFI pitfalls (
LESSONS.md) - CI pipeline: check, test, clippy, fmt, doc, MSRV, bench-compile
SECURITY.mdvulnerability disclosure policy
FAQ
Frequently asked questions about quack-rs and building DuckDB extensions in Rust.
General
What is quack-rs?
quack-rs is a Rust SDK for building DuckDB loadable extensions using DuckDB's
pure C Extension API. It provides safe, ergonomic builders for registering
scalar functions, aggregate functions, table functions, cast functions,
replacement scans, SQL macros, and copy functions (via the duckdb-1-5
feature), along with helpers for reading and writing DuckDB vectors, and
utilities for publishing community extensions.
Why does this exist?
Building a DuckDB extension in Rust means solving a set of undocumented FFI problems that each developer otherwise discovers independently. quack-rs documents all 31 known pitfalls, and its API prevents most of them. See the Pitfall Catalog.
What DuckDB version does quack-rs target?
quack-rs requires libduckdb-sys = ">=1.4.4, <2" and supports DuckDB 1.4.x
and 1.5.x. CI loads a built extension into DuckDB v1.4.4, v1.5.0, v1.5.5 and the
latest release.
The C API version passed to the dispatch-table initializer is "v1.2.0",
available as quack_rs::DUCKDB_API_VERSION. Every DuckDB 1.4.x and 1.5.x
release loads extensions built for it (1.5.6 declares C API v1.5.6, but still
accepts v1.2.0). The C API version is a separate identifier from the DuckDB
release and from the libduckdb-sys crate version (e.g. 1.10505.0 for
DuckDB 1.5.5).
What is the minimum supported Rust version (MSRV)?
Rust 1.86.0 or later. This is enforced in Cargo.toml with
rust-version = "1.86.0".
Is quack-rs production-ready?
It is pre-1.0, and you should judge it against your own requirements rather than take a yes. What the record shows:
- The API still changes. Minor releases before 1.0 can break it; each one lists its breaking changes, with migration notes, in the changelog.
- Audits keep finding real defects. The review released as 0.16.0 fixed 24
defects in earlier releases, including two heap-corruption paths. The 0.18.0 audits fixed further soundness holes,
process aborts and wrong answers, and found one serious defect in unreleased
code before it was published: aggregate functions returned wrong results in
release builds with Cargo's default profile.
AUDIT.mdin the repository records each audit, what it found, and whether each fix was reproduced against a real DuckDB or derived from DuckDB's source. - Some limits are DuckDB's. The C API has defects quack-rs can only document or work around; see Known Limitations.
It was extracted from duckdb-behavioral, a DuckDB community extension, where the first 16 of the pitfalls it now documents were discovered. If you ship an extension built on it, run end-to-end tests that load the extension into each DuckDB release you support (see the Testing Guide).
Functions
Can I expose SQL macros as an extension?
Yes, without any C++ wrapper code. Use quack_rs::sql_macro::SqlMacro:
#![allow(unused)] fn main() { use libduckdb_sys::duckdb_connection; fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> { use quack_rs::sql_macro::SqlMacro; // Scalar macro let m = SqlMacro::scalar("double_it", &["x"], "x * 2")?; unsafe { m.register(con) }?; // Table macro let m = SqlMacro::table("recent_events", &["n"], "SELECT * FROM events ORDER BY ts DESC LIMIT n")?; unsafe { m.register(con) }?; Ok(()) } }
Register them in your registration closure (the one passed to entry_point!
or init_extension) alongside your other functions. A table macro's body is
bound when it is created, so the
events table must already exist or register returns an error.
See SQL Macros.
Can I register multiple overloads of the same function?
Yes, using AggregateFunctionSetBuilder (for aggregates) or
ScalarFunctionSetBuilder (for scalars). Both support complex parameter types
via param_logical(LogicalType) and complex return types via
returns_logical(LogicalType), and in both, each overload may return a
different type — DuckDB resolves an overload from its parameter types and
arity alone. See
Overloading with Function Sets.
Can I register multiple functions in one extension?
Yes. The registration closure receives a duckdb_connection and can register
as many functions as needed:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_connection, duckdb_extension_access, duckdb_extension_info}; use quack_rs::error::ExtensionError; use quack_rs::sql_macro::SqlMacro; use quack_rs::DUCKDB_API_VERSION; unsafe fn register_word_count(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } unsafe fn register_sentence_count(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } unsafe fn demo(info: duckdb_extension_info, access: *const duckdb_extension_access) -> bool { quack_rs::entry_point::init_extension(info, access, DUCKDB_API_VERSION, |con| { unsafe { register_word_count(con) }?; unsafe { register_sentence_count(con) }?; unsafe { SqlMacro::scalar("double_it", &["x"], "x * 2")? .register(con)?; } Ok(()) }) } }
Can I use the duckdb crate instead of libduckdb-sys?
No. The duckdb crate's bundled feature embeds its own copy of DuckDB. A
loadable extension must link against the DuckDB that loads it, not bundle a
separate copy. Use libduckdb-sys with the loadable-extension feature.
Can I have a scalar function with no parameters?
Yes. Just do not call param:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::prelude::*; unsafe extern "C" fn quack_callback(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {} fn demo(con: duckdb_connection) -> Result<(), ExtensionError> { unsafe { ScalarFunctionBuilder::new("current_quack") .returns(TypeId::Varchar) .function(quack_callback) .register(con)?; } Ok(()) } }
Testing
Do I need a DuckDB instance to run unit tests?
No. AggregateTestHarness simulates the aggregate lifecycle in pure Rust
without any DuckDB dependency. You can run cargo test without loading a DuckDB
binary.
My unit tests all pass but the extension crashes. Why?
Unit tests cannot detect FFI wiring bugs. See Pitfall P3 and the Testing Guide. Always run end-to-end tests that load the packaged extension into a real DuckDB process.
How do I test SQL macros?
SqlMacro::to_sql() is pure Rust and requires no DuckDB connection:
#![allow(unused)] fn main() { use quack_rs::sql_macro::SqlMacro; let m = SqlMacro::scalar("triple", &["x"], "x * 3").unwrap(); assert_eq!(m.to_sql(), r#"CREATE OR REPLACE MACRO "triple"("x") AS (x * 3)"#); }
For an end-to-end test, call the macro from your SQLLogicTest file:
query I
SELECT double_it(21);
----
42
Publishing
How do I publish to the DuckDB community extensions registry?
- Scaffold your project with
generate_scaffold - Push to GitHub
- Submit a pull request to the
community-extensions repo
with your
description.yml
See Community Extensions for the full workflow.
My extension name is taken. What should I do?
Use a vendor-prefixed name: myorg_analytics instead of analytics. Extension
names must be unique among DuckDB community extensions; check
community-extensions.duckdb.org
before you pick one.
Do I need to set up CI manually?
No. generate_scaffold produces .github/workflows/extension-ci.yml which
builds and tests your extension on Linux, macOS, and Windows automatically.
Can my extension be installed with INSTALL ... FROM community?
Yes, once your pull request is merged into the community-extensions repository.
Until then, users load the .duckdb_extension file directly, starting DuckDB with
-unsigned because a local build is not signed:
LOAD './path/to/my_extension.duckdb_extension';
Troubleshooting
My aggregate returns wrong results with no error.
The most common cause is Pitfall L1: your combine callback is not propagating
all configuration fields. See
Pitfall L1
and test with AggregateTestHarness::combine.
The NULLs my function writes come back as values.
You are likely calling duckdb_vector_get_validity without first calling
duckdb_vector_ensure_validity_writable. A vector that has never held a NULL
has no validity mask, so duckdb_vector_get_validity returns a null pointer and
duckdb_validity_set_row_invalid silently does nothing (dereferencing that
pointer yourself crashes). Use VectorWriter::set_null instead. See
Pitfall L4.
My function is not found in SQL after LOAD.
If LOAD succeeded, the entry point ran, so check the name you gave the
builder. If you build a function set through the raw C API instead of
quack-rs's set builders, every member needs its own name, or DuckDB silently
skips it (Pitfall L6).
If LOAD itself fails, check the entry point symbol. DuckDB takes the file's
base name, lowercases it and calls <base name>_init_c_api, so
my_extension.duckdb_extension needs my_extension_init_c_api
(Pitfall P1).
make configure fails with a missing file error.
The extension-ci-tools submodule is missing. In a new project, such as one
fresh from the scaffold, add it once (git submodule update --init does nothing
until it has been added):
git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools
In a clone of a repository that already has the submodule:
git submodule update --init --recursive
See Pitfall P4.
My SQLLogicTest fails in CI but passes locally.
SQLLogicTest does exact string matching. The most common issue is a difference in NULL representation, decimal places, or line endings. Run the query in the same DuckDB version used by CI and copy the output verbatim.
How do I read a VARCHAR that is longer than 12 bytes?
VectorReader::read_str handles both the inline (≤ 12 bytes) and pointer
(> 12 bytes) formats automatically. No special handling needed.
What happens if I read from a NULL row?
You get garbage data from the vector's data buffer — and for VARCHAR or
BLOB, possibly a stale pointer into freed memory. Always check is_valid
before reading. See NULL Handling & Strings.
Architecture
Why use libduckdb-sys with loadable-extension instead of the duckdb crate?
The duckdb crate is designed for embedding DuckDB, not for extending it. Its
bundled feature includes a statically linked DuckDB binary, which conflicts
with the DuckDB runtime that loads your extension. libduckdb-sys with
loadable-extension provides lazy-initialized function pointers that are
populated by DuckDB at extension load time.
Why not use duckdb-loadable-macros?
duckdb-loadable-macros relies on extract_raw_connection which uses the
internal Rc<RefCell<InnerConnection>> layout. This is fragile and causes
SEGFAULTs when the layout changes between duckdb crate versions.
init_extension uses the correct C API entry sequence directly.
Why must the release profile use panic = "unwind"?
quack-rs wraps every entry point and every callback it generates in
catch_unwind and reports a panic as an ordinary SQL error (a raw
extern "C" callback you write yourself needs a *_callback! macro or
catch_ffi_panic; see Installation).
catch_unwind cannot catch anything under
panic = "abort": the process terminates at the panic site, taking the user's
DuckDB session with it. validate_release_profile rejects abort for this
reason. Returning Result and using ? is still the right style — the guards
are a safety net, not a substitute.
Can I use async Rust in my extension?
Not directly in FFI callbacks. DuckDB's callbacks are synchronous C functions.
You can run an async runtime such as Tokio and block on async tasks inside
callbacks (for example with Runtime::block_on), but the callbacks themselves
must return synchronously.
How does FfiState<T> prevent double-free?
Each state slot starts with a tag that init_callback writes once the T
is in place. destroy_callback drops the T (freeing its box, when T is
too large to store inline) only when the tag matches, and clears the tag
first. A second call to destroy_callback on the same state finds no tag and
does nothing.
Contributing
This page describes how to build, test and contribute to quack-rs: the toolchain, the quality gates every pull request must pass, the test strategy, code standards, the repository layout and the release policy. Bug reports, documentation fixes, newly discovered pitfalls and code are all welcome.
Development prerequisites
| Tool | Version | Purpose |
|---|---|---|
| Rust | ≥ 1.86.0 (MSRV) | Compiler |
rustfmt | stable | Formatting |
clippy | stable | Linting |
cargo-msrv | latest | MSRV verification |
Install the Rust toolchain via rustup.rs.
Building
# Build the library
cargo build
# Build in release mode (enables LTO + strip)
cargo build --release
# Build the hello-ext example extension
cargo build --release --manifest-path examples/hello-ext/Cargo.toml
Quality gates
All of the following must pass before merging any pull request:
# Tests — zero failures, zero ignored
cargo test
# Integration tests
cargo test --test integration_test
# Doctests — `--all-targets` does not include them
cargo test --doc --features duckdb-1-5-4
# Linting — zero warnings (warnings are errors)
cargo clippy --all-targets -- -D warnings
# Formatting
cargo fmt -- --check
# Documentation — zero broken links or missing docs
RUSTDOCFLAGS="-D warnings" cargo doc --no-deps
# MSRV — must compile on Rust 1.86.0 (excludes benches; matches CI)
cargo +1.86.0 check
These same checks run in CI on every push and pull request.
Test strategy
Unit tests
Unit tests live in #[cfg(test)] modules alongside the code. They test
pure-Rust logic that does not require a live DuckDB instance.
Important constraint: libduckdb-sys with features = ["loadable-extension"]
makes all DuckDB C API functions go through lazy AtomicPtr dispatch. These
pointers are only populated when duckdb_rs_extension_api_init is called from
within a real DuckDB extension load — or by testing::InMemoryDb::open() under
the bundled-test / bundled-test-prebuilt features. Without one of those,
calling any duckdb_* function in a unit test panics ("DuckDB API not
initialized or DuckDB feature omitted"). Put such tests in tests/ffi_roundtrip.rs
or its submodules under tests/ffi_roundtrip/, which open an InMemoryDb.
Integration tests
tests/integration_test.rs contains pure-Rust tests that cross module
boundaries — testing interval with AggregateTestHarness, verifying FfiState
lifecycle, and so on. These still cannot call duckdb_* functions.
Property-based tests
Selected modules include proptest-based tests:
interval.rs— overflow edge cases across the fulli32/i64rangetesting/harness.rs— sum associativity, identity element forAggregateState
Example-extension tests
examples/hello-ext/ contains #[cfg(test)] unit tests for its pure-Rust logic.
CI also tests it end to end: it builds the cdylib, appends the metadata footer
with append_metadata, and loads it into DuckDB 1.4.4, 1.5.0, 1.5.5 and the latest
release. See CONTRIBUTING.md for the exact commands.
Code standards
Safety documentation
Every unsafe block must have a // SAFETY: comment explaining:
- Which invariant the caller guarantees
- Why the operation is valid given that invariant
clippy::undocumented_unsafe_blocks enforces this in library code (CI treats
it as an error); test code is exempt.
#![allow(unused)] fn main() { struct Ffi { inner: *mut u64 } let ffi = Ffi { inner: Box::into_raw(Box::new(0_u64)) }; // SAFETY: `ffi.inner` came from `Box::into_raw` and has not been freed; // nothing else holds it, so reclaiming and dropping the box is sound. unsafe { drop(Box::from_raw(ffi.inner)) }; }
No panics across FFI
A panic must never escape a function DuckDB calls (callbacks and entry
points): every callback kind runs under catch_unwind. Beyond that, library code
uses Option/Result and ?, and panics only where its # Panics section says
so — VectorWriter::write_varchar on a string over 4 GiB, or a builder's
new(name) on an interior NUL, which has try_new(name) beside it.
Clippy lint policy
The crate enables the all, pedantic, nursery and cargo lint groups, plus
undocumented_unsafe_blocks. All warnings are treated as errors in CI. Lints are
suppressed only where they produce false positives for SDK API patterns:
[lints.clippy]
module_name_repetitions = "allow" # e.g., AggregateFunctionBuilder
must_use_candidate = "allow" # builder methods
missing_errors_doc = "allow" # unsafe extern "C" callbacks
return_self_not_must_use = "allow" # builder pattern
Documentation
Every public item must have a doc comment. Follow these conventions:
- First line: a one-sentence summary, ending with a period
# Safety: mandatory on everyunsafe fn# Panics: mandatory if the function can panic# Errors: mandatory on functions returningResult# Example: encouraged on public types and key methods
Repository structure
quack-rs/
├── src/
│ ├── abi.rs # `DuckDB` C Extension API ABI compatibility checking
│ ├── appender.rs # Bulk data appending
│ ├── arrow.rs # Arrow C Data Interface bridge (`duckdb-1-5-4` feature; floor set by the `libduckdb-sys` 1.10504.0 bindings)
│ ├── callback.rs # Panic-safe callback wrapper macros for `DuckDB` extension callbacks
│ ├── catalog.rs # Catalog entry lookup (`DuckDB` 1.5.0+)
│ ├── chunk_writer.rs # Auto-sizing chunk writer for table function scan callbacks
│ ├── client_context.rs # Client context access (`DuckDB` 1.5.0+)
│ ├── config.rs # RAII wrapper for `DuckDB` database configuration
│ ├── config_option.rs # Extension-defined configuration options (`DuckDB` 1.5.0+)
│ ├── connection.rs # [`Connection`] — version-agnostic extension registration facade
│ ├── data_chunk.rs # Ergonomic wrapper around `DuckDB` data chunks
│ ├── debug_repr.rs # Internal helpers for the crate's `Debug` implementations
│ ├── entry_point.rs # Extension entry point helper
│ ├── error.rs # Error types for `DuckDB` extension FFI error propagation
│ ├── error_data.rs # Structured error data (`DuckDB` 1.5.0+)
│ ├── expression.rs # Bound expressions (`DuckDB` 1.5.0+)
│ ├── extra_info.rs # Ownership of a function's `extra_info` allocation until `DuckDB` takes it
│ ├── file_system.rs # File system access (`DuckDB` 1.5.0+)
│ ├── instance_cache.rs # Database instance cache (`DuckDB` 1.5.0+)
│ ├── interval.rs # `DuckDB` `INTERVAL` type conversion utilities
│ ├── lib.rs # Crate root: module declarations and crate-level documentation
│ ├── prelude.rs # Convenience re-exports for the most commonly used `quack-rs` items
│ ├── query.rs # Running SQL from inside an extension
│ ├── secrets.rs # Credential handling for extensions
│ ├── selection_vector.rs # Selection vectors (`DuckDB` 1.5.0+)
│ ├── sql_macro.rs # SQL macro registration for `DuckDB` extensions
│ ├── table_description.rs # Table description metadata
│ ├── tls.rs # Type-erased TLS configuration provider for HTTP-capable extensions
│ ├── value.rs # RAII wrapper around `DuckDB` values (`duckdb_value`)
│ ├── warning.rs # Structured security warning API for extensions
│ ├── aggregate/
│ │ ├── callbacks.rs # Type aliases for the five required `DuckDB` aggregate callback signatures
│ │ ├── info.rs # Ergonomic wrapper around `duckdb_function_info` for aggregate function callbacks
│ │ ├── mod.rs # Builders for registering `DuckDB` aggregate functions
│ │ ├── state.rs # Generic `FfiState<T>` wrapper for safe aggregate state management
│ │ └── builder/
│ │ ├── mod.rs # Builder types for registering `DuckDB` aggregate functions
│ │ ├── overload.rs # One overload within an [`AggregateFunctionSetBuilder`]
│ │ ├── set.rs # Builder for registering a `DuckDB` aggregate function set (multiple overloads)
│ │ ├── single/
│ │ │ └── register.rs # `AggregateFunctionBuilder::register`
│ │ ├── single.rs # Builder for registering a single-signature `DuckDB` aggregate function
│ │ └── tests.rs # Unit tests
│ ├── appender/
│ │ ├── chunk.rs # Chunk-at-a-time appends: handing the [`Appender`] a whole [`DataChunk`]
│ │ ├── construct.rs # Creating an [`Appender`] and choosing the columns it appends to
│ │ ├── lifecycle.rs # Flushing, closing and (with `duckdb-1-5`) clearing an [`Appender`]
│ │ ├── rows.rs # Row-at-a-time appends: `row`, `end_row` and the non-numeric `append_*` methods
│ │ └── scalars.rs # The fixed-width numeric `append_*` methods, `append_bool` through `append_u128`
│ ├── arrow/
│ │ ├── array.rs # `ArrowArray` — an owned Arrow C Data Interface array
│ │ ├── convert.rs # The four conversions between `DuckDB` data chunks and the Arrow C Data Interface
│ │ ├── converted.rs # `ArrowConvertedSchema` — an Arrow schema translated into `DuckDB`'s own type descriptors
│ │ ├── export_check.rs # Values `DuckDB` would export as different values, found before the export
│ │ ├── import_check.rs # Structural checks on an imported array, and flattening what the readers cannot index
│ │ ├── import_layout/
│ │ │ └── tests.rs # Unit tests
│ │ ├── import_layout.rs # Arrow layouts `DuckDB` imports wrongly, found by walking the array with its schema
│ │ ├── options.rs # `ArrowOptions` — the Arrow production settings of a connection or a result
│ │ ├── schema.rs # `ArrowSchema` — an owned Arrow C Data Interface schema
│ │ └── tests.rs # Unit tests
│ ├── bin/
│ │ └── append_metadata/
│ │ ├── cli.rs # Command-line parsing and validation (std-only, no clap)
│ │ ├── footer.rs # The 512-byte `DuckDB` extension footer, and the optional 22-byte WebAssembly custom-section header that precedes it
│ │ ├── main.rs # Append a DuckDB extension metadata block to a compiled .so / .dylib / .dll file,
│ │ └── tests.rs # Unit tests
│ ├── callback/
│ │ └── payload.rs # Disposing of a caught panic payload without re-entering the unwinder
│ ├── cast/
│ │ ├── builder.rs # Builder for registering custom `DuckDB` cast functions
│ │ └── mod.rs # Builder for registering `DuckDB` custom cast functions
│ ├── copy_function/
│ │ ├── info.rs # Callback info wrappers for copy function callbacks
│ │ └── mod.rs # Copy function registration (`DuckDB` 1.5.0+)
│ ├── datetime/
│ │ ├── checks.rs # Pure-Rust mirrors of the checks `DuckDB` makes before it throws
│ │ ├── mod.rs # Calendar conversions for `DuckDB`'s temporal types
│ │ └── tests.rs # Unit tests
│ ├── query/
│ │ ├── bind.rs # Binding a [`PreparedStatement`]'s parameters: the typed `bind_*` methods and `bind_value`
│ │ ├── chunk.rs # [`OwnedDataChunk`]: a `duckdb_data_chunk` destroyed on drop
│ │ ├── connection.rs # [`OwnedConnection`] and its cross-thread [`InterruptHandle`]
│ │ ├── cstr.rs # The two C-string conversions the `query` module runs everything through
│ │ ├── live_tests.rs # Tests that need a live `DuckDB`
│ │ ├── prepared.rs # Inspecting and executing a [`PreparedStatement`]; `bind.rs` binds its parameters
│ │ └── result.rs # Reading a [`QueryResult`]: its columns, its chunks and what kind of outcome it is
│ ├── replacement_scan/
│ │ └── mod.rs # Builder for registering `DuckDB` replacement scans
│ ├── scaffold/
│ │ ├── escape.rs # Quoting configured free text for YAML and Rust doc comments
│ │ ├── mod.rs # Project scaffolding for `DuckDB` Rust extensions
│ │ ├── templates.rs # Template generators for scaffold file content
│ │ ├── tests.rs # Unit tests
│ │ ├── tests_escaping.rs # Free text reaches the generated files intact
│ │ └── tests_generated.rs # Unit tests
│ ├── scalar/
│ │ ├── info.rs # Ergonomic wrapper around `duckdb_function_info` for scalar function callbacks
│ │ ├── mod.rs # Builder for registering `DuckDB` scalar functions
│ │ ├── state.rs # Typed bind data and per-thread local state for scalar functions (`DuckDB` 1.5.0+)
│ │ ├── typed.rs # Scalar functions written as ordinary Rust closures
│ │ ├── typed_builder.rs # The builder the closure-based scalar constructors return, and the one `extern "C"` trampoline they all share
│ │ └── builder/
│ │ ├── collision.rs # Refusing a scalar signature that would replace an existing one or make calls ambiguous
│ │ ├── mod.rs # Builder for registering `DuckDB` scalar functions
│ │ ├── overload.rs # One overload within a [`ScalarFunctionSetBuilder`]
│ │ ├── set.rs # Builder for registering a `DuckDB` scalar function set (multiple overloads)
│ │ ├── signature.rs # Detecting overloads that accept the same call
│ │ ├── single.rs # Builder for registering a single-signature `DuckDB` scalar function
│ │ └── tests.rs # Unit tests
│ ├── table/
│ │ ├── bind_data.rs # Type-safe bind data management for table functions
│ │ ├── builder.rs # Builder for registering `DuckDB` table functions
│ │ ├── cstr.rs # Panic-free `&str` → `CString` conversion for the callback info wrappers
│ │ ├── info.rs # Ergonomic wrappers around `DuckDB` callback info handles
│ │ ├── init_data.rs # Type-safe init data management for table functions
│ │ ├── mod.rs # Builder for registering `DuckDB` table functions
│ │ ├── type_check.rs # Detects logical types that `DuckDB` refuses without saying so
│ │ ├── typed.rs # Closure-based table functions with typed scan state
│ │ └── typed/
│ │ └── trampolines.rs # `extern "C"` trampolines behind [`TypedTableFunctionBuilder`][super::TypedTableFunctionBuilder]
│ ├── testing/
│ │ ├── bundled_api_init.cpp # Compiled only when the `bundled-test` Cargo feature is active
│ │ ├── harness.rs # [`AggregateTestHarness`] — test aggregate logic without `DuckDB`
│ │ ├── in_memory_db.rs # In-memory `DuckDB` helper for integration tests
│ │ ├── mock_registrar.rs # [`MockRegistrar`] — a [`Registrar`] implementation for testing
│ │ ├── mock_vector.rs # In-memory mock types for `DuckDB` vectors
│ │ ├── mod.rs # Test utilities for `DuckDB` extension development
│ │ └── mock_vector/
│ │ ├── reader.rs # `MockVectorReader` — an in-memory mock input vector
│ │ ├── tests.rs # Unit tests
│ │ └── writer.rs # `MockVectorWriter` — an in-memory mock output vector
│ ├── types/
│ │ ├── logical_type.rs # RAII wrapper for `duckdb_logical_type`
│ │ ├── mod.rs # `DuckDB` type system wrappers
│ │ ├── null_handling.rs # NULL propagation behaviour for `DuckDB` functions
│ │ ├── type_id.rs # Ergonomic enum of all `DuckDB` column types
│ │ └── logical_type/
│ │ └── construct.rs # Every `LogicalType` constructor, as `try_*` plus a panicking wrapper
│ ├── validate/
│ │ ├── extension_name.rs # Extension name validation per `DuckDB` community extension rules
│ │ ├── function_name.rs # SQL function name validation for `DuckDB` extensions
│ │ ├── mod.rs # Validation utilities for `DuckDB` community extension compliance
│ │ ├── platform.rs # `DuckDB` build platform validation
│ │ ├── release_profile.rs # Release profile validation for `DuckDB` loadable extensions
│ │ ├── semver.rs # Semantic versioning validation for `DuckDB` community extensions
│ │ ├── spdx.rs # SPDX license identifier validation for `DuckDB` community extensions
│ │ ├── spdx_exceptions.rs # The SPDX license-exception identifiers accepted after `WITH`
│ │ └── description_yml/
│ │ ├── mod.rs # Validation of `DuckDB` community extension `description.yml` files
│ │ ├── model.rs # A validated representation of a `DuckDB` community extension `description.yml`
│ │ ├── parser.rs # Parses and validates a `description.yml` string
│ │ ├── tests.rs # Unit tests
│ │ ├── tests_corpus.rs # Unit tests
│ │ ├── tests_yaml.rs # Unit tests
│ │ ├── validator.rs # Validates a `description.yml` string and returns `Ok(())` if it passes all checks
│ │ ├── yaml.rs # A reader for the subset of YAML that `description.yml` files use
│ │ ├── yaml11.rs # Which plain scalars PyYAML (YAML 1.1) reads as booleans, numbers, dates or null
│ │ └── yaml/
│ │ └── scalar.rs # Scalar-level pieces of the `description.yml` YAML reader: decoding plain, quoted, block and flow values, and recognising keys and comments
│ ├── value/
│ │ ├── blob.rs # `Value::as_blob` — `BLOB` extraction
│ │ ├── checks.rs # Pure-Rust preconditions checked before a `Value` call reaches `DuckDB`
│ │ ├── composite.rs # Composite constructors: `STRUCT`, `LIST`, `ARRAY`, `ENUM`, `MAP`, `UNION`
│ │ ├── defaults.rs # The defaulting accessors — `Value::as_*_or`
│ │ ├── getters.rs # The typed scalar accessors — `Value::as_i64`, `as_timestamp`, `as_decimal`, …
│ │ ├── hugeint.rs # Conversions between Rust's 128-bit integers and `DuckDB`'s split-word `HUGEINT` / `UHUGEINT` records
│ │ ├── nested.rs # Reading nested values: `LIST` elements, `STRUCT` fields, `MAP` entries
│ │ ├── render_guard.rs # Refusing to render a value `DuckDB` would throw on (an out-of-range timestamp from SQL)
│ │ ├── scalars.rs # The non-temporal scalar constructors
│ │ ├── temporal.rs # Temporal constructors, validated against `DuckDB`'s ranges
│ │ └── temporal_checks.rs # Pure-Rust range checks for the temporal types, derived from `DuckDB`'s source
│ └── vector/
│ ├── complex.rs # Complex type vector operations: STRUCT fields, LIST elements, MAP entries
│ ├── list_builder.rs # Safe construction of `LIST` and `MAP` output vectors
│ ├── mod.rs # Safe helpers for reading from and writing to `DuckDB` data vectors
│ ├── nested_null.rs # The child validity masks a NULL in a nested vector must also clear
│ ├── ops.rs # Whole-vector operations (`DuckDB` 1.5.0+)
│ ├── reader.rs # Safe typed reading from `DuckDB` data vectors
│ ├── string.rs # `DuckDB` `VARCHAR` and `BLOB` (`duckdb_string_t`) reading utilities
│ ├── struct_reader.rs # Batched, typed reader for STRUCT input vectors
│ ├── struct_writer.rs # Batched, typed writer for STRUCT output vectors
│ ├── uuid.rs # Converting between a `UUID`'s textual bits and `DuckDB`'s vector storage
│ ├── validity.rs # Validity bitmap helpers for `DuckDB` NULL tracking
│ └── writer.rs # Safe typed writing to `DuckDB` result vectors
├── tests/
│ ├── aggregate_leaks.rs # Aggregate states `DuckDB` never destroys leak no Rust heap
│ ├── append_metadata_cli.rs # The `append_metadata` binary run end to end: exit status, output and the file it writes
│ ├── ffi_roundtrip.rs # End-to-end FFI round-trips against a real `DuckDB`
│ ├── file_handle_close.rs # `FileHandle`'s `Drop` when the close fails (stubbed C API)
│ ├── handle_leaks.rs # Every RAII handle frees what `DuckDB` allocated for it (glibc)
│ ├── integration_test.rs # Integration tests for `quack-rs`
│ ├── secret_zeroize.rs # `SecretEntry` never frees a buffer that still holds a secret
│ └── ffi_roundtrip/
│ ├── agg_states.rs # Every aggregate state is dropped, including the ones `DuckDB` moves
│ ├── agg_window.rs # Aggregates in the running-window and sorted-aggregate paths
│ ├── appender_api.rs # `Appender` methods a no-op replacement survived: schemas, column types, `clear_columns`, defaults
│ ├── appender_rows.rs # What happens to buffered rows when an append fails mid-row
│ ├── arrow_export.rs # `arrow::data_chunk_to_arrow` refuses values `DuckDB` would export wrongly
│ ├── arrow_import.rs # `arrow::data_chunk_from_arrow` checks against a live `DuckDB`
│ ├── arrow_layout.rs # Valid Arrow layouts `DuckDB` imports wrongly, refused, and their correct neighbours
│ ├── bind_expressions.rs # What a bind callback learns about its arguments from `Expression`
│ ├── chunk_writer.rs # `ChunkWriter` against a chunk `DuckDB` allocated
│ ├── collision.rs # The scalar signature-collision check, held to `DuckDB`'s own binder
│ ├── copy_from_columns.rs # A typed `COPY … FROM` reader that declares a column is refused
│ ├── file_errors.rs # `FileHandle` reports write, sync and seek failures, and refuses a seek past `i64::MAX`
│ ├── handles_api.rs # `StructWriter` child handles and `InMemoryDb::execute`'s row count
│ ├── lifecycle.rs # Aggregate NULL rows, name collisions, overload builders, bind-data sharing
│ ├── list_limits.rs # `ListBuilder` stops at `DuckDB`'s byte ceiling, not an element count
│ ├── mock_parity.rs # `MockRegistrar` refuses a bad type with the live registration's message
│ ├── nested_reserve.rs # which nested buffers a `LIST` reserve moves (the writer contracts)
│ ├── nested_validity.rs # `VectorWriter::set_valid` on nested rows, against a live `DuckDB`
│ ├── panic_guards.rs # The panic-guard macros and `set_error` methods, against a live `DuckDB`
│ ├── query_docs.rs # Pins the documented behaviour of `query`, `PreparedStatement`, `DbConfig`
│ ├── query_stream.rs # A streaming result that stops early must not look like a finished one
│ ├── scalar_agg.rs # Scalar and aggregate builder regressions
│ ├── table_cast.rs # Table function, cast, replacement scan, SQL macro and COPY regressions
│ ├── table_description.rs # `TableDescription` accessors on an index `DuckDB` cannot hold
│ ├── temporal_binds.rs # Temporal and over-4-GiB values refused by `PreparedStatement` binds and the `Appender`
│ ├── tooling.rs # Checks of quack-rs's tooling tables against the linked `DuckDB`
│ ├── value_getters.rs # Each typed `Value` getter at its own type; out-of-range TIMETZ / TIME_NS are not cast
│ ├── value_nested.rs # Nested `Value` construction and inspection against a live `DuckDB`
│ ├── value_query.rs # `Value` getters, DECIMAL binding, `Expression::fold`
│ ├── value_render.rs # `Value` rendering of values SQL builds and `DuckDB` cannot render
│ ├── value_temporal.rs # Every `Value` getter against every temporal source type, at every edge
│ └── vector_dt.rs # NULLs in nested output vectors; selection vectors
├── benches/
│ └── interval_bench.rs # Criterion benchmarks
├── examples/
│ └── hello-ext/ # Reference extension: aggregates, scalars, table functions, casts
├── book/ # mdBook documentation source
│ ├── src/ # Markdown pages (this site)
│ └── theme/custom.css
├── .github/workflows/ci.yml # CI pipeline
├── .github/workflows/docs.yml # GitHub Pages deployment
├── CONTRIBUTING.md
├── LESSONS.md # The DuckDB Rust FFI pitfalls (L1–L19, P1–P12)
├── CHANGELOG.md
└── README.md
Releasing
quack-rs uses libduckdb-sys = ">=1.4.4, <2" — a bounded range covering DuckDB 1.4.x
and 1.5.x, every one of which loads C API v1.2.0 extensions. The <2 upper bound
prevents silent adoption of a future major release that may change the C API.
Before broadening the range to a new major band:
- Read the DuckDB changelog for C API changes
- Check the new C API version string (used in
duckdb_rs_extension_api_init) - Update
DUCKDB_API_VERSIONinsrc/lib.rsif the C API version changed - Audit all callback signatures against the new
libduckdb-sysbindings - Update the
libduckdb-sysandduckdbversion requirements inCargo.toml
Versions follow Semantic Versioning as Cargo applies it.
While the crate is pre-1.0, a breaking change to the public API bumps the
minor version (0.17.x → 0.18.0) and is marked Breaking: in
CHANGELOG.md; see the semantic versioning policy in
RELEASING.md.
Reporting issues
Use GitHub Issues. For security
vulnerabilities, see SECURITY.md
for responsible disclosure policy.
License
quack-rs is licensed under the MIT License. Contributions are accepted under the same license. By submitting a pull request, you agree to license your contribution under MIT.