quack-rs: DuckDB extensions in Rust

A Rust SDK for building DuckDB loadable extensions — no C++ required.

CI Crates.io Documentation License: MIT MSRV: 1.86.0


What is quack-rs?

quack-rs is a Rust SDK for building DuckDB loadable extensions without C++ or CMake. It wraps DuckDB's C Extension API — the C interface DuckDB exposes to loadable extensions — in safe builders and RAII types, and guards against the FFI pitfalls documented in the Pitfall Catalog, so you can focus on extension logic.

DuckDB's own documentation acknowledges the gap:

"Writing a Rust-based DuckDB extension requires writing glue code in C++ and will force you to build through DuckDB's CMake & C++ based extension template. We understand that this is not ideal and acknowledge the fact that Rust developers prefer to work on pure Rust codebases."

— DuckDB Community Extensions FAQ

quack-rs closes that gap. No C++. No CMake. No glue code.


What you can build

Extension typequack-rs support
Scalar functions✅ ScalarFunctionBuilder
Overloaded scalars✅ ScalarFunctionSetBuilder
Aggregate functions✅ AggregateFunctionBuilder
Overloaded aggregates✅ AggregateFunctionSetBuilder (per-overload return types)
Table functions✅ TableFunctionBuilder (raw) + TypedTableFunctionBuilder<S> (closure-based, typed scan state)
Cast / TRY_CAST functions✅ CastFunctionBuilder
Replacement scans✅ ReplacementScanBuilder
SQL macros (scalar)✅ SqlMacro::scalar
SQL macros (table)✅ SqlMacro::table
Copy functions (COPY TO / COPY FROM)✅ CopyFunctionBuilder (requires duckdb-1-5)

Note: Window functions have no counterpart in DuckDB's public C Extension API and cannot be implemented from Rust (or any language) via that API. See Known Limitations.


Why does this exist?

quack-rs was extracted from duckdb-behavioral, a production DuckDB community extension. Building that extension revealed 16 undocumented pitfalls in DuckDB's Rust FFI surface — struct layouts, callback contracts, and initialization sequences that aren't covered anywhere in the DuckDB documentation or libduckdb-sys docs.

Three of those pitfalls caused extension-breaking bugs that passed 435 unit tests before being caught by end-to-end tests:

  1. A SEGFAULT on load (wrong entry point sequence)
  2. 6 of 7 functions silently not registered (undocumented function-set naming rule)
  3. Wrong aggregate results under parallel plans (combine callback not propagating configuration fields to fresh target states)

quack-rs rules out the first two: entry_point! performs the correct initialization sequence, and the function-set builders name every member. The third lives in your own combine callback, so no API can prevent it; AggregateTestHarness::combine lets you test for it without DuckDB. The full list is in the Pitfall Catalog.


Key features

  • Zero C++ — no CMakeLists.txt, no header files, no glue code
  • Every function kind the C API can register — scalar, aggregate, table, cast, replacement scan, copy function (duckdb-1-5) — plus SQL macros
  • Panic-safe FFI — the entry point and the callbacks quack-rs generates catch panics and report them as SQL errors; registration errors surface via Result
  • RAII memory management — LogicalType and FfiState<T> prevent leaks and double-frees
  • Type-safe builders — ScalarFunctionBuilder, AggregateFunctionBuilder, TableFunctionBuilder, CastFunctionBuilder, ReplacementScanBuilder
  • SQL macros — register CREATE MACRO statements without any FFI callbacks
  • Testable state — AggregateTestHarness<T> tests aggregate logic without a live DuckDB
  • Scaffold generator — generates a complete community extension project (Cargo.toml, Makefile, CI, description.yml, tests) from one function call
  • 31 pitfalls documented — every known DuckDB Rust FFI pitfall, with symptom, root cause and fix

New to DuckDB extensions? → Start with Quick Start

Adding quack-rs to an existing project? → See Installation

Writing your first function? → See Scalar Functions or Aggregate Functions

Want SQL macros without FFI callbacks? → See SQL Macros

Submitting a community extension? → See Community Extensions

Something broke? → See Pitfall Catalog

Quick Start

This page takes you from an empty crate to a DuckDB loadable extension written in Rust, loaded into the DuckDB CLI, in three steps.


Prerequisites

  • Rust ≥ 1.86.0 (MSRV) — install via rustup
  • DuckDB CLI (for testing the built extension) — download

Step 1 — Add quack-rs to your extension

In your extension's Cargo.toml:

[dependencies]
quack-rs = "0.18"
libduckdb-sys = { version = ">=1.4.4, <2", features = ["loadable-extension"] }

[lib]
name = "my_extension"       # must match your extension name — see Pitfall P1
crate-type = ["cdylib", "rlib"]

[profile.release]
panic = "unwind"            # required — quack-rs's panic guards need unwinding (see Installation)
lto = true
opt-level = 3
codegen-units = 1
strip = true

Starting from scratch? The scaffold generator generates a complete community extension project, including this Cargo.toml.


Step 2 — Write the extension

#![allow(unused)]
fn main() {
// src/lib.rs
use quack_rs::entry_point;
use quack_rs::error::ExtensionError;
use quack_rs::scalar::ScalarFunctionBuilder;
use quack_rs::types::TypeId;
use quack_rs::vector::{VectorReader, VectorWriter};
use libduckdb_sys::{duckdb_connection, duckdb_function_info, duckdb_data_chunk, duckdb_vector};

/// Scalar function: double_it(BIGINT) → BIGINT
unsafe extern "C" fn double_it(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    // SAFETY: input is a valid data chunk provided by DuckDB.
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };
    let row_count = reader.row_count();

    for row in 0..row_count {
        if unsafe { !reader.is_valid(row) } {
            unsafe { writer.set_null(row) };
            continue;
        }
        let value = unsafe { reader.read_i64(row) };
        unsafe { writer.write_i64(row, value * 2) };
    }
}

fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        ScalarFunctionBuilder::new("double_it")
            .param(TypeId::BigInt)
            .returns(TypeId::BigInt)
            .function(double_it)
            .register(con)?;
    }
    Ok(())
}

entry_point!(my_extension_init_c_api, |con| register(con));
}

Step 3 — Build and test

# Build the extension
cargo build --release

# DuckDB refuses to LOAD a bare .so: append the metadata footer first.
# quack-rs ships the `append_metadata` binary for this step.
cargo install quack-rs --bin append_metadata
append_metadata target/release/libmy_extension.so my_extension.duckdb_extension \
    --abi-type C_STRUCT --extension-version v0.1.0 \
    --duckdb-version v1.2.0 --platform linux_amd64

# Load it; -unsigned allows a locally built, unsigned extension.
duckdb -unsigned -c "LOAD './my_extension.duckdb_extension'; SELECT double_it(21);"
# ┌───────────────┐
# │ double_it(21) │
# │     int64     │
# ├───────────────┤
# │            42 │
# └───────────────┘

macOS: the library is libmy_extension.dylib and the platform osx_arm64 (or osx_amd64). Windows: my_extension.dll and windows_amd64.


What's next?

Installation

This page covers the Cargo.toml dependencies and release-profile settings a DuckDB loadable extension built with quack-rs needs, the minimum Rust version, and the optional test features.

Adding quack-rs to an existing extension

Add the following to your extension's Cargo.toml:

[dependencies]
quack-rs = "0.18"
libduckdb-sys = { version = ">=1.4.4, <2", features = ["loadable-extension"] }

Why >=1.4.4, <2? Every DuckDB 1.4.x and 1.5.x release loads extensions built for C API version v1.2.0 (1.5.6 declares v1.5.6 and accepts every earlier one), so quack-rs supports both with a single bounded range. The <2 upper bound prevents silent adoption of a future major release whose C API may change in breaking ways — making any such upgrade an explicit, auditable decision. See Extension Anatomy.


Required Cargo.toml settings

Every DuckDB extension requires specific Cargo settings to link and behave correctly:

[lib]
name = "my_extension"       # ← must match extension name exactly (Pitfall P1)
crate-type = ["cdylib", "rlib"]
#             ^^^^^^  cdylib produces the .so/.dylib/.dll DuckDB loads
#                      rlib   optional: lets doctests, examples and tests/ link the crate

[profile.release]
panic = "unwind"            # REQUIRED — quack-rs catches panics at every FFI boundary;
                            #   "abort" makes that impossible (see below)
lto = true                  # recommended — reduces binary size, improves performance
opt-level = 3               # recommended
codegen-units = 1           # recommended — better optimisation, slower build
strip = true                # recommended — reduces binary size

Why panic = "unwind", not "abort"?

Every callback quack-rs generates (the *_callback! macros, the closure-based builders, FfiState's callbacks), and every entry point, runs your code inside std::panic::catch_unwind and turns a panic into an ordinary SQL error that DuckDB reports to the user. catch_unwind can only catch a panic that unwinds: under panic = "abort" the process terminates at the panic site, before any guard runs, taking the user's whole DuckDB session with it.

A raw unsafe extern "C" fn that you pass to a builder yourself (such as double_it in the Quick Start) is installed as written, with no guard. Define it with the matching macro (scalar_callback!, aggregate_update_callback!, …), which reports a panic as a SQL error, or wrap its body in quack_rs::callback::catch_ffi_panic and report the Err it returns yourself.

(A panic that escapes an extern "C" function without being caught is not undefined behaviour on Rust ≥ 1.81 — the runtime aborts the process — but that is exactly the outcome the guards exist to prevent.)

validate_release_profile rejects panic = "abort", and the scaffold generator emits panic = "unwind".


Minimum Supported Rust Version

quack-rs requires Rust ≥ 1.86.0.

1.86.0 is a ceiling as much as a floor. DuckDB's community-extension build workflow (_extension_distribution.yml in duckdb/extension-ci-tools) pins Rust 1.86.0 for its WebAssembly jobs, so an extension — and therefore quack-rs — must build on it; CI's msrv-vs-duckdb-ci job re-derives that pin and fails if the MSRV rises above it. The msrv job checks the crate with cargo +1.86.0 check, and the benchmark dev-dependency (criterion) needs 1.86 as well.

Install or update via:

rustup update stable
rustup default stable

Verify:

rustc --version   # must be ≥ 1.86.0

Development dependencies

To run SQL against your functions inside cargo test, enable one of quack-rs's test features as a dev-dependency:

[dev-dependencies]
# Compiles DuckDB from C++ source: no setup, slow cold build.
quack-rs = { version = "0.18", features = ["bundled-test"] }
# ...or link a prebuilt libduckdb instead (set DUCKDB_DOWNLOAD_LIB=1 or DUCKDB_LIB_DIR):
# quack-rs = { version = "0.18", features = ["bundled-test-prebuilt"] }

Either one initialises the loadable-extension dispatch table from the linked DuckDB, so testing::InMemoryDb works and the whole C API — including your own registration code — can be exercised in a test. Without them, any duckdb_* call in a cargo test process panics, because nothing has filled the dispatch table. See the Testing Guide.


Starting a new extension from scratch

Use the scaffold generator to produce a complete project with these settings, the build files and CI already in place.

Your First Extension

This page builds a DuckDB extension in Rust step by step, using hello-ext, the example extension bundled with quack-rs. The table lists four of the functions it registers, one of each major kind (the full list is in its README); this page walks through the aggregate and the scalar. The table function and the cast are covered in Table Functions and Cast Functions.

SQLKindSignature
word_count(text)AggregateVARCHAR → BIGINT
first_word(text)ScalarVARCHAR → VARCHAR
generate_series_ext(n)TableBIGINT → TABLE(value BIGINT)
CAST(VARCHAR AS INTEGER)CastVARCHAR → INTEGER

Full source: examples/hello-ext/src/lib.rs


Build and try it

cargo build --release --manifest-path examples/hello-ext/Cargo.toml

# DuckDB only loads files ending in `.duckdb_extension` that carry its
# 512-byte metadata footer; a bare `.so` is refused. Append the footer:
cargo run --bin append_metadata -- \
    examples/hello-ext/target/release/libhello_ext.so \
    hello_ext.duckdb_extension \
    --abi-type C_STRUCT --extension-version v0.1.0 \
    --duckdb-version v1.2.0 --platform linux_amd64

# An unsigned local build needs -unsigned (it must be a startup flag).
duckdb -unsigned

Then in the DuckDB CLI:

LOAD './hello_ext.duckdb_extension';

-- Aggregate: total words across all rows
SELECT word_count(sentence) FROM (
    VALUES ('hello world'), ('one two three'), (NULL)
) t(sentence);
-- → 5  (2 + 3; NULL contributes 0)

-- Scalar: first word of each row
SELECT first_word(sentence) FROM (
    VALUES ('hello world'), ('  padded  '), (''), (NULL)
) t(sentence);
-- → 'hello', 'padded', '', NULL

Overview

An extension has four parts:

  1. State struct — holds data accumulated during aggregation (aggregate only)
  2. Callbacks — update, combine, finalize (aggregate; ffi_state::<T>() supplies state_size, state_init and state_destroy) or a single function callback (scalar)
  3. Registration — wire callbacks to DuckDB via AggregateFunctionBuilder / ScalarFunctionBuilder
  4. Entry point — DuckDB's initialization hook, generated by entry_point!

Part 1 — Aggregate function: word_count

An aggregate function accumulates state across many rows and emits one result per group.

1a. The state struct

#![allow(unused)]
fn main() {
use quack_rs::prelude::*;
#[derive(Default, Debug)]
struct WordCountState {
    count: i64,
}

impl AggregateState for WordCountState {}
}

AggregateState is a marker trait with no methods; its supertraits require the state to be Default + Send + Sync + 'static. FfiState<WordCountState> stores it in the bytes DuckDB allocates for each group (boxing it only when it is large or over-aligned) and manages its lifecycle (size, init, destroy).

1b. state_size, state_init and state_destroy

These three callbacks are always identical boilerplate, so you do not write them: .ffi_state::<WordCountState>() on the builder (see Part 3) installs FfiState::<WordCountState>::size_callback, init_callback and destroy_callback together. Installing all three for one type in a single call means DuckDB never allocates space sized for one state and has it initialised as another.

size_callback returns FfiState::<WordCountState>::size() — one usize tag word followed by the state. A T aligned no more strictly than usize and at most 256 bytes (like WordCountState) is stored inline in the bytes DuckDB allocates per group; a larger or more strictly aligned T is boxed, and the slot holds the Box pointer. init_callback writes WordCountState::default() into the slot (or boxes it), then sets the tag — a value salted per state type, so a destructor only drops states that were initialised for its own T.

destroy_callback drops the T in each state (and its box, if the state is too large to be stored inline), clearing the state's tag first so a second call is a no-op.

1c. update — accumulate one batch

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 }
unsafe extern "C" fn wc_update(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    states: *mut duckdb_aggregate_state,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let row_count = reader.row_count();

    for row in 0..row_count {
        if !unsafe { reader.is_valid(row) } {
            continue; // NULL input → skip (contributes 0 words)
        }
        let s = unsafe { reader.read_str(row) };
        let words = count_words(s);

        let state_ptr = unsafe { *states.add(row) };
        if let Some(st) = unsafe { FfiState::<WordCountState>::with_state_mut(state_ptr) } {
            st.count += words;
        }
    }
}
}

Key points:

  • Check is_valid(row) before reading — never dereference an invalid (NULL) row
  • VectorReader::new(chunk, col) gives column col from the chunk
  • count_words is pure Rust — no unsafe, easy to unit-test separately

1d. combine — merge parallel results

Pitfall L1: DuckDB creates fresh target states before calling combine, set up by state_init (here WordCountState::default()), not copies of the source. You must copy all fields — not just the result field. In an aggregate with config fields (e.g., a histogram with a bin_width) you must also copy those, or results will be silently corrupted.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
unsafe extern "C" fn wc_combine(
    _info: duckdb_function_info,
    source: *mut duckdb_aggregate_state,
    target: *mut duckdb_aggregate_state,
    count: idx_t,
) {
    for i in 0..count as usize {
        let src_ptr = unsafe { *source.add(i) };
        let tgt_ptr = unsafe { *target.add(i) };
        let src = unsafe { FfiState::<WordCountState>::with_state(src_ptr) };
        let tgt = unsafe { FfiState::<WordCountState>::with_state_mut(tgt_ptr) };
        if let (Some(s), Some(t)) = (src, tgt) {
            t.count += s.count;
            // If you add fields to WordCountState, combine them here too.
        }
    }
}
}

1e. finalize — write output

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
unsafe extern "C" fn wc_finalize(
    _info: duckdb_function_info,
    source: *mut duckdb_aggregate_state,
    result: duckdb_vector,
    count: idx_t,
    offset: idx_t,
) {
    let mut writer = unsafe { VectorWriter::new(result) };

    for i in 0..count as usize {
        let state_ptr = unsafe { *source.add(i) };
        match unsafe { FfiState::<WordCountState>::with_state(state_ptr) } {
            Some(st) => unsafe { writer.write_i64(offset as usize + i, st.count) },
            None     => unsafe { writer.set_null(offset as usize + i) },
        }
    }
}
}

offset is DuckDB's output row offset — always use offset as usize + i, not just i.


Part 2 — Scalar function: first_word

A scalar function processes one data chunk and returns one output value per row. The callback receives the full chunk and an output vector (not per-row state pointers).

Key rule: always propagate NULL

If the input row is NULL, write NULL to output — never read from an invalid row.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") }
unsafe extern "C" fn first_word_scalar(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };
    let row_count = reader.row_count();

    for row in 0..row_count {
        if !unsafe { reader.is_valid(row) } {
            unsafe { writer.set_null(row) }; // NULL in → NULL out
            continue;
        }
        let s = unsafe { reader.read_str(row) };
        unsafe { writer.write_varchar(row, first_word(s)) };
    }
}
}

The pure logic:

#![allow(unused)]
fn main() {
pub fn first_word(s: &str) -> &str {
    s.split_whitespace().next().unwrap_or("")
}
}

Note: set_null internally calls duckdb_vector_ensure_validity_writable before writing the null flag — this is required by DuckDB and handled for you by VectorWriter.


Part 3 — Registration

The snippets below register through the raw builders and entry_point!, which take a duckdb_connection. hello-ext itself uses the equivalent entry_point_v2!, whose closure receives a &Connection and registers each builder through the Registrar trait (con.register_aggregate(...), con.register_scalar(...)); see The Entry Point.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 }
fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") }
unsafe extern "C" fn wc_update(_info: duckdb_function_info, input: duckdb_data_chunk,
    states: *mut duckdb_aggregate_state) {
    let reader = unsafe { VectorReader::new(input, 0) };
    for row in 0..reader.row_count() {
        if !unsafe { reader.is_valid(row) } { continue; }
        let words = count_words(unsafe { reader.read_str(row) });
        if let Some(st) = unsafe { FfiState::<WordCountState>::with_state_mut(*states.add(row)) } {
            st.count += words;
        }
    }
}
unsafe extern "C" fn wc_combine(_info: duckdb_function_info, source: *mut duckdb_aggregate_state,
    target: *mut duckdb_aggregate_state, count: idx_t) {
    for i in 0..count as usize {
        let src = unsafe { FfiState::<WordCountState>::with_state(*source.add(i)) };
        let tgt = unsafe { FfiState::<WordCountState>::with_state_mut(*target.add(i)) };
        if let (Some(s), Some(t)) = (src, tgt) { t.count += s.count; }
    }
}
unsafe extern "C" fn wc_finalize(_info: duckdb_function_info, source: *mut duckdb_aggregate_state,
    result: duckdb_vector, count: idx_t, offset: idx_t) {
    let mut writer = unsafe { VectorWriter::new(result) };
    for i in 0..count as usize {
        match unsafe { FfiState::<WordCountState>::with_state(*source.add(i)) } {
            Some(st) => unsafe { writer.write_i64(offset as usize + i, st.count) },
            None => unsafe { writer.set_null(offset as usize + i) },
        }
    }
}
unsafe extern "C" fn first_word_scalar(_info: duckdb_function_info, input: duckdb_data_chunk,
    output: duckdb_vector) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };
    for row in 0..reader.row_count() {
        if !unsafe { reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; }
        unsafe { writer.write_varchar(row, first_word(reader.read_str(row))) };
    }
}
fn live_connection() -> libduckdb_sys::duckdb_connection {
    std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
    let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess);
    }
    con
}
/// First column of the first row, as BIGINT; `None` for NULL.
fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    let reader = unsafe { chunk.reader(0) };
    unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) }
}
unsafe fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        AggregateFunctionBuilder::new("word_count")
            .param(TypeId::Varchar)
            .returns(TypeId::BigInt)
            .ffi_state::<WordCountState>() // state_size + init + destructor
            .update(wc_update)
            .combine(wc_combine)
            .finalize(wc_finalize)
            .register(con)?;

        ScalarFunctionBuilder::new("first_word")
            .param(TypeId::Varchar)
            .returns(TypeId::Varchar)
            .function(first_word_scalar)
            .register(con)?;
    }
    Ok(())
}
let con = live_connection();
unsafe { register(con) }.unwrap();
assert_eq!(query_i64(con, "SELECT word_count(s) FROM (VALUES ('hello world'), (NULL), ('one two three')) t(s)"), Some(5));
assert_eq!(query_i64(con, "SELECT length(first_word('  quack rs'))::BIGINT"), Some(5));
assert_eq!(query_i64(con, "SELECT count(*) FROM (SELECT first_word(NULL) AS w) WHERE w IS NULL"), Some(1));
// Many groups, so DuckDB runs combine on partial states.
assert_eq!(query_i64(con, "SELECT sum(c)::BIGINT FROM (SELECT word_count('a b') AS c FROM range(100000) GROUP BY range % 997)"), Some(200000));
}

Both builders call the DuckDB C API internally. register returns Err if DuckDB reports a failure — this propagates to the entry point and is surfaced to the user.


Part 4 — Entry point

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe fn register(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
quack_rs::entry_point!(hello_ext_init_c_api, |con| unsafe { register(con) });
}

This one line expands to the equivalent of:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_connection, duckdb_extension_access, duckdb_extension_info};
use quack_rs::error::ExtensionError;
unsafe fn register(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
#[no_mangle]
pub unsafe extern "C" fn hello_ext_init_c_api(
    info: duckdb_extension_info,
    access: *const duckdb_extension_access,
) -> bool {
    unsafe {
        quack_rs::entry_point::init_extension_with_policy(
            info, access, quack_rs::DUCKDB_API_VERSION,
            quack_rs::abi::AbiPolicy::Strict,
            |con| unsafe { register(con) },
        )
    }
}
}

Pass the full symbol name — hello_ext_init_c_api here. DuckDB looks up this exact symbol when loading the extension. See The Entry Point for the full initialization sequence.


Unit tests (no DuckDB process needed)

Test pure logic directly:

#![allow(unused)]
fn main() {
fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 }
fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") }
#[test]
fn count_words_whitespace_variants() {
    assert_eq!(count_words("  hello  world  "), 2);
    assert_eq!(count_words("\t\nhello\tworld\n"), 2);
    assert_eq!(count_words("   "), 0); // all whitespace → 0
}

#[test]
fn first_word_empty_and_whitespace() {
    assert_eq!(first_word(""), "");
    assert_eq!(first_word("   "), "");
}
}

Test aggregate state with AggregateTestHarness:

#![allow(unused)]
fn main() {
use quack_rs::prelude::*;
use quack_rs::testing::AggregateTestHarness;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 }
#[test]
fn word_count_null_rows_are_skipped() {
    // DuckDB passes NULL rows to `update`; the callback's `is_valid` check
    // skips them, so they never reach the state.
    let mut h = AggregateTestHarness::<WordCountState>::new();
    h.update(|s| s.count += count_words("hello"));
    // NULL row omitted — models the callback's `is_valid` skip
    h.update(|s| s.count += count_words("world"));
    assert_eq!(h.finalize().count, 2);
}

#[test]
fn word_count_combine() {
    let mut h1 = AggregateTestHarness::<WordCountState>::new();
    h1.update(|s| s.count += count_words("hello world")); // 2

    let mut h2 = AggregateTestHarness::<WordCountState>::new();
    h2.update(|s| s.count += count_words("one two three four")); // 4

    h2.combine(&h1, |src, tgt| tgt.count += src.count);
    assert_eq!(h2.finalize().count, 6);
}
}

Run all tests with:

cargo test --manifest-path examples/hello-ext/Cargo.toml

See the Testing Guide for the full test strategy.

Project Scaffold

quack_rs::scaffold::generate_scaffold generates every file of a new DuckDB community extension project in Rust — Cargo.toml, Makefile, CI workflow, description.yml, an example function and its tests — from a single function call.


What it generates

my_extension/
├── Cargo.toml                          # cdylib crate, dependencies, release profile
├── Makefile                            # delegates to cargo + extension-ci-tools
├── extension_config.cmake              # required by extension-ci-tools
├── src/
│   ├── lib.rs                          # entry point, example function, unit test
│   └── wasm_lib.rs                     # WASM staticlib shim
├── description.yml                     # community extension metadata
├── test/
│   └── sql/
│       └── my_extension.test           # SQLLogicTest for the example function
├── .github/
│   └── workflows/
│       └── extension-ci.yml            # cross-platform CI workflow
├── .gitmodules                         # extension-ci-tools submodule
├── .gitignore
└── .cargo/
    └── config.toml                     # Windows CRT static linking

Usage

generate_scaffold returns paths relative to the project root (Cargo.toml, src/lib.rs, ...) and writes nothing itself. Join them under the directory the project should live in — writing them relative to the current directory would overwrite whatever Cargo.toml and src/lib.rs are already there.

use quack_rs::scaffold::{ScaffoldConfig, generate_scaffold};
use std::path::Path;

fn main() {
    let config = ScaffoldConfig {
        name: "my_extension".to_string(),
        description: "My DuckDB extension".to_string(),
        version: "0.1.0".to_string(),
        license: "MIT".to_string(),
        maintainer: "Your Name".to_string(),
        github_repo: "yourorg/duckdb-my-extension".to_string(),
        excluded_platforms: vec![],
        // `target_duckdb_version`, `use_unstable_c_api` and `git_ref` default
        // to the stable-ABI settings; see `concepts/abi.md`.
        ..ScaffoldConfig::default()
    };

    let files = generate_scaffold(&config).expect("scaffold generation failed");

    // Everything goes under ./my_extension/, never into the current directory.
    let root = Path::new(&config.name);
    for file in &files {
        let path = root.join(&file.path);
        if let Some(parent) = path.parent() {
            std::fs::create_dir_all(parent).unwrap();
        }
        std::fs::write(&path, &file.content).unwrap();
        println!("created {}", path.display());
    }
}

ScaffoldConfig fields

FieldTypeDescription
nameStringExtension name — must match [lib] name in Cargo.toml and description.yml
descriptionStringDescription for description.yml and the //! docs of src/lib.rs. Quoted and escaped, so :, # and quotes are fine; must not be empty, padded with whitespace, or contain control characters other than newline and tab
versionStringThe extension's version — validated by validate_extension_version
licenseStringSPDX license identifier (e.g., "MIT", "Apache-2.0")
maintainerStringYour name or org, listed in description.yml — one non-empty line
github_repoString"owner/repo", in the characters GitHub allows
excluded_platformsVec<String>Platforms to skip (e.g., ["wasm_mvp", "wasm_eh"])
git_refStringrepo.ref — a commit hash (or tag), not a branch. Defaults to REF_PLACEHOLDER so it cannot be submitted unset
target_duckdb_versionStringWritten as TARGET_DUCKDB_VERSION in the Makefile. Defaults to DUCKDB_API_VERSION (v1.2.0); with use_unstable_c_api, an exact DuckDB release such as v1.5.5
use_unstable_c_apiboolSet when the extension enables duckdb-1-5 / duckdb-1-5-3 / duckdb-1-5-4. Defaults to false

Name validation

Extension names must satisfy all of:

  • Match ^[a-z][a-z0-9_]*$ (no hyphens: the entry point is <name>_init_c_api)
  • Not exceed 64 characters
  • Be globally unique on community-extensions.duckdb.org

Use vendor-prefixed names to avoid collisions: myorg_analytics, not analytics.

The scaffold generator checks the first two rules before generating any files and returns an error if the name breaks one. Uniqueness is yours to check.


After scaffolding

cd my_extension
git init
git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools
make configure
make release
make test

git submodule add is not optional in a new project: the scaffold writes .gitmodules, but git submodule update --init does nothing until the submodule has been added once (Pitfall P4). The generated Makefile stops with this command if the checkout is missing.

cargo test runs the generated unit test and make test runs test/sql/my_extension.test against a real DuckDB. Replace the example function in src/lib.rs with your own, give each new function a test in both places, and push to GitHub — CI runs automatically.


Excluded platforms

Some extensions cannot be built for all platforms (e.g., extensions that depend on platform-specific system libraries, or WASM environments that lack threading).

#![allow(unused)]
fn main() {
use quack_rs::scaffold::ScaffoldConfig;

let config = ScaffoldConfig {
    excluded_platforms: vec![
        "wasm_mvp".to_string(),
        "wasm_eh".to_string(),
        "wasm_threads".to_string(),
    ],
    ..ScaffoldConfig::default()
};
}

Validate individual platform names with quack_rs::validate::validate_platform, or a semicolon-delimited string (as used in description.yml) with quack_rs::validate::validate_excluded_platforms_str.

Extension Anatomy

A DuckDB loadable extension is a shared library (.so / .dylib / .dll) that DuckDB loads at runtime. This page covers what DuckDB expects of a Rust extension built on the C Extension API: the entry-point symbol, the initialization sequence, the loadable-extension dispatch table, and version and binary compatibility.


The initialization sequence

When DuckDB loads your extension, it:

  1. Opens the shared library and looks up the symbol {name}_init_c_api, where {name} is the extension name
  2. Calls that function with an info handle and a pointer to a duckdb_extension_access struct (the set_error, get_database and get_api callbacks)
  3. Your function must:
    1. Call duckdb_rs_extension_api_init(info, access, api_version) to initialize the dispatch table
    2. Get the duckdb_database handle via access.get_database(info)
    3. Open a duckdb_connection via duckdb_connect
    4. Register functions on that connection
    5. Disconnect
    6. Return true (success) or false (failure), reporting any error through access.set_error

quack_rs::entry_point::init_extension performs this sequence. It also checks the C API layout before registration (see ABI Compatibility) and converts a panic in your registration code into a load error. The entry_point! macro generates the required #[no_mangle] extern "C" symbol:

#![allow(unused)]
fn main() {
use quack_rs::entry_point;
use quack_rs::error::ExtensionError;
fn register(_con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
entry_point!(my_extension_init_c_api, |con| register(con));
// emits: #[no_mangle] pub unsafe extern "C" fn my_extension_init_c_api(...)
}

Symbol naming

The symbol name must be {extension_name}_init_c_api. Extension names use lowercase ASCII letters, digits and underscores only (no hyphens, which cannot appear in a C symbol). If the symbol is missing or misnamed, DuckDB fails to load the extension.

Extension name: "word_count_ext"
Required symbol: word_count_ext_init_c_api

Pass the full symbol name to entry_point!. The exported name then appears verbatim at the call site; the macro does not build identifiers at compile time.


The loadable-extension feature

With features = ["loadable-extension"], libduckdb-sys does not link DuckDB; every C API function dispatches through a table of function pointers instead:

Without feature:  duckdb_query(...)  →  calls linked libduckdb directly
With feature:     duckdb_query(...)  →  dispatches through an AtomicPtr table

The AtomicPtr table starts as null. The extension's entry point fills it by calling duckdb_rs_extension_api_init, which copies the pointers out of the API struct DuckDB provides. This means:

  • Any call before duckdb_rs_extension_api_init panics with "DuckDB API not initialized or DuckDB feature omitted"
  • In a plain cargo test, you cannot call any duckdb_* function — no DuckDB host process ever initializes the table

This is why quack-rs offers AggregateTestHarness for testing: it simulates the aggregate lifecycle in pure Rust, without calling the DuckDB API. For SQL-level tests, the bundled-test / bundled-test-prebuilt features link a real DuckDB and fill the table when InMemoryDb::open() runs (see Testing Guide).


Dependency model

graph TD
    EXT["your-extension"]
    QR["quack-rs"]
    LDS["libduckdb-sys >=1.4.4, <2<br/>{loadable-extension}<br/>(headers only — no linked library)"]

    EXT --> QR
    EXT --> LDS
    QR  --> LDS

The loadable-extension feature produces a shared library that does not statically link DuckDB. Instead, it receives DuckDB's function pointers at load time, so the extension runs inside the host DuckDB process and uses that process's DuckDB instance, memory and threads.


Version support

libduckdb-sys = ">=1.4.4, <2" — the bounded range is intentional.

Every DuckDB 1.4.x and 1.5.x release loads extensions built for C API version v1.2.0 (the version string passed to duckdb_rs_extension_api_init, quack_rs::DUCKDB_API_VERSION). DuckDB 1.5.6 declares C API version v1.5.6 and still accepts v1.2.0 extensions. CI loads the example extension into DuckDB 1.4.4, 1.5.0, 1.5.5 and the latest release. Using a range rather than an exact pin means:

  • Extension authors can choose which libduckdb-sys release to build against (for example =1.4.4 for DuckDB 1.4.4, or ~1.10505.0 for DuckDB 1.5.5, which libduckdb-sys numbers 1.10505.x) and still resolve against quack-rs
  • quack-rs itself doesn't force a DuckDB downgrade on users

The <2 upper bound is equally intentional: it prevents silent adoption of a future major release that may introduce breaking C API changes. Upgrading beyond the 1.x band requires an explicit quack-rs release that audits the new C API surface.

For your own extension's Cargo.toml: an extension that uses only the stable C API (the default quack-rs features) can keep the same ">=1.4.4, <2" range; this is what scaffold::generate_scaffold writes. An extension that enables the duckdb-1-5* features must pin libduckdb-sys to the bindings of the one DuckDB release it is stamped for (the scaffold writes ~1.10505.0 for v1.5.5). See ABI Compatibility.


Binary compatibility

Which DuckDB releases accept an extension binary depends on the ABI type stamped into its metadata footer:

  • A C_STRUCT binary targeting C API v1.2.0 (the default, stable API only) loads into every DuckDB release whose C API version is at least v1.2.0 — all of 1.4.x and 1.5.x — on the platform it was built for
  • A C_STRUCT_UNSTABLE binary loads only into the exact DuckDB release it names. Stamp builds that use the duckdb-1-5* features this way, so DuckDB itself refuses a mismatched release; quack-rs's runtime layout check is the backstop when the stamp is missing
  • DuckDB checks the footer's platform and version fields at load time and refuses a mismatch
  • Core and community extensions are signed; a binary you build locally is not
  • To load an unsigned extension during development, start DuckDB with allow_unsigned_extensions enabled (duckdb -unsigned in the CLI); the setting cannot be changed on a running database
  • The community extension CI builds and signs each extension for every supported platform

The Entry Point

Every DuckDB loadable extension exports one C-callable entry-point function, which DuckDB calls when it loads the extension. quack-rs generates it with the entry_point_v2! or entry_point! macro, or you can write it by hand around init_extension.


Option A: entry_point_v2! with Connection (recommended)

Added in v0.4.0.

The entry_point_v2! macro gives your closure a &Connection instead of a raw duckdb_connection. The Connection type implements the Registrar trait, which registers every kind of function, macro, cast, copy function and config option; replacement scans, which belong to the database rather than a connection, are Connection's own methods:

use quack_rs::entry_point_v2;
use quack_rs::connection::{Connection, Registrar};
use quack_rs::error::ExtensionError;

unsafe fn register(con: &Connection) -> Result<(), ExtensionError> {
    unsafe {
        con.register_scalar(/* ScalarFunctionBuilder */)?;
        con.register_aggregate(/* AggregateFunctionBuilder */)?;
        con.register_table(/* TableFunctionBuilder */)?;
        con.register_cast(/* CastFunctionBuilder */)?;
        con.register_scalar_set(/* ScalarFunctionSetBuilder */)?;
        con.register_aggregate_set(/* AggregateFunctionSetBuilder */)?;
        con.register_sql_macro(/* SqlMacro */)?;
        con.register_replacement_scan(/* callback, data, destructor */);
        // con.register_copy_function(/* CopyFunctionBuilder */)?;  // requires duckdb-1-5
    }
    Ok(())
}

entry_point_v2!(my_extension_init_c_api, |con| unsafe { register(con) });

(The register_* arguments above are placeholders, so that block is not compilable as written.) The macro emits the equivalent of the following; the real expansion also evaluates the policy and closure arguments inside a panic guard:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_extension_access, duckdb_extension_info};
use quack_rs::connection::Connection;
use quack_rs::error::ExtensionError;
unsafe fn register(_con: &Connection) -> Result<(), ExtensionError> { Ok(()) }
#[no_mangle]
pub unsafe extern "C" fn my_extension_init_c_api(
    info: duckdb_extension_info,
    access: *const duckdb_extension_access,
) -> bool {
    unsafe {
        quack_rs::entry_point::init_extension_v2_with_policy(
            info, access, quack_rs::DUCKDB_API_VERSION,
            quack_rs::abi::AbiPolicy::Strict,
            |con| unsafe { register(con) },
        )
    }
}
}

entry_point_v2!(name, policy, |con| ...) passes a different AbiPolicy.

Pass the full symbol name to the macro. The symbol {name}_init_c_api must match the name field in description.yml and the [lib] name in Cargo.toml.

Why Connection over raw duckdb_connection?

entry_point! (raw)entry_point_v2! (Connection)
Receivesduckdb_connection&Connection
RegistrationCall builders' .register(con)Call con.register_*()
Type safetyRaw pointerTyped wrapper around the connection and database handles
Replacement scansNeed the database handle, which the closure does not receivecon.register_replacement_scan*()

Option B: The entry_point! macro

The original macro passes a raw duckdb_connection to your closure. It performs the same initialization, but you pass the connection to each builder's .register():

#![allow(unused)]
fn main() {
use quack_rs::entry_point;
use quack_rs::error::ExtensionError;

fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> {
    // Register each function on `con`, e.g. `unsafe { builder.register(con)? };`
    let _ = con;
    Ok(())
}

entry_point!(my_extension_init_c_api, |con| register(con));
}

Option C: Manual entry point

If you need full control (e.g., multiple registration functions, conditional logic):

#![allow(unused)]
fn main() {
use quack_rs::entry_point::init_extension;
use libduckdb_sys::{duckdb_extension_info, duckdb_extension_access};
use libduckdb_sys::duckdb_connection;
use quack_rs::error::ExtensionError;
fn register_scalar_functions(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
fn register_aggregate_functions(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
fn register_sql_macros(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }

#[no_mangle]
pub unsafe extern "C" fn my_extension_init_c_api(
    info: duckdb_extension_info,
    access: *const duckdb_extension_access,
) -> bool {
    unsafe {
        init_extension(info, access, quack_rs::DUCKDB_API_VERSION, |con| {
            register_scalar_functions(con)?;
            register_aggregate_functions(con)?;
            register_sql_macros(con)?;
            Ok(())
        })
    }
}
}

What init_extension does

flowchart TD
    A["<b>1. duckdb_rs_extension_api_init</b>(info, access, version)<br/>Fills the global AtomicPtr dispatch table"]
    L["<b>2. ABI layout check</b><br/>Applies the AbiPolicy (Strict by default)"]
    B["<b>3. access.get_database</b>(info)<br/>Returns the duckdb_database handle"]
    C["<b>4. duckdb_connect</b>(db, &amp;mut con)<br/>Opens a connection for function registration"]
    D["<b>5. register</b>(con) ← your closure<br/>A panic becomes an error"]
    E["<b>6. duckdb_disconnect</b>(&amp;mut con)<br/>Always runs, even if registration failed"]
    F{Error?}
    G["return <b>true</b>"]
    H["return <b>false</b><br/>error reported via access.set_error"]

    A --> L --> B --> C --> D --> E --> F
    L -->|refused| H
    F -->|no| G
    F -->|yes| H

    style G fill:#1c3b1c,stroke:#4a9e4a,color:#c8ecc8
    style H fill:#3b1c1c,stroke:#9e4a4a,color:#ecc8c8

An error from any step, including an Err or a panic from your closure in step 5, is reported to DuckDB via access.set_error, and the function returns false. DuckDB then fails the LOAD with that message. (When DuckDB itself detects the failure, as when get_database returns null, quack-rs returns false without overwriting DuckDB's own message.) The layout check in step 2 only runs when a duckdb-1-5* feature is enabled; see ABI Compatibility.

Registration is not transactional: functions registered before a failure stay registered for the life of the database. Do fallible setup work (reading configuration, building lookup tables) before the first register call.


The C API version constant

#![allow(unused)]
fn main() {
pub const DUCKDB_API_VERSION: &str = "v1.2.0";
}

Pitfall P2: This is the C API version, not the DuckDB release version. DuckDB 1.4.x and 1.5.0–1.5.5 declare C API version v1.2.0; 1.5.6 declares v1.5.6 and still loads extensions that target v1.2.0 (DuckDB accepts any C API version up to its own). A binary stamped with a DuckDB release version instead (-dv v1.5.5) is refused at LOAD by every DuckDB whose C API version is lower; quack-rs's append_metadata rejects such a value for C_STRUCT up front. See Pitfall P2.


No panics in the entry point

init_extension never panics. All error paths use Result and ?. If your registration closure returns Err or panics, the message is reported to DuckDB via access.set_error and the LOAD fails with an error instead of aborting the process. Catching a panic requires panic = "unwind" in your release profile (the scaffold's default).

Never use unwrap() or expect() in FFI callbacks. See Pitfall L3.

Error Handling

quack-rs reports every error through one type, ExtensionError, and the ExtResult<T> alias. This page covers creating and propagating errors, how they reach DuckDB, and why extension code must not panic.


ExtensionError

#![allow(unused)]
fn main() {
use quack_rs::error::{ExtensionError, ExtResult};
let (name, code) = ("my_fn", 1);
let some_std_error = std::fmt::Error;

// From a string literal
let e = ExtensionError::from("something went wrong");

// From a format string
let e = ExtensionError::new(format!("failed to register '{}': code {}", name, code));

// Wrapping another error
let e = ExtensionError::from_error(some_std_error);
}

ExtensionError implements:

  • std::error::Error
  • Display, Debug, Clone, PartialEq, Eq
  • From<&str>, From<String>, From<Box<dyn Error>>, From<Box<dyn Error + Send + Sync>>
  • From<std::io::Error>, From<std::ffi::NulError>, From<std::fmt::Error>

The From<std::io::Error> impl lets extensions that allocate runtime resources during initialization (a tokio runtime, for example) use ? without .map_err():

#![allow(unused)]
fn main() {
use quack_rs::connection::Connection;
use quack_rs::error::ExtensionError;
// A stand-in with tokio's signature: `Runtime::new() -> std::io::Result<Runtime>`.
mod tokio { pub mod runtime { pub struct Runtime;
    impl Runtime { pub fn new() -> std::io::Result<Self> { Ok(Runtime) } } } }
fn register_all(con: &Connection) -> Result<(), ExtensionError> {
    let _rt = tokio::runtime::Runtime::new()?; // ← io::Error → ExtensionError
    // ... register functions ...
    Ok(())
}
}

ExtResult<T>

A type alias for Result<T, ExtensionError>, used throughout the SDK:

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
pub type ExtResult<T> = Result<T, ExtensionError>;
}

Propagating errors with ?

In your registration function:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::prelude::*;
unsafe extern "C" fn my_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        ScalarFunctionBuilder::try_new("my_fn")?
            .param(TypeId::BigInt)
            .returns(TypeId::BigInt)
            .function(my_fn)
            .register(con)?;   // ← ? propagates registration errors

        SqlMacro::scalar("my_macro", &["x"], "x + 1")?
            .register(con)?;

        Ok(())
    }
}
}

If any registration call fails, ? returns the error from register, which init_extension then reports to DuckDB via access.set_error.


Error reporting to DuckDB

init_extension converts the ExtensionError to a C string for DuckDB's set_error callback, following the same rule as ExtensionError::to_c_string: a C string cannot hold a NUL byte, so each one is replaced with ? and the rest of the message is kept:

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;

let err = ExtensionError::new("bad\0input");
assert_eq!(err.to_c_string().to_str(), Ok("bad?input"));
}

DuckDB surfaces this string to the user as the extension load error.


No panics

The central rule of DuckDB extension development:

Never unwrap(), expect(), or panic!() in any code path that DuckDB may call.

A panic cannot unwind out of an extern "C" function: since Rust 1.81 the runtime aborts the process instead, taking the user's DuckDB session with it. quack-rs's callback macros and typed builders catch panics and turn them into SQL errors (which requires panic = "unwind" in the release profile), but that is a safety net for bugs, not an error-handling strategy.

Safe patterns

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_aggregate_state;
use quack_rs::aggregate::{AggregateState, FfiState};
use quack_rs::error::ExtensionError;
#[derive(Default)] struct MyState { count: u64 }
impl AggregateState for MyState {}
fn some_fallible_call() -> Result<u64, ExtensionError> { Ok(1) }
unsafe fn demo(state_ptr: duckdb_aggregate_state, maybe_count: Option<u64>)
    -> Result<(), ExtensionError> {
// ✅ Use Option methods
if let Some(s) = FfiState::<MyState>::with_state_mut(state_ptr) {
    s.count += 1;
}

// ✅ Use Result and ?
let value = some_fallible_call()?;

// ✅ Use unwrap_or / unwrap_or_else / map
let count = maybe_count.unwrap_or(0);

// ❌ Never in FFI callbacks
let s = FfiState::<MyState>::with_state_mut(state_ptr).unwrap(); // panics if None
Ok(())
}
}

In init_extension

init_extension reports every error via set_error and does not panic itself. It also runs your registration closure under catch_unwind, so a panic there becomes a load error rather than a process abort (with panic = "unwind").

Type System

quack-rs describes DuckDB column types with two types: TypeId, a plain enum of DuckDB's type ids, and LogicalType, an owned handle for full type descriptions such as DECIMAL(18, 3), LIST(VARCHAR) or a STRUCT. This page also maps each DuckDB type to its Rust type and vector read/write methods.


TypeId

TypeId is an enum covering DuckDB's column types (the GEOMETRY and VARIANT types added in DuckDB 1.5.x are exposed behind the duckdb-1-5-3 feature — see Known Limitations). The list below names the variants; it is a listing, not compilable code:

use quack_rs::types::TypeId;

TypeId::Boolean
TypeId::TinyInt     // i8
TypeId::SmallInt    // i16
TypeId::Integer     // i32
TypeId::BigInt      // i64
TypeId::UTinyInt    // u8
TypeId::USmallInt   // u16
TypeId::UInteger    // u32
TypeId::UBigInt     // u64
TypeId::HugeInt     // i128
TypeId::UHugeInt    // u128
TypeId::Float       // f32
TypeId::Double      // f64
TypeId::Timestamp
TypeId::TimestampTz
TypeId::TimestampS
TypeId::TimestampMs
TypeId::TimestampNs
TypeId::Date
TypeId::Time
TypeId::TimeTz
TypeId::Interval
TypeId::Varchar
TypeId::Blob
TypeId::Decimal
TypeId::Enum
TypeId::List
TypeId::Struct
TypeId::Map
TypeId::Uuid
TypeId::Union
TypeId::Bit
TypeId::Array
TypeId::TimeNs
TypeId::Any
TypeId::Varint           // SQL name BIGNUM (VARINT before DuckDB 1.4)
TypeId::SqlNull
TypeId::IntegerLiteral
TypeId::StringLiteral
TypeId::Geometry         // duckdb-1-5-3
TypeId::Variant          // duckdb-1-5-3

TypeId implements Copy, Clone, Debug, PartialEq, Eq, Hash and Display.

SQL name

#![allow(unused)]
fn main() {
use quack_rs::types::TypeId;
assert_eq!(TypeId::BigInt.sql_name(), "BIGINT");
assert_eq!(TypeId::Varchar.sql_name(), "VARCHAR");
assert_eq!(format!("{}", TypeId::Timestamp), "TIMESTAMP");
}

DuckDB constant

TypeId::to_duckdb_type() returns the DUCKDB_TYPE_* integer constant from libduckdb-sys. You rarely need this directly — it's called internally by LogicalType::new.

Reverse conversion

TypeId::from_duckdb_type(raw) converts a raw DUCKDB_TYPE constant back into a TypeId. It panics if the value does not match any known constant; TypeId::try_from_duckdb_type(raw) returns None instead.

#![allow(unused)]
fn main() {
use quack_rs::types::TypeId;

let type_id = TypeId::from_duckdb_type(libduckdb_sys::DUCKDB_TYPE_DUCKDB_TYPE_BIGINT);
assert_eq!(type_id, TypeId::BigInt);
}

LogicalType

LogicalType is an RAII wrapper around DuckDB's duckdb_logical_type. Creating one calls into DuckDB, so this block is compiled but not run:

#![allow(unused)]
fn main() {
use quack_rs::types::{LogicalType, TypeId};

let lt = LogicalType::new(TypeId::Varchar);
// lt.as_raw() returns the duckdb_logical_type pointer
// Drop calls duckdb_destroy_logical_type automatically
}

Pitfall L7: duckdb_create_logical_type allocates memory that must be freed with duckdb_destroy_logical_type. LogicalType's Drop implementation does this automatically, preventing the memory leak that occurs when calling the DuckDB C API directly. See Pitfall L7.

For a type a TypeId fully describes, pass the TypeId to a builder's param or returns method; the builder creates and destroys the LogicalType internally. For a parameterized type (DECIMAL, LIST, MAP, STRUCT, UNION, ENUM, ARRAY), build a LogicalType and pass it to param_logical or returns_logical.

Constructors

ConstructorCreates
LogicalType::new(type_id)Simple type from a TypeId
LogicalType::from_raw(ptr)Takes ownership of a raw duckdb_logical_type handle (unsafe)
LogicalType::decimal(width, scale)DECIMAL(width, scale)
LogicalType::list(element_type)LIST<element_type> from a TypeId
LogicalType::list_from_logical(element)LIST<element> from an existing LogicalType
LogicalType::map(key, value)MAP<key, value> from TypeIds
LogicalType::map_from_logical(key, value)MAP<key, value> from existing LogicalTypes
LogicalType::struct_type(fields)STRUCT from &[(&str, TypeId)]
LogicalType::struct_type_from_logical(fields)STRUCT from &[(&str, LogicalType)]
LogicalType::union_type(members)UNION from &[(&str, TypeId)]
LogicalType::union_type_from_logical(members)UNION from &[(&str, LogicalType)]
LogicalType::enum_type(members)ENUM from &[&str]
LogicalType::array(element_type, size)ARRAY<element_type>[size] from a TypeId
LogicalType::array_from_logical(element, size)ARRAY<element>[size] from an existing LogicalType

Every constructor except from_raw panics on invalid input (for example a composite TypeId passed to new, or a DECIMAL width outside 1–38) and has a try_* counterpart (try_new, try_decimal, try_struct_type, …) that returns Result<LogicalType, LogicalTypeError> instead. UNION types accept at most MAX_UNION_MEMBERS (255) members.

Introspection methods

All introspection methods are unsafe: they call into DuckDB, so the handle must be valid and the C API initialized.

MethodReturnsApplicable to
get_type_id()TypeIdAny
get_alias()Option<String>Any
set_alias(alias)()Any
decimal_width()u8DECIMAL
decimal_scale()u8DECIMAL
decimal_internal_type()TypeIdDECIMAL
enum_internal_type()TypeIdENUM
enum_dictionary_size()u32ENUM
enum_dictionary_value(index)StringENUM
list_child_type()LogicalTypeLIST
map_key_type()LogicalTypeMAP
map_value_type()LogicalTypeMAP
struct_child_count()u64STRUCT
struct_child_name(index)StringSTRUCT
struct_child_type(index)LogicalTypeSTRUCT
union_member_count()u64UNION
union_member_name(index)StringUNION
union_member_type(index)LogicalTypeUNION
array_size()u64ARRAY
array_child_type()LogicalTypeARRAY

Rust type ↔ DuckDB type mapping

When reading from or writing to vectors, use the corresponding VectorReader/VectorWriter method. The most common types:

DuckDB typeTypeIdReader methodWriter method
BOOLEANBooleanread_boolwrite_bool
TINYINTTinyIntread_i8write_i8
SMALLINTSmallIntread_i16write_i16
INTEGERIntegerread_i32write_i32
BIGINTBigIntread_i64write_i64
UTINYINTUTinyIntread_u8write_u8
USMALLINTUSmallIntread_u16write_u16
UINTEGERUIntegerread_u32write_u32
UBIGINTUBigIntread_u64write_u64
FLOATFloatread_f32write_f32
DOUBLEDoubleread_f64write_f64
HUGEINTHugeIntread_i128write_i128
UHUGEINTUHugeIntread_u128write_u128
VARCHARVarcharread_strwrite_varchar
BLOBBlobread_blobwrite_blob
UUIDUuidread_uuidwrite_uuid
DATEDateread_datewrite_date
TIMETimeread_timewrite_time
TIMESTAMPTimestampread_timestampwrite_timestamp
INTERVALIntervalread_intervalwrite_interval

The timestamp variants (TIMESTAMP_S, _MS, _NS, TIMESTAMPTZ), TIMETZ and DECIMAL have matching read_*/write_* methods as well.

NULLs are handled separately — see NULL Handling & Strings.

ABI Compatibility

This page explains DuckDB C Extension API ABI compatibility: which DuckDB releases a quack-rs extension binary can load into, and how quack-rs guards against a layout mismatch. A loadable extension does not link against DuckDB's symbols. DuckDB hands it a pointer to a duckdb_ext_api_v1 struct — an array of function pointers — and the extension calls through it.

Two things have to agree about that struct's layout: the header your extension was compiled against (whichever libduckdb-sys version Cargo resolved), and the DuckDB binary that is loading it. When they disagree, every call lands on the wrong slot.

The struct has two halves

RegionSlotsGuarantee
Stable0 .. 357Frozen since DuckDB v1.2.0 — same slots, same order, same signatures in every release through v1.5.6 (two slots, 114 and 138, were renamed varint → bignum in v1.4.0 with an identical struct layout)
Unstable357 ..DuckDB inserts new entries in the middle, shifting every later slot

The stable prefix is what makes "build once, load anywhere" possible. The unstable tail is not append-only:

DuckDBTotal slotsWhat moved
v1.2.0 – v1.2.2408baseline
v1.3.0 – v1.3.2428appended
v1.4.0 – v1.4.5459duckdb_create_varint → duckdb_create_bignum; appended
v1.5.0 – v1.5.1545duckdb_appender_clear inserted at slot 410
v1.5.2 – v1.5.6546duckdb_geometry_type_get_crs inserted at slot 493

DuckDB v1.5.6 declares all 546 slots stable for extensions that target C API v1.5.6 (earlier releases declared only the first 357 stable). The layout itself is unchanged from v1.5.2, and quack-rs targets C API v1.2.0, so the guard below still applies.

Every family since v1.2 changed the unstable tail, and twice (v1.5.0 and v1.5.2) an insertion in the middle shifted every later slot.

Which half are you using?

Everything quack-rs exposes by default lives in the stable prefix: scalar, aggregate, table and cast functions, vectors, data chunks, values, SQL macros, replacement scans, the query API and the datetime conversions.

The duckdb-1-5, duckdb-1-5-3 and duckdb-1-5-4 features wrap the unstable half — 130 of its 189 functions, covering scalar bind/init, copy functions in both directions, the Arrow C Data Interface bridge, catalog access, ErrorData, FileSystem, Expression, SelectionVector, config options, table descriptions, TIME_NS values and the client context.

abi::uses_unstable_api() reports which half a build uses. The book's own test build enables duckdb-1-5-4 (and therefore duckdb-1-5), so this block is compiled there but not run:

#![allow(unused)]
fn main() {
use quack_rs::abi;

// false unless `duckdb-1-5` is enabled.
assert!(!abi::uses_unstable_api());
}

Why DuckDB does not catch this for you

DuckDB validates the ABI metadata in your extension's footer:

ABI typeVersion field (-dv) meansAccepted by
C_STRUCTthe C API version (v1.2.0)any DuckDB whose C API version is at least that — then handed the whole struct, unstable region included
C_STRUCT_UNSTABLEan exact DuckDB release (v1.5.5)that release only

So a C_STRUCT binary that touches the unstable region loads happily into the wrong DuckDB and then mis-dispatches. DuckDB's own extension-template-c says:

WARNING: When set to 1, the duckdb_extension.h from the TARGET_DUCKDB_VERSION must be used, using any other version of the header is unsafe.

Built against v1.5.0's headers and loaded into v1.5.5, an extension calling ClientContext::from_connection invokes duckdb_destroy_client_context on a duckdb_connection. In practice:

double free or corruption (out)
Aborted (core dumped)

What quack-rs does

Two layers.

Build metadata. If you enable duckdb-1-5, stamp the binary C_STRUCT_UNSTABLE with the DuckDB release you built against, so DuckDB refuses the wrong engine at install time. With extension-ci-tools, set in the Makefile:

USE_UNSTABLE_C_API=1
TARGET_DUCKDB_VERSION=v1.5.5

generate_scaffold writes this pairing from a ScaffoldConfig, and rejects a target_duckdb_version that does not match use_unstable_c_api:

#![allow(unused)]
fn main() {
use quack_rs::scaffold::ScaffoldConfig;

let config = ScaffoldConfig {
    name: "my_ext".to_string(),
    use_unstable_c_api: true,
    target_duckdb_version: "v1.5.5".to_string(),
    ..ScaffoldConfig::default()
};
}

Runtime guard. When a duckdb-1-5* feature is enabled, abi::check compares the compiled-in slot count against the layout the running engine uses, resolved from duckdb_library_version() — which lives at stable slot 7 and is therefore always dispatched correctly. (Without those features the check reports StableOnly and always passes.) The entry point applies it according to an AbiPolicy:

PolicyBehaviour
Strict (default)Refuse to load, with a message naming both layouts and the fix
AllowUnknownEngineRefuse a layout the table knows is different; allow a release the table has no entry for
WarnPrint the diagnostic to stderr, then load anyway (set_error would fail the load)
TrustSkip the check

AllowUnknownEngine and Trust are only as safe as your knowledge of the engine: calling into the unstable region of a layout the extension was not built for is undefined behaviour. Before DuckDB v1.4.5 was in quack-rs's table, a v1.5.5 build loaded into v1.4.5 under AllowUnknownEngine and segfaulted.

The two lines below are alternatives: both define the same exported symbol, so they do not compile together.

use quack_rs::abi::AbiPolicy;

// Default: Strict.
quack_rs::entry_point!(my_ext_init_c_api, register);

// Explicit — appropriate when the binary is stamped C_STRUCT_UNSTABLE, because
// DuckDB already refuses to load it into the wrong release.
quack_rs::entry_point!(my_ext_init_c_api, AbiPolicy::Trust, register);

A refused load looks like this (a v1.5.0 build loaded into DuckDB v1.5.5):

DuckDB C extension API layout mismatch: this extension was built against a
duckdb_ext_api_v1 with 545 slots, but DuckDB v1.5.5 provides 546. The extension
uses the unstable region of the C API (quack-rs feature `duckdb-1-5`), whose slot
indices differ between these releases, so loading it would dispatch to the wrong
functions. Rebuild the extension against DuckDB v1.5.5. Stamp every build that
uses the unstable region with `--abi-type C_STRUCT_UNSTABLE --duckdb-version <the
DuckDB release it was built against>` (or `USE_UNSTABLE_C_API=1` with
extension-ci-tools) so DuckDB refuses a mismatched binary at install time.

For an engine older than v1.5.0 the advice changes: such an engine cannot run a duckdb-1-5 build at all, so the message says to build a variant without those features or to upgrade DuckDB.

Unknown DuckDB versions

Strict also refuses a DuckDB release quack-rs has no verified layout for — a newer release, or a -dev build. That is deliberate: DuckDB changed the unstable region in every minor release from 1.3.0 on, and in the 1.5.2 patch release, so "unknown" is not evidence of "compatible". The refusal lists the fixes, best first:

  1. rebuild against the release you are targeting and set QUACK_RS_TARGET_DUCKDB_VERSION to it, which turns the check into a positive match without waiting for a quack-rs release;
  2. upgrade quack-rs to a version whose layout table lists that release;
  3. if the extension does not need the unstable region, build it without the duckdb-1-5 features: the stable prefix is laid out identically in every DuckDB since v1.2.0.

It never suggests AllowUnknownEngine or Trust, for the reason above.

scripts/check-abi-table.py re-derives quack-rs's layout table from every upstream release header and runs in CI, so the table tracks DuckDB.

Scalar Functions

A DuckDB scalar function returns one output value per input row, like the built-in length(), upper() or sin(). This page shows how to write one in Rust with quack-rs: as a raw callback registered through ScalarFunctionBuilder, as a safe closure with map1 / map2, or as a set of overloads with ScalarFunctionSetBuilder.


Function signature

DuckDB calls your scalar function once per data chunk (not once per row). The signature is:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn my_fn(
    info: duckdb_function_info,     // function metadata (rarely needed)
    input: duckdb_data_chunk,       // input data — one or more columns
    output: duckdb_vector,          // output vector — one value per input row
)
{}
}

Inside the function, you:

  1. Create a VectorReader for each input column
  2. Create a VectorWriter for the output
  3. Loop over rows, checking for NULLs and transforming values

A panic that unwinds out of an extern "C" function aborts the process. Generate the function with scalar_callback!, which reports a panic as a SQL error, or use the closure constructors below (Error Handling).


Registration

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn my_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
use quack_rs::scalar::ScalarFunctionBuilder;
use quack_rs::types::TypeId;

unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        ScalarFunctionBuilder::new("my_fn")
            .param(TypeId::BigInt)      // first parameter type
            .param(TypeId::BigInt)      // second parameter type (if any)
            .returns(TypeId::BigInt)    // return type
            .function(my_fn)            // callback
            .register(con)?;
    }
    Ok(())
}
}

Before calling duckdb_register_scalar_function, register checks that returns and function are set and that no parameter or return type is a bare composite TypeId (List, Decimal, …; use the *_logical methods for those). It also refuses a signature that duckdb_functions() already lists under the same name, built-ins included: DuckDB would otherwise silently replace that overload for every connection, or make every call ambiguous. If DuckDB reports failure, register returns Err.

Validated registration

For user-configurable function names (e.g., from a config file), use try_new:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn my_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe fn demo(con: duckdb_connection, name: &str) -> Result<(), ExtensionError> {
ScalarFunctionBuilder::try_new(name)?   // validates name before building
    .param(TypeId::Varchar)
    .returns(TypeId::Varchar)
    .function(my_fn)
    .register(con)?;
Ok(())
}
}

try_new validates the name as an unquoted SQL identifier: [A-Za-z_][A-Za-z0-9_]*, at most 256 characters, and not a DuckDB keyword that cannot be called as a function unquoted (order, coalesce). Mixed case is allowed (DuckDB itself ships formatReadableSize). new does not validate: it only panics if the name contains an interior NUL byte, so use it for names known at compile time.

Closures

ScalarFunctionBuilder::map1 / map2 / map1_str / map2_str / map1_opt / map2_opt build the whole function from a Rust closure and return a TypedScalarFunctionBuilder. Its signature is fixed by the closure's types, so it deliberately offers only name(), volatile() and register(con) — no returns, param, function, extra_info, bind or init, any of which could make DuckDB hand the closure vectors of a different type than it reads and writes. Register it with register(con), or through a Registrar with register_typed_scalar.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
ScalarFunctionBuilder::map1("double_it", |x: i64| x * 2)?.register(con)?;
Ok(())
}
}

Complete example: double_it(BIGINT) → BIGINT

#![allow(unused)]
fn main() {
use quack_rs::prelude::*;
fn live_connection() -> libduckdb_sys::duckdb_connection {
    std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
    let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess);
    }
    con
}
/// First column of the first row, as BIGINT; `None` for NULL.
fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    let reader = unsafe { chunk.reader(0) };
    unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) }
}
use quack_rs::vector::{VectorReader, VectorWriter};
use libduckdb_sys::{duckdb_function_info, duckdb_data_chunk, duckdb_vector};

unsafe extern "C" fn double_it(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    // SAFETY: DuckDB provides valid chunk and vector pointers.
    let reader = unsafe { VectorReader::new(input, 0) };   // column 0
    let mut writer = unsafe { VectorWriter::new(output) };
    let row_count = reader.row_count();

    for row in 0..row_count {
        if unsafe { !reader.is_valid(row) } {
            // NULL input → NULL output
            // SAFETY: row < row_count, writer is valid.
            unsafe { writer.set_null(row) };
            continue;
        }
        let value = unsafe { reader.read_i64(row) };
        unsafe { writer.write_i64(row, value * 2) };
    }
}
let con = live_connection();
unsafe {
    ScalarFunctionBuilder::new("double_it").param(TypeId::BigInt).returns(TypeId::BigInt)
        .function(double_it).register(con).unwrap();
}
assert_eq!(query_i64(con, "SELECT double_it(21)"), Some(42));
// From a column, not a literal: `double_it(NULL::BIGINT)` is constant-folded.
assert_eq!(query_i64(con, "SELECT double_it(i) FROM (VALUES (NULL::BIGINT)) t(i)"), None);
}

Multi-parameter example: add(BIGINT, BIGINT) → BIGINT

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn add(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let col0 = unsafe { VectorReader::new(input, 0) };  // first param
    let col1 = unsafe { VectorReader::new(input, 1) };  // second param
    let mut writer = unsafe { VectorWriter::new(output) };

    for row in 0..col0.row_count() {
        if unsafe { !col0.is_valid(row) || !col1.is_valid(row) } {
            unsafe { writer.set_null(row) };
            continue;
        }
        let a = unsafe { col0.read_i64(row) };
        let b = unsafe { col1.read_i64(row) };
        unsafe { writer.write_i64(row, a + b) };
    }
}
}

VARCHAR example: shout(VARCHAR) → VARCHAR

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn shout(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };

    for row in 0..reader.row_count() {
        if unsafe { !reader.is_valid(row) } {
            unsafe { writer.set_null(row) };
            continue;
        }
        let s = unsafe { reader.read_str(row) };
        let upper = s.to_uppercase();
        unsafe { writer.write_varchar(row, &upper) };
    }
}
}

Overloading with function sets

If your function accepts different parameter types or arities, use ScalarFunctionSetBuilder to register multiple overloads under a single name:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn add_ints(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe extern "C" fn add_doubles(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
use quack_rs::scalar::{ScalarFunctionSetBuilder, ScalarOverloadBuilder};
use quack_rs::types::TypeId;

unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        ScalarFunctionSetBuilder::new("my_add")
            .overload(
                ScalarOverloadBuilder::new()
                    .param(TypeId::Integer).param(TypeId::Integer)
                    .returns(TypeId::Integer)
                    .function(add_ints)
            )
            .overload(
                ScalarOverloadBuilder::new()
                    .param(TypeId::Double).param(TypeId::Double)
                    .returns(TypeId::Double)
                    .function(add_doubles)
            )
            .register(con)?;
    }
    Ok(())
}
}

Like AggregateFunctionSetBuilder, this builder calls duckdb_scalar_function_set_name on every individual function before adding it to the set (Pitfall L6).

ScalarOverloadBuilder has the same per-function settings as ScalarFunctionBuilder, applied to that overload only: null_handling, extra_info, varargs / varargs_logical, volatile, and (DuckDB 1.5+) bind / init. register checks every overload for a return type and a callback before it creates any DuckDB handle, and the error names the overload's index. It also refuses two overloads that accept the same call: the same argument types, or, with varargs, the same types at some argument count. f(BIGINT) and f(BIGINT, BIGINT...) both accept f(1), so they cannot share a set; DuckDB would accept the set and then fail every such call as ambiguous. An overload that matches an existing function's signature is refused as for ScalarFunctionBuilder.


NULL handling

Your callback receives NULL rows whatever the setting: under the default, DefaultNullHandling, it promises NULL-in-NULL-out and must write the NULLs itself — call chunk.propagate_nulls(&mut writer) at the end, or use the closure constructors (map1, map2 and their _str forms), which do it for you (Pitfall L8, NULL handling). A function that means to return non-NULL for NULL input (e.g., a COALESCE-like function) sets SpecialNullHandling:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn my_coalesce_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
use quack_rs::types::NullHandling;

ScalarFunctionBuilder::new("coalesce_custom")
    .param(TypeId::BigInt)
    .returns(TypeId::BigInt)
    .null_handling(NullHandling::SpecialNullHandling)
    .function(my_coalesce_fn)
    .register(con)?;
Ok(())
}
}

With SpecialNullHandling, the callback must check VectorReader::is_valid(row) itself and decide what each NULL input produces. The closure constructors map1_opt / map2_opt register SpecialNullHandling for you and pass the closure an Option (None for NULL); returning None writes NULL.


Complex parameter and return types

For scalar functions that accept or return parameterized types like LIST(BIGINT), use param_logical and returns_logical:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn flatten_list_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
use quack_rs::scalar::ScalarFunctionBuilder;
use quack_rs::types::{LogicalType, TypeId};

ScalarFunctionBuilder::new("flatten_list")
    .param_logical(LogicalType::list(TypeId::BigInt))  // LIST(BIGINT) input
    .returns(TypeId::BigInt)
    .function(flatten_list_fn)
    .register(con)?;
Ok(())
}
}

These methods are also available on ScalarOverloadBuilder for function sets:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn my_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
fn demo() {
let _ =
ScalarOverloadBuilder::new()
    .param(TypeId::Varchar)
    .returns_logical(LogicalType::list(TypeId::Timestamp))  // LIST(TIMESTAMP) output
    .function(my_fn)
;
}
}

Key points

  • VectorReader::new(input, column_index) — the column index is zero-based
  • Always check is_valid(row) before reading — skipping this reads garbage for NULL rows, and for VARCHAR/BLOB can follow a stale pointer
  • set_null must be called for NULL outputs — it calls ensure_validity_writable automatically (Pitfall L4)
  • read_bool returns bool — handles DuckDB's non-0/1 boolean bytes correctly (Pitfall L5)
  • read_str handles both inline and pointer string formats automatically (Pitfall P7)

Varargs and volatility

These ScalarFunctionBuilder methods map to functions in DuckDB's stable C API (v1.2.0), so they need no feature flag and work on DuckDB 1.4.x and 1.5.x alike:

varargs(type_id: TypeId)

Declares that the function accepts a variable number of trailing arguments, all of the given TypeId. Maps to duckdb_scalar_function_set_varargs. A composite TypeId such as List or Decimal makes register return an error naming the varargs slot; use varargs_logical for those.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn concat_all_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
ScalarFunctionBuilder::new("concat_all")
    .varargs(TypeId::Varchar)
    .returns(TypeId::Varchar)
    .function(concat_all_fn)
    .register(con)?;
Ok(())
}
}

varargs_logical(logical_type: LogicalType)

Like varargs, but accepts a LogicalType for parameterized variadic arguments. Maps to duckdb_scalar_function_set_varargs.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn merge_lists_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
ScalarFunctionBuilder::new("merge_lists")
    .varargs_logical(LogicalType::list(TypeId::BigInt))
    .returns_logical(LogicalType::list(TypeId::BigInt))
    .function(merge_lists_fn)
    .register(con)?;
Ok(())
}
}

volatile()

Marks the function as volatile: DuckDB re-evaluates it for every row, even when its arguments are constant (as for random()), and never merges two identical calls in one query. Without it, DuckDB may evaluate a call with constant arguments only once. Maps to duckdb_scalar_function_set_volatile. Also available on the closure-built TypedScalarFunctionBuilder.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn random_int_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
ScalarFunctionBuilder::new("random_int")
    .returns(TypeId::Integer)
    .volatile()
    .function(random_int_fn)
    .register(con)?;
Ok(())
}
}

DuckDB 1.5.0 additions (duckdb-1-5)

The following ScalarFunctionBuilder methods are available when the duckdb-1-5 feature is enabled:

bind(bind_fn)

Sets a custom bind callback that runs at plan time. Use this to inspect argument types and set the return type dynamically. Maps to duckdb_scalar_function_set_bind.

The callback takes a RawScalarBindInfo (wrap it with ScalarBindInfo::new), not a bare duckdb_bind_info: DuckDB passes a table function's bind callback a different, larger struct, and a callback written for one kind of function corrupts memory on the other, so the types keep them apart. init takes a RawScalarInitInfo for the same reason. Generate panic-safe callbacks with scalar_bind_callback! and scalar_init_callback!; the table-function macros do not type-check here.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn dynamic_return_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe extern "C" fn my_bind_fn(_: quack_rs::scalar::RawScalarBindInfo) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
ScalarFunctionBuilder::new("dynamic_return")
    .varargs(TypeId::Varchar)
    .returns(TypeId::Varchar)   // default; overridden in bind
    .bind(my_bind_fn)
    .function(dynamic_return_fn)
    .register(con)?;
Ok(())
}
}

init(init_fn)

Sets a local-init callback invoked once per thread before execution begins. Use this to allocate per-thread state. Maps to duckdb_scalar_function_set_init.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn stateful_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe extern "C" fn my_init_fn(_: quack_rs::scalar::RawScalarInitInfo) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
ScalarFunctionBuilder::new("stateful_fn")
    .param(TypeId::BigInt)
    .returns(TypeId::BigInt)
    .init(my_init_fn)
    .function(stateful_fn)
    .register(con)?;
Ok(())
}
}

Typed bind data and local state

ScalarBindData<T> and ScalarLocalState<T> store a Rust value from the bind and init callbacks with a generated, panic-safe destructor, so no hand-written Box::from_raw is needed. A function that multiplies its argument by a factor fixed at bind time, and counts the rows each thread processes:

#![allow(unused)]
fn main() {
use quack_rs::prelude::*;
use quack_rs::scalar::{
    ScalarBindData, ScalarBindInfo, ScalarFunctionInfo, ScalarInitInfo, ScalarLocalState,
};

#[derive(Clone)]
struct Factor(i64);

quack_rs::scalar_bind_callback!(scaled_bind, |info| {
    // SAFETY: `info` is the argument of the running bind callback.
    let bind = unsafe { ScalarBindInfo::new(info) };
    ScalarBindData::set(&bind, Factor(10));
});

quack_rs::scalar_init_callback!(scaled_init, |info| {
    // SAFETY: `info` is the argument of the running init callback.
    let init = unsafe { ScalarInitInfo::new(info) };
    ScalarLocalState::set(&init, 0_u64); // rows this thread has processed
});

quack_rs::scalar_callback!(scaled, |info, input, output| {
    // SAFETY: DuckDB passes valid handles to the running callback.
    let fn_info = unsafe { ScalarFunctionInfo::new(info) };
    let chunk = unsafe { DataChunk::from_raw(input) };
    // SAFETY: `scaled_bind` stored a `Factor`, and nothing else did.
    let factor = unsafe { ScalarBindData::<Factor>::get(&fn_info) }.map_or(1, |f| f.0);
    // SAFETY: `scaled_init` stored a `u64` on this thread; no other borrow is live.
    if let Some(rows) = unsafe { ScalarLocalState::<u64>::get_mut(&fn_info) } {
        *rows += chunk.size() as u64;
    }
    let reader = unsafe { chunk.reader(0) };
    let mut writer = unsafe { VectorWriter::from_vector(output) };
    for row in 0..chunk.size() {
        // A NULL row holds an arbitrary value; `wrapping_mul` keeps it from
        // overflowing, and `propagate_nulls` below overwrites it with NULL.
        unsafe { writer.write_i64(row, reader.read_i64(row).wrapping_mul(factor)) };
    }
    unsafe { chunk.propagate_nulls(&mut writer) };
});

unsafe fn demo(con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> {
ScalarFunctionBuilder::new("scaled")
    .param(TypeId::BigInt)
    .returns(TypeId::BigInt)
    .bind(scaled_bind)
    .init(scaled_init)
    .function(scaled)
    .register(con)?;
Ok(())
}
}
  • Bind data must be Clone + Send + Sync. Every executing thread reads the same value concurrently, and DuckDB copies the bound expression whenever the optimizer duplicates it (filter pushdown through a projection does). Without a copy callback the copy has no bind data at all, so set registers one that clones T. Wrap data that is expensive or impossible to clone in an Arc<T>.
  • Local state must be Send: it is per thread, but may be freed on another.
  • Call set at most once per callback. DuckDB overwrites the stored pointer on a second call without freeing the first value, so that value is leaked (never dropped).
  • Bind data must depend only on the call's arguments (their values when constant, their types) and extra_info. DuckDB's CScalarFunctionBindData::Equals ignores bind data, so two calls with the same arguments — SELECT f(i), f(i) — are merged and share the first call's bind data. A bind callback that reads a counter, a clock or a random source needs the function marked volatile(), which stops the merge.

Extra info

Attach arbitrary data to a scalar function using extra_info, for example a locale or a configuration struct that parameterises its behaviour. The method is available on both ScalarFunctionBuilder and ScalarOverloadBuilder. The pointee must be Send + Sync: DuckDB passes the same pointer to the callback on every thread that runs the function, and the destructor runs on whichever thread releases it.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe extern "C" fn locale_upper_fn(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe extern "C" fn my_destroy(p: *mut std::os::raw::c_void) {
    drop(unsafe { Box::from_raw(p.cast::<String>()) });
}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
use std::os::raw::c_void;

let config = Box::into_raw(Box::new("en_US".to_string())).cast::<c_void>();
unsafe {
    ScalarFunctionBuilder::new("locale_upper")
        .param(TypeId::Varchar)
        .returns(TypeId::Varchar)
        .extra_info(config, Some(my_destroy))
        .function(locale_upper_fn)
        .register(con)?;
}
Ok(())
}
}

Inside the callback, retrieve the extra info with ScalarFunctionInfo::get_extra_info().


ScalarFunctionInfo

ScalarFunctionInfo wraps the duckdb_function_info handle provided to a scalar function callback. It exposes:

  • get_extra_info() -> *mut c_void — retrieves the extra-info pointer set during registration
  • set_error(message) — reports an error, which fails the query
#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
use quack_rs::scalar::ScalarFunctionInfo;

unsafe extern "C" fn my_fn(
    info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let info = unsafe { ScalarFunctionInfo::new(info) };
    let extra = unsafe { info.get_extra_info() };
    // ... use extra info, or report errors via info.set_error("...") ...
}
}

With the duckdb-1-5 feature, ScalarFunctionInfo also provides:

  • get_bind_data() -> *mut c_void — retrieves bind data set during the bind callback
  • get_state() -> *mut c_void — retrieves per-thread state set during the init callback

ScalarBindInfo (duckdb-1-5)

ScalarBindInfo wraps the duckdb_bind_info handle provided to a scalar function bind callback. It exposes:

  • argument_count() -> u64 — number of arguments
  • argument(index) -> Option<Expression> — the argument at index as an RAII Expression, before DuckDB casts it to the parameter type. On DuckDB before 1.5.5 it fails the bind and returns None, because those releases abort the process when the argument is a scalar subquery
  • get_argument(index) -> duckdb_expression — the raw handle, with that hazard left to the caller
  • get_extra_info() -> *mut c_void — the extra-info pointer from registration
  • set_bind_data(data, destroy) — stores per-query data retrievable during execution
  • set_bind_data_copy(copy) — the callback DuckDB uses to duplicate that data when it copies the bound expression; without one the copy's bind data is NULL
  • set_error(message) — reports an error
  • get_client_context() -> ClientContext — access to the connection's catalog and config

ScalarInitInfo (duckdb-1-5)

ScalarInitInfo wraps the duckdb_init_info handle provided to a scalar function init callback. It exposes:

  • get_extra_info() -> *mut c_void — the extra-info pointer from registration
  • get_bind_data() -> *mut c_void — the bind data from the bind callback
  • set_state(state, destroy) — stores per-thread state retrievable during execution
  • set_error(message) — reports an error
  • get_client_context() -> ClientContext — access to the connection's catalog and config

Aggregate Functions

This page shows how to write a DuckDB aggregate function in Rust with quack-rs: the callbacks DuckDB calls, their signatures, and how AggregateFunctionBuilder registers them. An aggregate function reduces many rows to one value per group, like SUM(), COUNT() or AVG(). Because DuckDB aggregates in parallel, it also has a combine step that merges partial results from parallel workers.

Known DuckDB limitation

Out-of-bounds state reads. Two query shapes make DuckDB call every C-API aggregate's update with a state array holding one state while passing count > 1 rows, so the callback reads states[1..count] past the end of the array (undefined behaviour, in any C-API aggregate, whether built with quack-rs or by hand):

  • Window aggregates whose frame is the whole partition, e.g. agg(x) OVER () — WindowConstantAggregator (src/function/window/window_constant_aggregator.cpp, ~lines 106 and 296–299 in DuckDB 1.5.5).
  • Ordered aggregates, e.g. agg(x ORDER BY y) — src/function/aggregate/sorted_aggregate_function.cpp, ~lines 630–633.

Both paths pass a CONSTANT_VECTOR of states because the function has no simple_update (which the C API cannot set), and CAPIAggregateUpdate (src/main/capi/aggregate_function-c.cpp, ~lines 92–110) passes the vector's data pointer to the extension without flattening it. This is a defect in DuckDB's C API, not in quack-rs, and it cannot be detected from inside the callback: reading states[1] to check is itself the out-of-bounds read. Until DuckDB fixes it, do not use C-API aggregates in those two query shapes.

Reported upstream as duckdb/duckdb#26109.


The aggregate lifecycle

flowchart TD
    REG["<b>Registration</b><br/>AggregateFunctionBuilder<br/>→ duckdb_register_aggregate_function"]

    REG     --> SIZE
    SIZE    --> INIT
    INIT    --> UPDATE
    UPDATE  --> COMBINE
    COMBINE --> FINAL
    FINAL   --> DESTROY

    SIZE["<b>state_size</b>()<br/>How many bytes to allocate per group?"]
    INIT["<b>state_init</b>(state)<br/>Initialise a fresh state"]
    UPDATE["<b>update</b>(chunk, states[])<br/>Process one input batch<br/>(NULL rows included — check is_valid)"]
    COMBINE["<b>combine</b>(src[], tgt[], count)<br/>Merge partial results from parallel workers<br/>⚠️ Pitfall L1: target starts fresh — copy ALL config fields"]
    FINAL["<b>finalize</b>(states[], out, count, offset)<br/>Write count results at out[offset..], once per result batch"]
    DESTROY["<b>state_destroy</b>(states[], count)<br/>Free memory — after finalize and for combine<br/>sources after the merge (not every state: see Known Limitations)"]

    style COMBINE fill:#fff3cd,stroke:#e6ac00,color:#333

DuckDB may call combine many times as it merges partial results. A combine target holds whatever state_init set up, not a copy of the source, so combine must carry every field across (Pitfall L1). state_size is called whenever an operator sizes its state buffers, not once at registration, so it must always return the same value. destroy runs after finalize, and on combine's source states once they have been merged.

combine must leave its source states unchanged (Pitfall L15). A window's segment tree combines the same state into every frame that covers it, from several threads at once. A combine that moves data out of its source (mem::take, or zeroing a counter) is right for the first frame and wrong for the rest: in quack-rs's regression test, a sliding-window sum written that way was wrong on 4985 of 5000 rows. Read the source and copy or clone what the target needs.


Registration

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { count: i64 }
impl AggregateState for MyState {}
unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
use quack_rs::aggregate::AggregateFunctionBuilder;
use quack_rs::types::TypeId;

unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        AggregateFunctionBuilder::new("my_agg")
            .param(TypeId::BigInt)        // input type(s)
            .returns(TypeId::BigInt)      // output type
            .ffi_state::<MyState>()       // state_size + init + destructor
            .update(update)
            .combine(combine)
            .finalize(finalize)
            .register(con)?;
    }
    Ok(())
}
}

register returns an error if the return type or any of the five required callbacks (state_size, init, update, combine, finalize) is missing. The builder treats the destructor as optional (without one, register installs a no-op destructor; see Pitfall L13), but FfiState<T> needs its destroy_callback to drop each T. .ffi_state::<MyState>() sets state_size, init and destructor together from FfiState<MyState>, so the three cannot describe different states; see State Management. update, combine and finalize read the state through FfiState::<MyState>::with_state / with_state_mut with the same type.


Callback signatures

With FfiState<T> you do not write state_size, init or destroy yourself: ffi_state::<T>() installs FfiState::<T>::size_callback, init_callback and destroy_callback. The wrappers below show the signature DuckDB calls each one with, and what it does. update, combine and finalize are yours to write.

state_size

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { config_field: i64, accumulator: i64 }
impl MyState {
    fn accumulate(&mut self, v: i64) { self.accumulator += v; }
    fn result(&self) -> i64 { self.accumulator }
}
impl AggregateState for MyState {}
unsafe extern "C" fn state_size(info: duckdb_function_info) -> idx_t {
    unsafe { FfiState::<MyState>::size_callback(info) }
}
}

Returns the number of bytes DuckDB allocates per group, FfiState::<MyState>::size(): a tag word followed by MyState itself, padded to whole words (or, for a state larger than 256 bytes or aligned more strictly than usize, a Box<MyState> pointer).

state_init

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { config_field: i64, accumulator: i64 }
impl MyState {
    fn accumulate(&mut self, v: i64) { self.accumulator += v; }
    fn result(&self) -> i64 { self.accumulator }
}
impl AggregateState for MyState {}
unsafe extern "C" fn state_init(info: duckdb_function_info, state: duckdb_aggregate_state) {
    unsafe { FfiState::<MyState>::init_callback(info, state) };
}
}

Writes MyState::default() into the DuckDB-allocated state slot (or, for a boxed state, a Box holding it), then writes the tag that marks the slot initialised. A panic in default() is caught and reported to DuckDB as a query error.

update

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { config_field: i64, accumulator: i64 }
impl MyState {
    fn accumulate(&mut self, v: i64) { self.accumulator += v; }
    fn result(&self) -> i64 { self.accumulator }
}
impl AggregateState for MyState {}
unsafe extern "C" fn update(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    states: *mut duckdb_aggregate_state,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let row_count = reader.row_count();

    for row in 0..row_count {
        if unsafe { !reader.is_valid(row) } { continue; }
        let value = unsafe { reader.read_i64(row) };

        let state_ptr = unsafe { *states.add(row) };
        if let Some(st) = unsafe { FfiState::<MyState>::with_state_mut(state_ptr) } {
            st.accumulate(value);
        }
    }
}
}

states[row] is the state of that row's group; rows in the same group share one state. update receives NULL rows too, so skip rows where is_valid is false (see NULL Handling).

combine

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { config_field: i64, accumulator: i64 }
impl MyState {
    fn accumulate(&mut self, v: i64) { self.accumulator += v; }
    fn result(&self) -> i64 { self.accumulator }
}
impl AggregateState for MyState {}
unsafe extern "C" fn combine(
    _info: duckdb_function_info,
    source: *mut duckdb_aggregate_state,
    target: *mut duckdb_aggregate_state,
    count: idx_t,
) {
    for i in 0..count as usize {
        let src = unsafe { FfiState::<MyState>::with_state(*source.add(i)) };
        let tgt = unsafe { FfiState::<MyState>::with_state_mut(*target.add(i)) };
        if let (Some(s), Some(t)) = (src, tgt) {
            // ⚠️  MUST copy ALL fields — see Pitfall L1
            t.config_field = s.config_field;   // configuration
            t.accumulator  += s.accumulator;    // data
        }
    }
}
}

Pitfall L1 — critical: Target states are fresh states, set up by state_init (with FfiState<T>::init_callback, a T::default()), not copies of the source. You must copy every field, including configuration fields set during update. Forgetting even one config field produces silently wrong results. See Pitfall L1.

finalize

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { config_field: i64, accumulator: i64 }
impl MyState {
    fn accumulate(&mut self, v: i64) { self.accumulator += v; }
    fn result(&self) -> i64 { self.accumulator }
}
impl AggregateState for MyState {}
unsafe extern "C" fn finalize(
    _info: duckdb_function_info,
    source: *mut duckdb_aggregate_state,
    result: duckdb_vector,
    count: idx_t,
    offset: idx_t,
) {
    let mut writer = unsafe { VectorWriter::new(result) };
    for i in 0..count as usize {
        let state_ptr = unsafe { *source.add(i) };
        match unsafe { FfiState::<MyState>::with_state(state_ptr) } {
            Some(st) => unsafe { writer.write_i64(offset as usize + i, st.result()) },
            None     => unsafe { writer.set_null(offset as usize + i) },
        }
    }
}
}

offset is non-zero when DuckDB writes the results into part of a larger vector. Always add it to the output index.

state_destroy

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { config_field: i64, accumulator: i64 }
impl AggregateState for MyState {}
unsafe extern "C" fn state_destroy(states: *mut duckdb_aggregate_state, count: idx_t) {
    unsafe { FfiState::<MyState>::destroy_callback(states, count) };
}
}

destroy_callback drops the T in each state whose tag matches (freeing its box, if T is boxed), clearing the tag first, so a second call on the same state is a no-op. See Pitfall L2.


Complex parameter and return types

For functions that accept or return parameterized types like LIST(BIGINT), MAP(VARCHAR, INTEGER), or STRUCT(...), use param_logical and returns_logical instead of param and returns:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { count: i64 }
impl AggregateState for MyState {}
unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
use quack_rs::aggregate::AggregateFunctionBuilder;
use quack_rs::types::{LogicalType, TypeId};

unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        AggregateFunctionBuilder::new("retention")
            .param(TypeId::Boolean)
            .param(TypeId::Boolean)
            .returns_logical(LogicalType::list(TypeId::Boolean))  // LIST(BOOLEAN)
            .ffi_state::<MyState>()                                // state_size + init + destructor
            .update(update)
            .combine(combine)
            .finalize(finalize)
            .register(con)?;
    }
    Ok(())
}
}

param_logical and param can be interleaved — the parameter position is determined by the total number of calls made so far:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
fn demo() {
let _ =
AggregateFunctionBuilder::new("my_func")
    .param(TypeId::Varchar)                          // position 0: VARCHAR
    .param_logical(LogicalType::list(TypeId::BigInt)) // position 1: LIST(BIGINT)
    .param(TypeId::Integer)                           // position 2: INTEGER
    .returns(TypeId::BigInt)
    // ...
;
}
}

If both returns and returns_logical are called, the logical type takes precedence.


Extra info

extra_info attaches arbitrary data to an aggregate function, for example configuration that parameterises its behaviour. DuckDB calls the destroy callback to free the data when the function is dropped:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { count: i64 }
impl AggregateState for MyState {}
unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
unsafe extern "C" fn my_destroy(p: *mut std::os::raw::c_void) {
    drop(unsafe { Box::from_raw(p.cast::<u64>()) });
}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
use std::os::raw::c_void;

let config = Box::into_raw(Box::new(42u64)).cast::<c_void>();
unsafe {
    AggregateFunctionBuilder::new("my_agg")
        .param(TypeId::BigInt)
        .returns(TypeId::BigInt)
        .extra_info(config, Some(my_destroy))
        .ffi_state::<MyState>()  // state_size + init + destructor
        .update(update)
        .combine(combine)
        .finalize(finalize)
        .register(con)?;
}
Ok(())
}
}

Inside callbacks, retrieve the extra info with AggregateFunctionInfo::get_extra_info().


AggregateFunctionInfo

AggregateFunctionInfo wraps the duckdb_function_info handle that DuckDB passes to every aggregate callback except the destructor. It exposes:

  • get_extra_info() -> *mut c_void: the extra-info pointer set at registration.
  • set_error(message): fails the current query with message. Called from finalize, it also leaves some states undestroyed; see Known Limitations.
#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
use quack_rs::aggregate::AggregateFunctionInfo;

unsafe extern "C" fn update(
    info: duckdb_function_info,
    input: duckdb_data_chunk,
    states: *mut duckdb_aggregate_state,
) {
    let info = unsafe { AggregateFunctionInfo::new(info) };
    let extra = unsafe { info.get_extra_info() };
    // ... use extra info, or report errors via info.set_error("...") ...
}
}

Next steps

Aggregate State

This page covers aggregate state in a DuckDB aggregate function written in Rust: the AggregateState trait and FfiState<T>, which manages each state's lifecycle (allocation, initialisation, access and destruction) so that you do not write raw-pointer code for it.

Known DuckDB limitation. Every C-API aggregate, and so every aggregate that uses FfiState<T>, reads out of bounds under agg(x) OVER () (whole-partition window frames) and agg(x ORDER BY y). This is a DuckDB C API defect; see Aggregate Functions for the details and DuckDB source lines. Do not use C-API aggregates in those two query shapes.

Reported upstream as duckdb/duckdb#26109.


AggregateState trait

Any type that is Default + Send + Sync + 'static can be used as aggregate state by implementing the AggregateState marker trait. The Sync bound is new in 0.18.0: a window's segment tree lets several threads read the same state as a combine source at once, so a state containing a Cell or RefCell would race. Use atomics or a Mutex instead, or keep that data outside the state.

#![allow(unused)]
fn main() {
use quack_rs::aggregate::AggregateState;

#[derive(Default, Debug)]
struct MyState {
    config: usize,    // set in update, must be propagated in combine
    total: i64,       // accumulated data
}

impl AggregateState for MyState {}
}

AggregateState has no required methods. state_init uses Default to create each fresh state.


FfiState<T>

FfiState<T> names the layout of the bytes DuckDB allocates for each group's state, and the callbacks that manage them. The type itself is never constructed.

A small T, aligned no more strictly than usize and at most 256 bytes, is stored in those bytes directly. A larger or more strictly aligned T is boxed, and the slot holds the pointer. On wasm32, where usize is 4 bytes but u64, i64 and f64 are 8-byte aligned, a state containing one of them is therefore boxed. Either way the slot starts with a tag.

Memory layout

DuckDB-allocated slot (state_size bytes, a multiple of sizeof(usize)):
  [ tag: usize ][ T, padded to whole words ]      T stored inline
  [ tag: usize ][ *mut T ]                        T boxed
                     │
                     └──→  Box<T>  (on the Rust heap)

Storing T inline matters because DuckDB 1.4.4 to 1.5.5 does not destroy every state. When a grouped aggregate's result scan stops early (a LIMIT above it, an error, an interrupt), the states it never reached are never destroyed; so is one state per row of a window frame with EXCLUDE. An inline T's bytes belong to DuckDB, which frees them with the hash table; only what T itself owns on the heap, or a boxed T's box, leaks. See Known Limitations.

The tag marks the slot initialised. When one state_init call fails (a panicking T::default(), say), DuckDB 1.4.4 to 1.5.5 still runs the destructor over every state it created, including states whose state_init never ran, so a slot can hold arbitrary bytes. destroy_callback drops only a slot carrying the tag init_callback wrote, and clears the tag first. The tag is derived from a hash of T's TypeId (and, for a boxed T, the box's address), not from the slot's address, because DuckDB moves states by copying their bytes. The check turns dropping garbage from a certainty into a matter of chance: uninitialised bytes that happen to equal the tag. It mitigates the DuckDB defect (described in the repository's docs/upstream-duckdb-reports.md); it cannot guarantee against it.

Lifecycle callbacks

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::aggregate::{AggregateState, FfiState};
#[derive(Default, Debug)] struct MyState { config: usize, total: i64 }
impl AggregateState for MyState {}
unsafe fn demo(_info: duckdb_function_info, info: duckdb_function_info,
    state: duckdb_aggregate_state, states: *mut duckdb_aggregate_state, count: idx_t) {
// state_size: DuckDB calls this whenever an operator sizes its state buffers
FfiState::<MyState>::size_callback(_info);
// Returns: FfiState::<MyState>::size() (on a 64-bit target, a tag word, then the
// usize and i64 inline)

// state_init: DuckDB calls this for every state slot it allocates, combine
// targets included
FfiState::<MyState>::init_callback(info, state);
// Effect: writes MyState::default() into the slot (or a box holding it), then the tag

// destructor: DuckDB calls this after finalize, on combine's source states once
// merged, and (after a failed state_init) on states never initialised, which
// the tag makes it skip; not on every state (see Known Limitations)
FfiState::<MyState>::destroy_callback(states, count);
// Effect: for each state whose tag matches: clear the tag, then drop the T
}
}

Wiring them up: ffi_state::<T>()

The recommended way to register those three callbacks is ffi_state::<T>() (new in 0.18.0). It installs size_callback, init_callback and destroy_callback for the same T in one call, so the size DuckDB allocates and the state init writes cannot disagree. Wired one by one with the state_size, init and destructor setters, a size callback for one type paired with an init callback for a larger one writes past DuckDB's allocation.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_connection, duckdb_data_chunk,
    duckdb_function_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct MyState { config: usize, total: i64 }
impl AggregateState for MyState {}
unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        AggregateFunctionBuilder::new("my_agg")
            .param(TypeId::BigInt)
            .returns(TypeId::BigInt)
            .ffi_state::<MyState>()   // state_size + init + destructor
            .update(update)
            .combine(combine)
            .finalize(finalize)
            .register(con)?;
    }
    Ok(())
}
}

AggregateOverloadBuilder has the same method, for each overload of an AggregateFunctionSetBuilder; the set builder itself has none. update, combine and finalize still read the state through FfiState::<T>::with_state / with_state_mut with the same T.

Accessing state in callbacks

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::aggregate::{AggregateState, FfiState};
#[derive(Default, Debug)] struct MyState { config: usize, total: i64 }
impl AggregateState for MyState {}
unsafe fn demo(state_ptr: duckdb_aggregate_state, delta: i64) {
// Immutable access (in finalize, combine source):
if let Some(st) = FfiState::<MyState>::with_state(state_ptr) {
    let value = st.total;
}

// Mutable access (in update, combine target):
if let Some(st) = FfiState::<MyState>::with_state_mut(state_ptr) {
    st.total += delta;
}
}
}

The methods return Option<&T> and Option<&mut T> respectively: None if the slot's tag does not match, which happens after destroy_callback has run or when T::default() panicked in state_init. Returning Option instead of panicking keeps a panic from unwinding across the FFI boundary (Pitfall L3).


The double-free problem — solved

Without quack-rs, a naive destructor looks like:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, idx_t};
struct MyState;
// A hand-written boxed layout: this is the code *without* quack-rs.
#[repr(C)] struct FfiState<T> { inner: *mut T }
// ❌ Naive — causes double-free if DuckDB calls destroy twice
unsafe extern "C" fn destroy(states: *mut duckdb_aggregate_state, count: idx_t) {
    for i in 0..count as usize {
        let ffi = &mut *(*states.add(i) as *mut FfiState<MyState>);
        drop(Box::from_raw(ffi.inner));   // inner is now dangling — crash on second call
    }
}
}

FfiState::destroy_callback clears the slot's tag before dropping the T, and drops only a slot whose tag matches. If DuckDB calls destroy again, the tag no longer matches, the slot is skipped, and with_state returns None.


Testing state logic without DuckDB

AggregateTestHarness<S> simulates the DuckDB aggregate lifecycle in pure Rust:

#![allow(unused)]
fn main() {
use quack_rs::aggregate::AggregateState;
#[derive(Default, Debug)] struct MyState { config: usize, total: i64 }
impl AggregateState for MyState {}
use quack_rs::testing::AggregateTestHarness;

#[test]
fn combine_propagates_config() {
    let mut source = AggregateTestHarness::<MyState>::new();
    source.update(|s| {
        s.config = 5;    // config field set during update
        s.total += 100;
    });

    let mut target = AggregateTestHarness::<MyState>::new();
    target.combine(&source, |src, tgt| {
        tgt.config = src.config;   // must propagate config — Pitfall L1
        tgt.total  += src.total;
    });

    let result = target.finalize();
    assert_eq!(result.config, 5, "config must be propagated in combine");
    assert_eq!(result.total, 100);
}
}

See the Testing Guide for the full test strategy.

Overloading with Function Sets

A DuckDB aggregate function can have several signatures under one name, registered together as a function set. This page shows how to overload an aggregate function in Rust with AggregateFunctionSetBuilder, including variadic aggregates such as retention(c1, c2, ..., c32) and overloads with different return types.

Known DuckDB limitation. C-API aggregates — including every overload in a set — read out of bounds under agg(x) OVER () (whole-partition window frames) and agg(x ORDER BY y). This is a DuckDB C API defect; see Aggregate Functions for the details and DuckDB source lines. Do not use C-API aggregates in those two query shapes.

Reported upstream as duckdb/duckdb#26109.

Note: For scalar function overloads, see ScalarFunctionSetBuilder.


When to use function sets

Use AggregateFunctionSetBuilder when you need:

  • Multiple type signatures for the same function name (e.g., my_agg(INT) and my_agg(BIGINT))
  • Variadic arity under one name (e.g., retention(2 columns), retention(3 columns), ...)
  • Overloads that return different types (see Per-overload return types)

For a single signature, use AggregateFunctionBuilder directly.


Registration

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct RetentionState { hits: u32 }
impl AggregateState for RetentionState {}
unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
use quack_rs::aggregate::AggregateFunctionSetBuilder;
use quack_rs::types::TypeId;

unsafe fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        AggregateFunctionSetBuilder::new("retention")
            .returns(TypeId::Varchar)
            .overloads(2..=3, |n, builder| {
                // Each overload gets `n` BOOLEAN parameters
                let b = (0..n).fold(builder, |b, _| b.param(TypeId::Boolean));
                b.ffi_state::<RetentionState>() // state_size + init + destructor
                    .update(update)
                    .combine(combine)
                    .finalize(finalize)
            })
            .register(con)?;
    }
    Ok(())
}
}

overloads takes a RangeInclusive<usize> and a closure, called once per arity n with a fresh AggregateOverloadBuilder, that returns the configured overload. The set builder gives every member the set's name when it registers them.


Per-overload return types

DuckDB resolves an aggregate overload from its parameter types and arity alone — the return type plays no part in resolution. Members of one set are therefore free to return different types. This is how DuckDB's own arg_max works:

arg_max(ANY, ANY)      -> ANY
arg_max(ANY, ANY, ANY) -> ANY[]

Set the return type on the overload with AggregateOverloadBuilder::returns (or returns_logical), and add each one with overload:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct IntState { sum: i64 }
impl AggregateState for IntState {}
unsafe extern "C" fn int_update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn int_combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn int_finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
#[derive(Default)] struct StrState { longest: String }
impl AggregateState for StrState {}
unsafe extern "C" fn str_update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn str_combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn str_finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
use quack_rs::aggregate::{AggregateFunctionSetBuilder, AggregateOverloadBuilder};
use quack_rs::types::TypeId;

unsafe {
    AggregateFunctionSetBuilder::new("my_agg")
        .overload(
            AggregateOverloadBuilder::new()
                .param(TypeId::Integer)
                .returns(TypeId::Integer)      // my_agg(INTEGER) -> INTEGER
                .ffi_state::<IntState>()
                .update(int_update)
                .combine(int_combine)
                .finalize(int_finalize),
        )
        .overload(
            AggregateOverloadBuilder::new()
                .param(TypeId::Varchar)
                .returns(TypeId::Varchar)      // my_agg(VARCHAR) -> VARCHAR
                .ffi_state::<StrState>()
                .update(str_update)
                .combine(str_combine)
                .finalize(str_finalize),
        )
        .register(con)?;
}
Ok(())
}
}

Each overload carries its own callbacks, so overloads with different parameter types can use different state types: call ffi_state::<T>() on each overload with that overload's state type (AggregateFunctionSetBuilder itself has no ffi_state), and read it in that overload's update, combine and finalize with FfiState::<T>::with_state / with_state_mut for the same T.

Which return type wins

For each overload, in order:

  1. AggregateOverloadBuilder::returns_logical
  2. AggregateOverloadBuilder::returns
  3. AggregateFunctionSetBuilder::returns_logical (the set-level default)
  4. AggregateFunctionSetBuilder::returns (the set-level default)

Registration fails, naming the overload index, if an overload reaches the end of that list with nothing set or is missing a required callback. It also fails if two overloads take the same parameter types. All of this is checked for every overload before any DuckDB handle is created.

Other per-overload settings

Each overload can also carry its own extra_info (AggregateOverloadBuilder::extra_info), read in that overload's callbacks with AggregateFunctionInfo::get_extra_info. Ownership works as on AggregateFunctionBuilder: DuckDB frees it once registration has handed it over, even if registration then fails; the builder frees it if it never gets that far.

AggregateOverloadBuilder::null_handling sets an overload's NULL handling.

overload and overloads may be mixed on one builder; overloads register in the order they were added.

Note: AggregateOverloadBuilder was called OverloadBuilder before v0.18.0. The old name is still exported as a deprecated alias at quack_rs::aggregate::builder::OverloadBuilder.


The silent name bug — solved

Pitfall L6: When using a function set, the name must be set on each individual duckdb_aggregate_function via duckdb_aggregate_function_set_name, not just on the set. If any member lacks a name, it is silently not registered — no error is returned.

DuckDB's C API documentation does not mention this. It was found by reading DuckDB's C++ test code at test/api/capi/test_capi_aggregate_functions.cpp. In duckdb-behavioral, 6 of 7 functions failed to register silently due to this bug.

AggregateFunctionSetBuilder::register calls duckdb_aggregate_function_set_name on every member, whether it was added with overload or overloads.

See Pitfall L6.


Complex return types

If all overloads share one complex return type, set it once on the set builder as a default, rather than repeating it on every overload:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct RetentionState { hits: u32 }
impl AggregateState for RetentionState {}
unsafe extern "C" fn update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
use quack_rs::aggregate::AggregateFunctionSetBuilder;
use quack_rs::types::{LogicalType, TypeId};

unsafe {
    AggregateFunctionSetBuilder::new("retention")
        .returns_logical(LogicalType::list(TypeId::Boolean))  // default for every overload
        .overloads(2..=32, |n, builder| {
            (0..n).fold(builder, |b, _| b.param(TypeId::Boolean))
                .ffi_state::<RetentionState>()
                .update(update)
                .combine(combine)
                .finalize(finalize)
        })
        .register(con)?;
}
Ok(())
}
}

Individual overloads can also use param_logical for complex parameter types:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
fn demo() {
let _ = AggregateFunctionSetBuilder::new("retention")
.overloads(2..=8, |n, builder| {
    builder
        .param(TypeId::Interval)
        .param_logical(LogicalType::list(TypeId::Timestamp)) // LIST(TIMESTAMP) parameter
        // ...
})
;
}
}

Why not varargs?

DuckDB's C API has no duckdb_aggregate_function_set_varargs. A variadic aggregate must therefore be registered as one overload per supported arity, which overloads does in one call.

Note: Scalar functions do support varargs, through ScalarFunctionBuilder::varargs() (stable C API, no feature flag needed).

Table Functions

A DuckDB table function returns a result set rather than a single value, and is called in the FROM clause: SELECT * FROM my_function(args). This page shows how to write one in Rust with quack-rs. DuckDB drives a table function through three callbacks: bind, init and scan.

quack-rs provides two layers for registering table functions:

  1. TypedTableFunctionBuilder<S> (recommended for new extensions) — closure-based API that hides bind/init/scan trampolines behind safe Rust closures and gives every execution a fresh, typed scan state built from what bind produced.
  2. TableFunctionBuilder — the underlying raw builder used by TypedTableFunctionBuilder internally. Reach for it when you need fine-grained control: parallel scans (InitInfo::set_max_threads above 1, usually with local_init for per-thread state), projection pushdown with column filtering, or callback shapes that don't fit the "produce state in bind, mutate it in scan" model.

Both builders are backed by the helper types BindInfo, InitInfo, FunctionInfo, FfiBindData<T>, FfiInitData<T>, and FfiLocalInitData<T>.

Lifecycle

PhaseCallbackCalled whenTypical work
bindbind_fnQuery is planned (once per plan)Extract parameters; declare output columns; store configuration in bind data
initinit_fnEach execution of the plan startsAllocate per-scan state (cursor, row index, etc.)
scanscan_fnEach output batchFill the output chunk with rows; set its size with DataChunk::set_size

DuckDB calls the scan callback repeatedly until it sets the chunk size to 0, which signals the end of the results.

Bind once, init many times. DuckDB keeps the bind data for as long as the bound plan lives and runs init against it on every execution: each EXECUTE of a prepared statement, each iteration of a recursive CTE that references the function. Treat bind data as immutable after bind and build anything a scan consumes (cursors, open files) in init.

Closure-based typed state (with_state)

For the common "take parameters at bind, stream rows until exhausted" pattern, TypedTableFunctionBuilder<S> replaces all three callback trampolines with two closures. With with_state, the state returned by bind is a template: every execution of the plan scans a fresh clone() of it, so S must be Clone + Send.

#![allow(unused)]
fn main() {
use quack_rs::prelude::*;

#[derive(Clone)]
struct State {
    remaining: u64,
}

fn register(reg: &impl Registrar) -> ExtResult<()> {
    let builder = TableFunctionBuilder::new("count_down")
        .param(TypeId::BigInt)
        // 1. bind closure: declare the output schema, read parameters,
        //    return the template scan state (cloned for every execution).
        .with_state::<State, _>(|bind| {
            bind.add_result_column("n", TypeId::BigInt);
            let raw = unsafe { bind.get_parameter_value(0) };
            Ok(State { remaining: raw.as_i64_or(0).max(0) as u64 })
        })
        // 2. scan closure: mutate state, write rows, set chunk size.
        .scan(|state, chunk| {
            if state.remaining == 0 {
                unsafe { chunk.set_size(0) };
                return Ok(());
            }
            let mut writer = unsafe { chunk.writer(0) };
            unsafe { writer.write_i64(0, state.remaining as i64) };
            state.remaining -= 1;
            unsafe { chunk.set_size(1) };
            Ok(())
        })
        .build()?;
    unsafe { reg.register_table(builder) }
}
}

Separate bind data and scan state (with_bind_init)

When the scan state is expensive or impossible to clone (it owns a file handle, a large buffer, a connection), or the parameters and the cursor are naturally separate, use with_bind_init. bind returns immutable bind data B (Send + Sync); init builds a fresh scan state S from &B for every execution:

#![allow(unused)]
fn main() {
use quack_rs::prelude::*;

struct Params { n: i64 }
struct Cursor { next: i64, end: i64 }   // no Clone needed

fn register(reg: &impl Registrar) -> ExtResult<()> {
    let builder = TableFunctionBuilder::new("count_up")
        .param(TypeId::BigInt)
        .with_bind_init(
            |bind| {
                bind.add_result_column("n", TypeId::BigInt);
                let n = unsafe { bind.get_parameter_value(0) }.as_i64_or(0);
                Ok(Params { n })
            },
            |params: &Params| Ok(Cursor { next: 1, end: params.n }),
        )
        .scan(|cursor, chunk| {
            if cursor.next > cursor.end {
                unsafe { chunk.set_size(0) };
                return Ok(());
            }
            unsafe {
                chunk.writer(0).write_i64(0, cursor.next);
                chunk.set_size(1);
            }
            cursor.next += 1;
            Ok(())
        })
        .build()?;
    unsafe { reg.register_table(builder) }
}
}

What you get for free

  • No hand-written unsafe extern "C" fn trampolines. TypedTableFunctionBuilder generates them internally.
  • Typed scan state. The scan closure receives &mut S, freshly built for each execution (a clone of the with_state template, or init(&B) for with_bind_init) — no manual FfiBindData / FfiInitData shuffling, and a prepared statement can be executed any number of times.
  • Panic safety. User closures run inside catch_unwind. A panic is reported as a bind, init or scan error, and the scan sets the chunk size to zero, so the query fails cleanly instead of unwinding across the FFI boundary.
  • Error propagation. Return Err(ExtensionError::new("...")) from any closure to report a SQL error to DuckDB.

Trade-offs and threading

  • S must be Send + 'static (plus Clone for with_state). Sync is not required, so TypedTableFunctionBuilder forces scans to run on a single worker by calling InitInfo::set_max_threads(1) internally.
  • The typed builder does not offer projection pushdown: with pushdown on, the scan's chunk holds only the projected columns and a closure written against the declared schema would write the wrong column. build() returns an error if projection_pushdown(true) was set on the raw builder before with_state / with_bind_init. Use the raw builder for pushdown.
  • The bind closure must declare at least one column. With duckdb-1-5, a bind that declares none is reported as an ordinary bind error; DuckDB itself raises an INTERNAL Error with a C++ stack trace for it (and does so for a raw bind callback, which quack-rs cannot check after it returns — call set_error yourself).
  • Extensions that need multi-worker parallelism (set_max_threads above 1, with local_init + thread-local buffers) should use the raw TableFunctionBuilder directly.
  • TypedTableFunctionBuilder::build() returns a fully configured TableFunctionBuilder, so you can still pass it through any Registrar — including MockRegistrar for unit tests. Registration fails if you then enable projection_pushdown on it or replace its bind, init, local_init, scan or extra_info: the generated callbacks read one another's data.

Builder API

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk,
    duckdb_function_info, duckdb_init_info};
unsafe extern "C" fn my_bind_callback(_: duckdb_bind_info) {}
unsafe extern "C" fn my_init_callback(_: duckdb_init_info) {}
unsafe extern "C" fn my_scan_callback(_: duckdb_function_info, _: duckdb_data_chunk) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
use quack_rs::table::{TableFunctionBuilder, BindInfo, FfiBindData, FfiInitData};
use quack_rs::types::TypeId;

TableFunctionBuilder::new("my_function")
    .param(TypeId::BigInt)                 // positional parameter types
    .bind(my_bind_callback)               // declare output columns inside bind
    .init(my_init_callback)
    .scan(my_scan_callback)
    .register(con)?;
Ok(())
}
}

Output columns are declared inside the bind callback using BindInfo::add_result_column, not on the builder itself.

State management

Bind data

Bind data persists from the bind phase through all scan batches — and through every later execution of the same plan (see Bind once, init many times above), possibly read from several threads at once, so FfiBindData::set requires T: Send + Sync. Use FfiBindData<T> to allocate it safely:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_bind_info;
use quack_rs::table::{BindInfo, FfiBindData};
struct MyBindData {
    limit: i64,
}

unsafe extern "C" fn my_bind(info: duckdb_bind_info) {
    // `get_parameter_value` returns an RAII `Value`; a NULL argument reads as the default.
    let n = unsafe { BindInfo::new(info).get_parameter_value(0) }.as_i64_or(0);
    unsafe { FfiBindData::<MyBindData>::set(info, MyBindData { limit: n }) };
}
}

FfiBindData::set stores the value and registers a destructor so DuckDB frees it at the right time — no Box::into_raw / Box::from_raw needed.

Init (scan) state

Per-scan state (e.g., a current row index) uses FfiInitData<T> (T: Send + Sync, since concurrent scan threads share it):

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_init_info;
use quack_rs::table::FfiInitData;
struct MyScanState {
    pos: i64,
}

unsafe extern "C" fn my_init(info: duckdb_init_info) {
    unsafe { FfiInitData::<MyScanState>::set(info, MyScanState { pos: 0 }) };
}
}

Complete example: generate_series_ext

The hello-ext example registers generate_series_ext(n BIGINT) which emits integers 0 .. n-1. See examples/hello-ext/src/lib.rs for the full source.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk,
    duckdb_function_info, duckdb_init_info, DuckDBSuccess};
use quack_rs::data_chunk::DataChunk;
use quack_rs::table::{BindInfo, FfiBindData, FfiInitData, TableFunctionBuilder};
use quack_rs::types::TypeId;
struct GsBindData { total: i64 }
struct GsScanState { pos: i64 }
// Bind: extract `n`, register one output column
unsafe extern "C" fn gs_bind(info: duckdb_bind_info) {
    let bind_info = unsafe { BindInfo::new(info) };
    // Value is RAII — automatically destroyed when dropped.
    // A NULL argument reads as the default rather than aborting.
    let n = unsafe { bind_info.get_parameter_value(0) }.as_i64_or(0);

    bind_info.add_result_column("value", TypeId::BigInt);
    unsafe { FfiBindData::<GsBindData>::set(info, GsBindData { total: n }) };
}

// Init: zero-initialise the scan cursor
unsafe extern "C" fn gs_init(info: duckdb_init_info) {
    unsafe { FfiInitData::<GsScanState>::set(info, GsScanState { pos: 0 }) };
}

// Scan: emit a batch of rows using DataChunk wrapper
unsafe extern "C" fn gs_scan(info: duckdb_function_info, output: duckdb_data_chunk) {
    let chunk = unsafe { DataChunk::from_raw(output) };
    // Never unwrap in a callback: a missing state ends the scan instead.
    let bind = unsafe { FfiBindData::<GsBindData>::get_from_function(info) };
    let state = unsafe { FfiInitData::<GsScanState>::get_mut(info) };
    let (Some(bind), Some(state)) = (bind, state) else {
        unsafe { chunk.set_size(0) };
        return;
    };

    let remaining = bind.total - state.pos;
    let batch = remaining.min(2048).max(0) as usize;

    let mut writer = unsafe { chunk.writer(0) };
    for i in 0..batch {
        unsafe { writer.write_i64(i, state.pos + i as i64) };
    }
    unsafe { chunk.set_size(batch) };
    state.pos += batch as i64;
}
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
unsafe {
    assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), DuckDBSuccess);
    assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), DuckDBSuccess);
    TableFunctionBuilder::new("generate_series_ext")
        .param(TypeId::BigInt)
        .bind(gs_bind)
        .init(gs_init)
        .scan(gs_scan)
        .register(con)
        .unwrap();
}
let sum = |sql: &str| -> i64 {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    unsafe { chunk.reader(0).read_i64(0) }
};
assert_eq!(sum("SELECT sum(value)::BIGINT FROM generate_series_ext(5)"), 10);
assert_eq!(sum("SELECT count(*) FROM generate_series_ext(5000)"), 5000);
}

Registration

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk,
    duckdb_function_info, duckdb_init_info};
use quack_rs::table::TableFunctionBuilder;
use quack_rs::types::TypeId;
unsafe extern "C" fn gs_bind(_: duckdb_bind_info) {}
unsafe extern "C" fn gs_init(_: duckdb_init_info) {}
unsafe extern "C" fn gs_scan(_: duckdb_function_info, _: duckdb_data_chunk) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
TableFunctionBuilder::new("generate_series_ext")
    .param(TypeId::BigInt)
    .bind(gs_bind)
    .init(gs_init)
    .scan(gs_scan)
    .register(con)?;
Ok(())
}
}

Advanced features

Named parameters

Named parameters let callers pass optional arguments by name (e.g., step := 10):

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk,
    duckdb_function_info, duckdb_init_info};
use quack_rs::table::TableFunctionBuilder;
use quack_rs::types::TypeId;
unsafe extern "C" fn gs_v2_bind(_: duckdb_bind_info) {}
unsafe extern "C" fn gs_v2_init(_: duckdb_init_info) {}
unsafe extern "C" fn gs_v2_scan(_: duckdb_function_info, _: duckdb_data_chunk) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
TableFunctionBuilder::new("gen_series_v2")
    .param(TypeId::BigInt)                    // positional: n
    .named_param("step", TypeId::BigInt)      // named: step := <value>
    .bind(gs_v2_bind)
    .init(gs_v2_init)
    .scan(gs_v2_scan)
    .register(con)?;
Ok(())
}
}

In the bind callback, read the named parameter with BindInfo::get_named_parameter_value("step"). Named parameters are optional: if the query omits step := …, the returned Value wraps a null handle (is_null() is true), so use a defaulting accessor such as as_i64_or(1).

Registering a name twice

The C API has no table function sets: a second registration under a name that already exists — your own earlier one, another extension's, or a built-in such as range — is dropped by DuckDB while duckdb_register_table_function still reports success, and the old function keeps answering. register therefore checks duckdb_functions() first and returns an error if the name already belongs to a table function or table macro (compared case-insensitively). Table functions registered through the C API live in the in-memory system catalog and are never persisted, so reloading an extension into a database file never trips this check.

Local init (per-thread state)

local_init allocates per-thread state for a scan that runs on several threads. It does not make the scan parallel by itself — that is InitInfo::set_max_threads (see Thread control):

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk,
    duckdb_function_info, duckdb_init_info};
use quack_rs::table::TableFunctionBuilder;
use quack_rs::types::TypeId;
unsafe extern "C" fn gs_v2_bind(_: duckdb_bind_info) {}
unsafe extern "C" fn gs_v2_init(_: duckdb_init_info) {}
unsafe extern "C" fn gs_v2_local_init(_: duckdb_init_info) {}
unsafe extern "C" fn gs_v2_scan(_: duckdb_function_info, _: duckdb_data_chunk) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
TableFunctionBuilder::new("gen_series_v2")
    .param(TypeId::BigInt)
    .bind(gs_v2_bind)
    .init(gs_v2_init)
    .local_init(gs_v2_local_init)            // per-thread state allocation
    .scan(gs_v2_scan)
    .register(con)?;
Ok(())
}
}

The local init callback receives duckdb_init_info and can use FfiLocalInitData::<T>::set to store per-thread state (T: Send).

Thread control

Use InitInfo::set_max_threads in the global init callback to tell DuckDB how many threads can scan concurrently. The default is 1. Above 1, DuckDB calls the scan from that many threads at the same time whether or not local_init is set — and all of them share the one global init data and bind data. Do not use FfiInitData::get_mut then; keep shared mutable state behind a Mutex or atomics and read it with FfiInitData::get:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_init_info;
use quack_rs::table::{FfiInitData, InitInfo};
struct MyState { pos: i64 }
unsafe extern "C" fn gs_v2_init(info: duckdb_init_info) {
    let init_info = unsafe { InitInfo::new(info) };
    init_info.set_max_threads(1);
    unsafe { FfiInitData::<MyState>::set(info, MyState { pos: 0 }) };
}
}

Projection pushdown

Enable projection pushdown to let DuckDB skip unrequested columns:

#![allow(unused)]
fn main() {
use quack_rs::table::TableFunctionBuilder;
fn demo() {
let _ =
TableFunctionBuilder::new("my_func")
    .projection_pushdown(true)
    // ...
;
}
}

Caution: When projection pushdown is enabled, your scan callback must check which columns DuckDB actually needs using InitInfo::projected_column_count and InitInfo::projected_column_index. Writing to non-projected columns causes crashes. projected_column_index returns None past the end of the projection (the C API itself answers 0 there, which is indistinguishable from the first column).

See examples/hello-ext/src/lib.rs for a complete example using named_param, local_init, and set_max_threads.

Complex parameter types

For parameterised types that TypeId cannot express (e.g. LIST(BIGINT), MAP(VARCHAR, INTEGER), STRUCT(...)), use param_logical and named_param_logical:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_bind_info, duckdb_connection, duckdb_data_chunk,
    duckdb_function_info, duckdb_init_info};
use quack_rs::table::TableFunctionBuilder;
use quack_rs::types::TypeId;
unsafe extern "C" fn bind_fn(_: duckdb_bind_info) {}
unsafe extern "C" fn init_fn(_: duckdb_init_info) {}
unsafe extern "C" fn scan_fn(_: duckdb_function_info, _: duckdb_data_chunk) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
use quack_rs::types::LogicalType;

TableFunctionBuilder::new("read_data")
    .param_logical(LogicalType::list(TypeId::Varchar))        // positional LIST param
    .named_param_logical("options", LogicalType::map(          // named MAP param
        TypeId::Varchar, TypeId::Varchar,
    ))
    .bind(bind_fn)
    .init(init_fn)
    .scan(scan_fn)
    .register(con)?;
Ok(())
}
}

BindInfo helpers

BindInfo wraps duckdb_bind_info and exposes these methods:

MethodDescription
add_result_column(name, TypeId)Declares an output column (a type DuckDB would silently drop, like ANY, is a bind error instead)
add_result_column_with_type(name, &LogicalType)Output column with complex type (same check, including nested ANY/INVALID)
set_cardinality(rows, is_exact)Cardinality hint for the optimizer. DuckDB 1.5.5 treats is_exact = false as an estimate and an upper bound, true as an estimate only — the reverse of duckdb.h
set_error(message)Report a bind-time error (an empty message is replaced by a placeholder)
parameter_count()Number of positional parameters
get_parameter_value(index)Positional parameter as an RAII Value
get_named_parameter_value(name)Named parameter as an RAII Value; a null handle if the query omitted it
get_parameter(index) / get_named_parameter(name)The same as a raw duckdb_value, which the caller must destroy
get_extra_info()Returns the extra-info pointer set on the function
get_client_context()Returns a ClientContext (duckdb-1-5)
result_column_count() / result_column_name(i) / result_column_type(i)The target table's columns in a COPY … FROM reader (duckdb-1-5); zero columns otherwise

InitInfo helpers

InitInfo wraps duckdb_init_info:

MethodDescription
projected_column_count()Number of projected columns (with pushdown)
projected_column_index(idx)Declared column index at projection position; None when idx is out of range
set_max_threads(n)Maximum concurrent scan threads (default 1; shared global state above 1)
set_error(message)Report an init-time error (an empty message is replaced by a placeholder)
get_extra_info()Returns the extra-info pointer set on the function

FunctionInfo helpers

FunctionInfo wraps duckdb_function_info (scan callbacks):

MethodDescription
set_error(message)Report a scan-time error (an empty message is replaced by a placeholder)
get_extra_info()Returns the extra-info pointer set on the function

Extra info

Use TableFunctionBuilder::extra_info to attach function-level data that is accessible from all callbacks (bind, init, and scan) via get_extra_info(). The pointee must be Send + Sync: DuckDB passes the same pointer to callbacks running on several threads at once, and frees it on whichever thread releases the function.

Example output

SELECT * FROM generate_series_ext(5);
-- 0
-- 1
-- 2
-- 3
-- 4

SELECT value * value AS sq FROM generate_series_ext(4);
-- 0
-- 1
-- 4
-- 9

See also

Replacement Scans

A DuckDB replacement scan lets users query a file by its path alone:

SELECT * FROM 'myfile.myformat'

When DuckDB finds no table with that name, it calls each registered replacement scan in registration order, and one of them can redirect the query to a table function (in the example below, read_myformat('myfile.myformat')). DuckDB's own CSV, Parquet and JSON readers use the same mechanism. This page shows how to register one from a Rust extension with quack-rs.

quack-rs provides ReplacementScanBuilder (a static registration helper) and ReplacementScanInfo (an ergonomic wrapper for callbacks).

Registration API

Unlike the other builders in quack-rs, ReplacementScanBuilder uses a single static call because the DuckDB C API takes all arguments at once:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_database, duckdb_replacement_scan_info};
use std::os::raw::{c_char, c_void};
unsafe extern "C" fn my_scan_callback(_: duckdb_replacement_scan_info, _: *const c_char, _: *mut c_void) {}
fn demo(db: duckdb_database, my_state: String) {
use quack_rs::replacement_scan::ReplacementScanBuilder;

// Low-level: pass raw extra_data and an optional delete callback.
unsafe {
    ReplacementScanBuilder::register(
        db,                            // duckdb_database
        my_scan_callback,              // ReplacementScanFn
        std::ptr::null_mut(),          // extra_data (or a raw pointer)
        None,                          // delete_callback
    );
}

// Ergonomic: pass owned Rust data; boxing and destructor are handled for you.
unsafe {
    ReplacementScanBuilder::register_with_data(db, my_scan_callback, my_state);
}
}
}

Note: Replacement scans are registered on a database handle (duckdb_database), not a connection, and apply to every connection to that database. In an entry point, Connection::register_replacement_scan and register_replacement_scan_with_data do the same through the Connection you are given.

A raw delete_callback must accept a null argument: unlike DuckDB's other destructor slots, it is called with extra_data even when that is null.

Callback signature

The raw callback receives duckdb_replacement_scan_info, but you can wrap it with ReplacementScanInfo for ergonomic, safe access:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_replacement_scan_info;
use quack_rs::replacement_scan::ReplacementScanInfo;

unsafe extern "C" fn my_scan_callback(
    info: duckdb_replacement_scan_info,
    table_name: *const ::std::os::raw::c_char,
    _data: *mut ::std::os::raw::c_void,
) {
    let path = unsafe { std::ffi::CStr::from_ptr(table_name) }
        .to_str()
        .unwrap_or("");

    if !path.ends_with(".myformat") {
        return; // not ours: DuckDB tries the next replacement scan
    }

    // Use ReplacementScanInfo for ergonomic access
    unsafe {
        ReplacementScanInfo::new(info)
            .set_function("read_myformat")
            .add_varchar_parameter(path);
    }
}
}

table_name is the last part of the table reference: FROM myschema."f.myformat" reaches the callback as f.myformat, so a callback cannot see or honour a schema.

ReplacementScanInfo methods

MethodDescription
set_function(name)Redirect to the named table function
add_varchar_parameter(value)Add a VARCHAR parameter to the redirected call
add_i64_parameter(value)Add a BIGINT (i64) parameter
add_bool_parameter(value)Add a BOOLEAN parameter
add_parameter_raw(duckdb_value)Add a parameter of any type; DuckDB copies it, so the caller still destroys the value
set_error(message)Report an error, which fails the query. DuckDB ignores an empty message, so an empty one is replaced with a placeholder

Data passed to ReplacementScanBuilder::register_with_data must be Send + Sync: it lives in the database-wide configuration, is read by the callback from any connection's thread (concurrently), and is dropped by whichever thread closes the database. The raw register has the same requirement, stated in its # Safety.

When to use replacement scans vs table functions

ScenarioUse
SELECT * FROM my_function('file.ext')Table function
SELECT * FROM 'file.ext' (bare path)Replacement scan → delegates to a table function
File type auto-detectionReplacement scan

Most extensions implement both: a table function that does the actual work, and a replacement scan that detects the file extension and transparently routes bare-path queries to the table function.

See also

Cast Functions

A DuckDB cast function defines how values of one type are converted to another. This page shows how to register one from a Rust extension with quack-rs' CastFunctionBuilder. Once registered, CAST(x AS T) and TRY_CAST(x AS T) use your callback, and so do implicit conversions if you give the cast an implicit cost.

When to use cast functions

  • Your extension introduces a new logical type and needs CAST to/from standard types.
  • You want to override DuckDB's built-in cast behaviour for a specific type pair.
  • You need to control implicit cast priority relative to other registered casts.

Registering a cast

#![allow(unused)]
fn main() {
use quack_rs::cast::{CastFunctionBuilder, CastFunctionInfo, CastMode};
use quack_rs::types::TypeId;
use quack_rs::vector::{VectorReader, VectorWriter};
use libduckdb_sys::{duckdb_function_info, duckdb_vector, idx_t};

unsafe extern "C" fn varchar_to_int(
    info: duckdb_function_info,
    count: idx_t,
    input: duckdb_vector,
    output: duckdb_vector,
) -> bool {
    let cast_info = unsafe { CastFunctionInfo::new(info) };
    let reader = unsafe { VectorReader::from_vector(input, count as usize) };
    let mut writer = unsafe { VectorWriter::new(output) };

    for row in 0..count as usize {
        if !unsafe { reader.is_valid(row) } {
            unsafe { writer.set_null(row) };
            continue;
        }
        let s = unsafe { reader.read_str(row) };
        match s.parse::<i32>() {
            Ok(v) => unsafe { writer.write_i32(row, v) },
            Err(e) => {
                let msg = format!("cannot cast {:?} to INTEGER: {e}", s);
                if cast_info.cast_mode() == CastMode::Try {
                    // TRY_CAST: record a per-row error, which also sets the row to NULL
                    unsafe { cast_info.set_row_error(&msg, row as idx_t, output) };
                } else {
                    // Regular CAST: fail the whole query
                    cast_info.set_error(&msg);
                    return false;
                }
            }
        }
    }
    true
}

fn register(con: libduckdb_sys::duckdb_connection)
    -> Result<(), quack_rs::error::ExtensionError>
{
    unsafe {
        CastFunctionBuilder::new(TypeId::Varchar, TypeId::Integer)
            .function(varchar_to_int)
            .register(con)
    }
}
}

Implicit casts

Provide an implicit_cost to allow DuckDB to use the cast automatically in expressions where the types do not match:

#![allow(unused)]
fn main() {
use quack_rs::cast::CastFunctionBuilder;
use quack_rs::types::TypeId;
use libduckdb_sys::{duckdb_function_info, duckdb_vector, idx_t};
unsafe extern "C" fn my_cast(_: duckdb_function_info, _: idx_t, _: duckdb_vector, _: duckdb_vector) -> bool { true }
fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
unsafe {
    CastFunctionBuilder::new(TypeId::Varchar, TypeId::Integer)
        .function(my_cast)
        .implicit_cost(100) // lower = higher priority
        .register(con)
}
}
}

Extra info

Attach arbitrary data to a cast function using extra_info. This is useful for parameterising the cast behaviour (e.g., a rounding mode):

#![allow(unused)]
fn main() {
use quack_rs::cast::CastFunctionBuilder;
use quack_rs::types::TypeId;
use libduckdb_sys::{duckdb_function_info, duckdb_vector, idx_t};
use std::os::raw::c_void;
unsafe extern "C" fn my_cast(_: duckdb_function_info, _: idx_t, _: duckdb_vector, _: duckdb_vector) -> bool { true }
unsafe extern "C" fn my_destroy(_: *mut c_void) {}
fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
let mode = Box::into_raw(Box::new("round".to_string())).cast::<c_void>();
unsafe {
    CastFunctionBuilder::new(TypeId::Double, TypeId::BigInt)
        .function(my_cast)
        .implicit_cost(100)
        .extra_info(mode, Some(my_destroy))
        .register(con)
}
}
}

Inside the cast callback, retrieve the extra info with CastFunctionInfo::get_extra_info(). The pointee must be Send + Sync: DuckDB passes the same pointer to the callback on every thread that runs the cast, and the destructor runs on whichever thread releases it. CastFunctionBuilder is Send, so an unregistered builder may also run the destructor on the thread that drops it.

If register returns an error — a missing callback or type, a bare composite TypeId (use new_logical), a null connection, or a source or target type that is or contains ANY/INVALID, which DuckDB refuses — the builder still owns the extra info and runs its destructor exactly once. These cases are checked in Rust before DuckDB is called, because duckdb_register_cast_function rejects them before taking ownership of the pointer.

TRY_CAST vs CAST

Inside your callback, check CastFunctionInfo::cast_mode() to distinguish between the two modes:

ModeUser wroteExpected behaviour on error
CastMode::NormalCAST(x AS T)Call set_error and return false
CastMode::TryTRY_CAST(x AS T)Call set_row_error for each failed row, continue

set_row_error records the message and sets that row of the output to NULL (FlatVector::SetNull inside the C API); row must be less than the callback's count — DuckDB does not check it in release builds.

The return value does nothing in TRY_CAST mode. DuckDB discards it (execute_cast.cpp calls the cast and ignores the result), so returning false does not turn the chunk into NULLs: every row you did not null keeps whatever the output vector held — possibly a previous chunk's value. Null each failed row yourself. The one exception is a panic inside a cast_callback! body: the macro then sets every row of the chunk to NULL, since nothing the body wrote can be trusted.

In Normal mode, set a message before returning false. A cast_callback! body that sets none fails the query with "cast function failed without reporting an error message"; a hand-written extern "C" callback that sets none makes DuckDB report Conversion Error: followed by nothing. An empty message passed to set_error / set_row_error is replaced by a placeholder.

Working example

The examples/hello-ext extension registers two cast functions:

  • CAST(VARCHAR AS INTEGER) / TRY_CAST(VARCHAR AS INTEGER) — basic cast
  • CAST(DOUBLE AS BIGINT) — with implicit_cost(100) and extra_info for rounding mode

See examples/hello-ext/src/lib.rs for complete, copy-paste-ready references.

Complex source and target types

For casts involving complex types like DECIMAL(18, 3) or LIST(VARCHAR), use the new_logical constructor instead of new:

#![allow(unused)]
fn main() {
use quack_rs::cast::CastFunctionBuilder;
use quack_rs::types::{LogicalType, TypeId};
use libduckdb_sys::{duckdb_function_info, duckdb_vector, idx_t};
unsafe extern "C" fn my_cast(_: duckdb_function_info, _: idx_t, _: duckdb_vector, _: duckdb_vector) -> bool { true }
fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
unsafe {
    CastFunctionBuilder::new_logical(
        LogicalType::list(TypeId::Varchar),   // LIST(VARCHAR) source
        LogicalType::list(TypeId::Integer),   // LIST(INTEGER) target
    )
    .function(my_cast)
    .register(con)
}
}
}

The source() and target() accessor methods return Option<TypeId> — they return None when the type was set via new_logical (since a LogicalType cannot always be expressed as a simple TypeId).

API reference

NULL Handling in Functions

This page explains how NULL inputs reach the scalar and aggregate functions of a DuckDB extension written in Rust, and how to give them SQL NULL semantics.

The one thing to take away: for a scalar function, DefaultNullHandling does not make DuckDB return NULL for you. Your callback is invoked for NULL rows too, and if it writes a value there, that value is the answer. Call DataChunk::propagate_nulls, or use ScalarFunctionBuilder::map1 / map2, which do it for you. An aggregate's update likewise receives NULL rows under either setting.


What DuckDB actually does

DuckDB's FunctionNullHandling has two settings, and quack-rs mirrors them as NullHandling. The names suggest that the default makes the engine handle NULL propagation. For functions registered through the C API it does not: a scalar function's output for a NULL row is kept as written, and an aggregate's update is handed NULL rows (see Aggregate functions). For scalars the result is silent wrong answers rather than an error.

For scalar functions, two pieces of DuckDB source settle it (quoted from v1.5.4):

src/main/capi/scalar_function-c.cpp — the C API bridge calls your callback for the whole flattened chunk and never looks at the result's validity:

void CAPIScalarFunction(DataChunk &input, ExpressionState &state, Vector &result) {
    ...
    input.Flatten();
    ...
    c_bind_info.info.function(c_function_info, c_input, c_result);
    if (!function_info.success) {
        throw InvalidInputException(function_info.error);
    }
    ...
}

src/execution/expression_executor/execute_function.cpp — the only NULL check is a debug-only assertion that your function already did the right thing:

static void VerifyNullHandling(const BoundFunctionExpression &expr, DataChunk &args, Vector &result) {
#ifdef DEBUG
    if (args.data.empty() || expr.function.GetNullHandling() != FunctionNullHandling::DEFAULT_NULL_HANDLING) {
        return;
    }
    // ... D_ASSERT(!result_data.validity.RowIsValid(idx));
#endif
}

Every DuckDB a user installs is a release build, so that assertion is compiled out.

Why the obvious test passes anyway

SELECT my_func(NULL);   -- NULL, even for a broken function

A literal NULL is constant-folded during binding: DuckDB evaluates the expression once, sees a NULL argument to a DEFAULT_NULL_HANDLING function, and substitutes NULL without calling anything. The wrong answers only appear when the argument comes from a column:

CREATE TABLE t(i BIGINT);
INSERT INTO t VALUES (1), (NULL), (3);
SELECT i, my_func(i) FROM t;

Against DuckDB 1.5.4, a function that unconditionally writes 999 returns:

imy_func(i)
1999
NULL999 ← not NULL
3999

quack-rs pins this behaviour in tests/ffi_roundtrip.rs, so a future DuckDB that starts propagating will show up as a test failure rather than as a surprise.


Doing it right

The safe route: typed scalar functions

ScalarFunctionBuilder::map1 / map2 take an ordinary Rust closure and handle validity for you — a NULL argument short-circuits to a NULL result without ever calling your code:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
fn live_connection() -> libduckdb_sys::duckdb_connection {
    std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
    let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess);
    }
    con
}
/// First column of the first row, as BIGINT; `None` for NULL.
fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    let reader = unsafe { chunk.reader(0) };
    unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) }
}
let con = live_connection();
let run = || -> Result<(), ExtensionError> { unsafe {
ScalarFunctionBuilder::map1("double_it", |x: i64| x * 2)?
    .register(con)?;
} Ok(()) };
run().unwrap();
assert_eq!(query_i64(con, "SELECT double_it(21)"), Some(42));
// From a column, not a literal: `double_it(NULL::BIGINT)` is constant-folded
// to NULL whatever the function does, so it cannot tell a broken one apart.
assert_eq!(query_i64(con, "SELECT double_it(i) FROM (VALUES (NULL::BIGINT)) t(i)"), None);
}

Use map1_opt / map2_opt when the function needs to see NULLs; those register SpecialNullHandling for you and hand the closure Option<T>.

The raw route: propagate_nulls

When you write the extern "C" callback yourself, restore SQL semantics with one call at the end:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
fn live_connection() -> libduckdb_sys::duckdb_connection {
    std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
    let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess);
    }
    con
}
/// First column of the first row, as BIGINT; `None` for NULL.
fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    let reader = unsafe { chunk.reader(0) };
    unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) }
}
quack_rs::scalar_callback!(double_it, |_info, input, output| {
    let chunk = unsafe { DataChunk::from_raw(input) };
    let reader = unsafe { chunk.reader(0) };
    let mut writer = unsafe { VectorWriter::from_vector(output) };
    for row in 0..chunk.size() {
        unsafe { writer.write_i64(row, reader.read_i64(row) * 2) };
    }
    // Without this, double_it(i) for a NULL `i` from a column is 0, not NULL.
    unsafe { chunk.propagate_nulls(&mut writer) };
});
let con = live_connection();
unsafe { ScalarFunctionBuilder::new("double_it").param(TypeId::BigInt).returns(TypeId::BigInt)
    .function(double_it).register(con).unwrap(); }
assert_eq!(query_i64(con, "SELECT double_it(21)"), Some(42));
// From a column, not a literal: `double_it(NULL::BIGINT)` is constant-folded
// to NULL whatever the function does, so it cannot tell a broken one apart.
assert_eq!(query_i64(con, "SELECT double_it(i) FROM (VALUES (NULL::BIGINT)) t(i)"), None);
}

propagate_nulls resolves each column's validity pointer once and marks the output NULL wherever any input column is NULL. A column with no validity mask has no NULLs and costs nothing. DataChunk::any_null(row) is the per-row form when you need the decision inline.


NullHandling enum

#![allow(unused)]
fn main() {
use quack_rs::types::NullHandling;

// Default: the function promises NULL in -> NULL out.
// Scalar: you must keep that promise (see above).
// Aggregate: `update` still receives NULL rows; skip them yourself.
NullHandling::DefaultNullHandling;

// The function means to see NULLs and may return non-NULL for them.
NullHandling::SpecialNullHandling;
}

Aggregate functions

Aggregates behave like scalar functions here: under either setting, update receives every row of the chunk, NULL rows included. CAPIAggregateUpdate in DuckDB's aggregate_function-c.cpp flattens the inputs and passes the whole chunk through; nothing on the way filters by validity. An aggregate that ignores NULLs skips them itself:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe fn demo(chunk: &DataChunk, reader: &VectorReader) {
for row in 0..chunk.size() {
    if !unsafe { reader.is_valid(row) } {
        continue; // a NULL row: its data slot holds no meaningful value
    }
    // ... accumulate reader.read_i64(row) into *states.add(row) ...
}
}
}

SpecialNullHandling declares that the aggregate may return non-NULL for NULL input (a count_with_nulls, say). For an aggregate DuckDB reads the setting in one place only — the correlated-subquery decorrelator, to pick an INNER or LEFT join — and no query we tried (correlated scalar subqueries, with and without arithmetic or coalesce around the aggregate, LATERAL, a correlated subquery in WHERE) answered differently under the two settings on DuckDB 1.5.5. Set it anyway when it is true; it is what DuckDB expects.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct CountState { count: i64 }
impl AggregateState for CountState {}
unsafe extern "C" fn my_update(_: duckdb_function_info, _: duckdb_data_chunk, _: *mut duckdb_aggregate_state) {}
unsafe extern "C" fn my_combine(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: *mut duckdb_aggregate_state, _: idx_t) {}
unsafe extern "C" fn my_finalize(_: duckdb_function_info, _: *mut duckdb_aggregate_state, _: duckdb_vector, _: idx_t, _: idx_t) {}
unsafe fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
use quack_rs::aggregate::AggregateFunctionBuilder;
use quack_rs::types::{TypeId, NullHandling};

unsafe {
    AggregateFunctionBuilder::new("count_with_nulls")
        .param(TypeId::BigInt)
        .returns(TypeId::BigInt)
        .null_handling(NullHandling::SpecialNullHandling)
        .ffi_state::<CountState>()
        .update(my_update)   // counts rows whose value is NULL, too
        .combine(my_combine)
        .finalize(my_finalize)
        .register(con)?;
}
Ok(())
}
}

Empty groups in a correlated subquery

One difference from an uncorrelated query holds under both settings. In SELECT (SELECT my_count(x) FROM t2 WHERE t2.k = t1.k) FROM t1, an outer row with no matching t2 rows gets NULL: the decorrelated plan joins the aggregate's groups back to the outer rows, and an outer row with no group never has an empty state finalized. DuckDB rewrites that NULL to 0 for its own count and count(*) only. A count-like aggregate of yours that returns 0 for empty input therefore returns NULL here; write coalesce((SELECT ...), 0) if the query needs 0. (SELECT my_count(x) FROM t2 WHERE false, uncorrelated, does finalize an empty state and returns 0.)


When to use special NULL handling

Use caseNULL handlingWho propagates
Scalar function, NULL in → NULL outDefaultNullHandlingyou (propagate_nulls, or map1/map2)
Scalar function that inspects NULLs (COALESCE-like, IS_NULL-like)SpecialNullHandlingyou
Aggregate, ignore NULL rowsDefaultNullHandling (the default)you (skip rows where is_valid is false)
Aggregate that counts NULLsSpecialNullHandlingyou

If you don't call .null_handling(), DefaultNullHandling is used.

SQL Macros

A DuckDB SQL macro packages a SQL expression or query as a named function, with no FFI callback behind it. This page shows how to create scalar and table macros from a Rust extension with quack-rs' SqlMacro: you give the macro body as a string and call .register(con), which runs CREATE OR REPLACE MACRO.


Two macro types

TypeSQL generatedReturns
ScalarCREATE OR REPLACE MACRO "name"("a", "b") AS (expression)one value per row
TableCREATE OR REPLACE MACRO "name"("a", "b") AS TABLE querya result set

Scalar macros

A scalar macro wraps a SQL expression. Think of it as a parameterized SQL alias:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_connection, DuckDBSuccess};
use quack_rs::error::ExtensionError;
fn live_connection() -> duckdb_connection {
    std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
    let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), DuckDBSuccess);
    }
    con
}
fn query_i64(con: duckdb_connection, sql: &str) -> i64 {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    unsafe { chunk.reader(0).read_i64(0) }
}
use quack_rs::sql_macro::SqlMacro;

fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        // clamp(x, lo, hi) → greatest(lo, least(hi, x))
        SqlMacro::scalar("clamp", &["x", "lo", "hi"], "greatest(lo, least(hi, x))")?
            .register(con)?;

        // golden_ratio() → 1.61803398874989
        SqlMacro::scalar("golden_ratio", &[], "1.61803398874989")?
            .register(con)?;

        // safe_div(a, b) → CASE WHEN b = 0 THEN NULL ELSE a / b END
        SqlMacro::scalar(
            "safe_div",
            &["a", "b"],
            "CASE WHEN b = 0 THEN NULL ELSE a / b END",
        )?
        .register(con)?;
    }
    Ok(())
}
let con = live_connection();
register(con).unwrap();
assert_eq!(query_i64(con, "SELECT clamp(9, 1, 5)::BIGINT"), 5);
assert_eq!(query_i64(con, "SELECT clamp(-3, 1, 5)::BIGINT"), 1);
assert_eq!(query_i64(con, "SELECT (golden_ratio() * 1000)::BIGINT"), 1618);
assert_eq!(query_i64(con, "SELECT count(*) FROM (SELECT safe_div(1, 0) AS v) WHERE v IS NULL"), 1);
}

Use in DuckDB:

SELECT clamp(rating, 1, 5) FROM reviews;
SELECT safe_div(revenue, orders) FROM monthly_stats;

Table macros

A table macro wraps a SQL query that returns rows:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_connection, DuckDBSuccess};
use quack_rs::error::ExtensionError;
use quack_rs::sql_macro::SqlMacro;
fn live_connection() -> duckdb_connection {
    std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
    let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), DuckDBSuccess);
    }
    con
}
fn query_i64(con: duckdb_connection, sql: &str) -> i64 {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    unsafe { chunk.reader(0).read_i64(0) }
}
fn register(con: duckdb_connection) -> Result<(), ExtensionError> {
unsafe {
    // active_users(tbl) → SELECT * FROM query_table(tbl) WHERE active = true
    SqlMacro::table(
        "active_users",
        &["tbl"],
        "SELECT * FROM query_table(tbl) WHERE active = true",
    )?
    .register(con)?;

    // recent_orders(days) → last N days of orders
    SqlMacro::table(
        "recent_orders",
        &["days"],
        "SELECT * FROM orders WHERE order_date >= current_date - INTERVAL (days) DAY",
    )?
    .register(con)?;
}
Ok(())
}
let con = live_connection();
unsafe { quack_rs::query::execute(con, "CREATE TABLE users AS SELECT * FROM (VALUES (1, true), (2, false), (3, true)) t(id, active)") }.unwrap();
unsafe { quack_rs::query::execute(con, "CREATE TABLE orders AS SELECT current_date - 3 AS order_date UNION ALL SELECT current_date - 30") }.unwrap();
register(con).unwrap();
assert_eq!(query_i64(con, "SELECT count(*) FROM recent_orders(7)"), 1);
assert_eq!(query_i64(con, "SELECT count(*) FROM active_users(users)"), 2);
}

A table macro's body is bound when the macro is created, so every table it names directly (orders above) must already exist when register runs, or registration fails with a catalog error. To take a table as a parameter, read it through query_table(tbl): a bare FROM tbl looks for a table literally named tbl.

Use in DuckDB:

SELECT * FROM active_users(users);
SELECT count(*) FROM recent_orders(7);

Inspecting the generated SQL

to_sql() returns the CREATE OR REPLACE MACRO statement without requiring a live connection. Use it for logging, debugging, or assertions in tests:

#![allow(unused)]
fn main() {
use quack_rs::sql_macro::SqlMacro;
let m = SqlMacro::scalar("add", &["a", "b"], "a + b")?;
assert_eq!(
    m.to_sql(),
    r#"CREATE OR REPLACE MACRO "add"("a", "b") AS (a + b)"#
);

let t = SqlMacro::table("active_users", &["tbl"], "SELECT * FROM query_table(tbl) WHERE active = true")?;
assert_eq!(
    t.to_sql(),
    r#"CREATE OR REPLACE MACRO "active_users"("tbl") AS TABLE SELECT * FROM query_table(tbl) WHERE active = true"#
);
Ok::<(), quack_rs::error::ExtensionError>(())
}

Name and parameter validation

Macro names are validated with validate_function_name, the same rules as function names, and parameter names with validate_parameter_name. A name must:

  • start with an ASCII letter or underscore, followed by ASCII letters, digits or underscores;
  • be at most 256 characters long;
  • not be a DuckDB keyword that cannot be used in that position unquoted. For a macro, that is a keyword it cannot be called by (order, coalesce); for a parameter, one the body cannot refer to it by (order, left). The two lists differ: a parameter may be called columns, which a macro may not.

Case is not restricted — DuckDB identifiers are case-insensitive, so a macro registered as MyMacro is callable as mymacro(...) or MYMACRO(...).

#![allow(unused)]
fn main() {
use quack_rs::sql_macro::SqlMacro;
assert!(SqlMacro::scalar("1f", &[], "1").is_err());       // ❌ starts with a digit
assert!(SqlMacro::scalar("my-macro", &[], "1").is_err()); // ❌ hyphen
assert!(SqlMacro::scalar("f", &["a b"], "1").is_err());   // ❌ space in param
assert!(SqlMacro::scalar("f", &["order"], "1").is_err()); // ❌ reserved keyword as param
assert!(SqlMacro::scalar("f", &["left"], "1").is_err());  // ❌ a body cannot refer to it
assert!(SqlMacro::scalar("columns", &[], "1").is_err());  // ❌ cannot be called as a macro
assert!(SqlMacro::scalar("f", &["columns"], "1").is_ok()); // ✅ fine as a parameter
assert!(SqlMacro::scalar("MyMacro", &[], "1").is_ok());   // ✅ mixed case allowed
assert!(SqlMacro::scalar("f", &["X"], "1").is_ok());      // ✅ mixed-case param allowed
assert!(SqlMacro::scalar("f", &["_x"], "1").is_ok());     // ✅ underscore prefix allowed
}

SQL injection safety

Macro and parameter names are restricted to ASCII letters, digits and underscores, preventing SQL injection at the identifier level. to_sql() additionally emits every name as a double-quoted identifier ("name") — the validated character set cannot contain ", so no escaping is needed. Quoting does not make names case-sensitive in DuckDB.

The body (expression or query) is your own extension code — it is included verbatim. Never build macro bodies from untrusted user input. register does refuse a body that turns the statement into several (1); DROP TABLE t; SELECT (1) — it counts statements with DuckDB's own parser (duckdb_extract_statements) and executes nothing if there is more than one — but a body can still change the meaning of the single statement it is part of.

A scalar body containing -- gets a newline before the closing parenthesis, so a trailing line comment ("x + 1 -- plus one") does not comment it out.


Where a macro lives

A macro is not a function registration: register runs CREATE OR REPLACE MACRO, so the macro is an ordinary catalog object in the connection's default database and schema — the user's database:

  • It persists. In a database file the macro is still there in the next session, even if the extension is never loaded again. Reloading is fine: OR REPLACE replaces it.
  • It needs a writable database. On a read-only database the CREATE fails; an entry point that propagates that error with ? makes LOAD fail. Decide whether a macro is essential there.
  • It silently replaces a user's macro of the same name. Prefix macro names with your extension's name.
  • It can shadow a built-in. A macro named abs in the default schema is found before the built-in abs, so abs(-1) calls the macro.

How it works under the hood

SqlMacro::register counts the statements in to_sql() with duckdb_extract_statements, refuses more than one, and executes the CREATE OR REPLACE MACRO statement via duckdb_query.

The query result is zero-initialized before duckdb_query, any error message is read with duckdb_result_error, and duckdb_destroy_result runs on success and failure alike.


Choosing between macros and scalar functions

ScenarioUse
Logic expressible in SQLSQL macro — simpler, no FFI
Logic needs Rust code (algorithms, external crates, etc.)Scalar function
Simple expressionsSQL macro (expanded inline, no callback per chunk)
Type-specific overloadsScalar function set (ScalarFunctionSetBuilder)
Returning a tableSQL table macro

Copy Functions

A DuckDB copy function implements a custom file format for the COPY statement. This page shows how to register one from a Rust extension with quack-rs' CopyFunctionBuilder.

Requires the duckdb-1-5 feature flag (DuckDB 1.5.0+).

A format can support writing, reading, or both:

DirectionYou supplyDuckDB calls
COPY t TO 'f' (FORMAT my_format)bind + sink + finalize (and optionally global_init)those callbacks
COPY t FROM 'f' (FORMAT my_format)copy_from(table_function)your table function's bind, init and scan

duckdb_register_copy_function decides which directions a format supports by looking at the sink and the reader independently, so a read-only format leaves the writing callbacks unset entirely. Set bind, sink and finalize together or not at all: register refuses a builder with only some of them, and one that implements neither direction.

Lifecycle (COPY … TO)

  1. Bind — called when the statement is bound: once for a plain COPY, and again on every EXECUTE of a prepared one. Inspect output columns, configure the export.
  2. Global init — called once per output file: once for a plain COPY, once per file with PER_THREAD_OUTPUT or PARTITION_BY. Open the file, allocate that file's global state. With USE_TMP_FILE the path is a temporary name that DuckDB renames afterwards.
  3. Sink — called for each data chunk; with PER_THREAD_OUTPUT or PARTITION_BY, from several threads at once (see Threads). Write rows to the output.
  4. Finalize — called once per output file, after its last sink. Flush buffers, close the file. It is not called when a sink reports an error, so release resources in the global state's destructor as well.

Builder API

#![allow(unused)]
fn main() {
use quack_rs::copy_function::CopyFunctionBuilder;

fn demo(
    my_bind_fn: quack_rs::copy_function::CopyBindFn,
    my_global_init_fn: quack_rs::copy_function::CopyGlobalInitFn,
    my_sink_fn: quack_rs::copy_function::CopySinkFn,
    my_finalize_fn: quack_rs::copy_function::CopyFinalizeFn,
) -> Result<(), quack_rs::error::ExtensionError> {
let builder = CopyFunctionBuilder::try_new("my_format")?
    .bind(my_bind_fn)
    .global_init(my_global_init_fn)
    .sink(my_sink_fn)
    .finalize(my_finalize_fn);

// Register on a connection, for example in the entry point:
// unsafe { builder.register(con)?; }
Ok(())
}
}

COPY … FROM

Reading is a table function, attached to the copy function rather than registered on its own. Build it with TableFunctionBuilder::build_handle, then hand it to copy_from:

#![allow(unused)]
fn main() {
use quack_rs::copy_function::CopyFunctionBuilder;
use quack_rs::table::TableFunctionBuilder;
use quack_rs::types::TypeId;

fn demo(
    bind: quack_rs::table::BindFn,
    init: quack_rs::table::InitFn,
    scan: quack_rs::table::ScanFn,
) -> Result<(), quack_rs::error::ExtensionError> {
// SAFETY: the callbacks match their declared signatures.
let reader = unsafe {
    TableFunctionBuilder::new("my_format")   // name it after the format
        .param(TypeId::Varchar)              // the file path — exactly one
        .named_param("skip_rows", TypeId::BigInt) // a COPY option
        .bind(bind)
        .init(init)
        .scan(scan)
        .build_handle()
}?;

let format = CopyFunctionBuilder::try_new("my_format")?.copy_from(reader)?;
// unsafe { format.register(con)?; }
Ok(())
}
}

Four things about the reader are not like an ordinary table function:

  • The file path is positional parameter 0, always VARCHAR. duckdb.h requires the function to declare exactly that one parameter, and DuckDB does not check it — copy_from does, and returns an error naming the mismatch.
  • COPY options are named parameters. (FORMAT my_format, SKIP_ROWS 1) arrives as skip_rows; matching is case-insensitive. An option the function never declared is a binder error before your bind callback runs — and the message names the table function, which is why the example gives it the format's name.
  • Those options arrive uncast. DuckDB passes each value as written, not cast to the declared type: SKIP_ROWS 'abc' reaches a BIGINT parameter as the VARCHAR 'abc', and SKIP_ROWS 3 as an INTEGER. An option written without a value ((FORMAT my_format, HEADER)) is not passed at all. Check Value::type_id() before trusting a value — as_i64_or(0) on 'abc' quietly returns the default.
  • The schema is already fixed, because COPY … FROM loads into an existing table. The bind callback must not call add_result_column (a typed reader built with with_state or with_bind_init fails its bind if it does). Read the target's schema instead:
#![allow(unused)]
fn main() {
use quack_rs::table::BindInfo;
fn demo(bind: &BindInfo) {
for i in 0..bind.result_column_count() {
    let name = bind.result_column_name(i);
    // SAFETY: `i` is in range, and this runs during the bind callback.
    let ty = unsafe { bind.result_column_type(i) };
    let _ = (name, ty);
}
}
}

Registering a format name twice

duckdb_register_copy_function drops a copy function whose name already exists — your earlier registration, another extension's, or a built-in format such as csv — and still reports success. register checks first and returns an error. There is no catalog view of copy functions, so the check asks the binder: it runs COPY (SELECT <missing column>) TO '' (FORMAT '<name>'), where a binder error means the format resolved and a catalog error means it did not. Nothing is written and no copy callback runs; a format name owned by an autoloadable extension (parquet, json) may be autoloaded by that lookup, exactly as the same COPY typed by a user would. Copy functions registered through the C API are never persisted, so reloading into a database file does not trip the check.

Options

CopyBindInfo::options() returns the COPY … TO options as one STRUCT value (None only if DuckDB returns a null handle). How DuckDB 1.5.5 builds it:

  • Option names are upper-cased: compression 'zstd' arrives as COMPRESSION. FORMAT itself is not among them.
  • With no options besides FORMAT, the value is SQL NULL, not an empty STRUCT — check is_sql_null() first.
  • An option given without a value (HEADER) is a NULL field.
  • Several values (LST (1, 2)) arrive as a LIST, or an unnamed STRUCT when their types differ.
  • An explicit NULL value is rejected by the binder before your callback runs.
  • Field order follows DuckDB's internal hash map, not the statement — look fields up by name:
#![allow(unused)]
fn main() {
use quack_rs::copy_function::CopyBindInfo;
fn demo(bind: &CopyBindInfo) -> Option<String> {
let options = bind.options()?;
if options.is_sql_null() {
    return None; // COPY ... (FORMAT my_format) with no other options
}
let names = options.struct_field_names();
let idx = names.iter().position(|n| n == "COMPRESSION")?;
options.struct_child(idx)?.as_str().ok()
}
}

Threads

The data pointers are untyped, so the compiler cannot check this for you:

  • extra_info lives as long as the database and is read from every connection's thread — treat it as T: Send + Sync.
  • Bind data is shared by every sink call. COPY … TO with PER_THREAD_OUTPUT or PARTITION_BY runs the sink on several threads at once — treat it as T: Send + Sync and never mutate it without a lock.
  • Global state is one per output file: per thread with PER_THREAD_OUTPUT, per partition with PARTITION_BY (reachable from more than one thread). Guard any mutation with a Mutex unless you know neither option is in use.

Callback signatures

PhaseSignature
Bindunsafe extern "C" fn(info: duckdb_copy_function_bind_info)
Global initunsafe extern "C" fn(info: duckdb_copy_function_global_init_info)
Sinkunsafe extern "C" fn(info: duckdb_copy_function_sink_info, chunk: duckdb_data_chunk)
Finalizeunsafe extern "C" fn(info: duckdb_copy_function_finalize_info)

Callback info wrappers

Each phase provides an ergonomic wrapper type around its raw info handle. Wrap the handle at the top of your callback to access helper methods:

CopyBindInfo

MethodDescription
column_count()Number of output columns
column_type(index)LogicalType of the column at index, or None if out of range
options()The COPY … TO options, as one STRUCT Value (see Options)
get_extra_info()Extra-info pointer set on the copy function
set_bind_data(data, destroy)Store bind data and its destructor
set_error(message)Report a bind-time error
get_client_context()Returns a ClientContext for catalog/config access

CopyGlobalInitInfo

MethodDescription
get_bind_data()Retrieve the bind data pointer
get_extra_info()Extra-info pointer set on the copy function
get_file_path()Output file path for the COPY operation; an error if it is not valid UTF-8
get_file_path_bytes()The same path as its exact bytes, for a path that is not UTF-8
set_global_state(state, destroy)Store global state and its destructor
set_error(message)Report an init-time error
get_client_context()Returns a ClientContext

CopySinkInfo

MethodDescription
get_bind_data()Retrieve the bind data pointer
get_extra_info()Extra-info pointer set on the copy function
get_global_state()Retrieve the global state pointer
set_error(message)Report a sink-time error
get_client_context()Returns a ClientContext

CopyFinalizeInfo

MethodDescription
get_bind_data()Retrieve the bind data pointer
get_extra_info()Extra-info pointer set on the copy function
get_global_state()Retrieve the global state pointer
set_error(message)Report a finalize-time error
get_client_context()Returns a ClientContext

All four wrappers are re-exported from quack_rs::copy_function:

#![allow(unused)]
fn main() {
use quack_rs::copy_function::{CopyBindInfo, CopyGlobalInitInfo, CopySinkInfo, CopyFinalizeInfo};
}

Reading & Writing Vectors

DuckDB passes data to and from your extension as vectors: columnar arrays of typed values, each with a separate validity (NULL) bitmap. VectorReader and VectorWriter give typed access to these vectors from scalar, aggregate and table function callbacks; DataChunk, StructReader, StructWriter and ChunkWriter build on them.


VectorReader

Construction

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(input: duckdb_data_chunk, column_index: usize) {
// In a scalar function callback:
let reader = unsafe { VectorReader::new(input, column_index) };

// In an aggregate update callback:
let reader = unsafe { VectorReader::new(input, 0) };   // first column
}
}

VectorReader::new takes the duckdb_data_chunk and a zero-based column index. The reader holds raw pointers into the chunk, so it must not outlive the callback.

Row count

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader) {
let n = reader.row_count();   // number of rows in this chunk
}
}

Chunk sizes vary. Always loop over 0..reader.row_count(); never assume a fixed size.

NULL check

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, writer: &mut VectorWriter) {
for row in 0..reader.row_count() {
if unsafe { !reader.is_valid(row) } {
    // row is NULL — skip or propagate NULL to output
    unsafe { writer.set_null(row) };
    continue;
}
}
}
}

Always check is_valid before reading. Reading a fixed-width value from a NULL row returns garbage data; reading a VARCHAR or BLOB from one can follow a stale pointer into freed memory (see NULL Handling & Strings).

Reading values

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, row: usize) {
let i: i8  = unsafe { reader.read_i8(row) };
let i: i16 = unsafe { reader.read_i16(row) };
let i: i32 = unsafe { reader.read_i32(row) };
let i: i64 = unsafe { reader.read_i64(row) };
let u: u8  = unsafe { reader.read_u8(row) };
let u: u16 = unsafe { reader.read_u16(row) };
let u: u32 = unsafe { reader.read_u32(row) };
let u: u64 = unsafe { reader.read_u64(row) };
let f: f32 = unsafe { reader.read_f32(row) };
let f: f64 = unsafe { reader.read_f64(row) };
let b: bool = unsafe { reader.read_bool(row) };   // safe: uses u8 != 0
let s: &str = unsafe { reader.read_str(row) };    // handles inline + pointer format
let iv = unsafe { reader.read_interval(row) };    // returns DuckInterval

// Temporal and binary types (v0.10.0+):
let d: i32 = unsafe { reader.read_date(row) };      // days since epoch
let ts: i64 = unsafe { reader.read_timestamp(row) }; // microseconds since epoch
let t: i64 = unsafe { reader.read_time(row) };       // microseconds since midnight
let blob: &[u8] = unsafe { reader.read_blob(row) };  // binary data
let uuid: u128 = unsafe { reader.read_uuid(row) };   // UUID's textual 128 bits
}
}

VectorWriter

Construction

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(output: duckdb_vector, result: duckdb_vector) {
// In a scalar function callback:
let mut writer = unsafe { VectorWriter::new(output) };

// In an aggregate finalize callback:
let mut writer = unsafe { VectorWriter::new(result) };
}
}

Writing values

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
use quack_rs::interval::DuckInterval;
fn demo(writer: &mut VectorWriter, row: usize, s: &str, interval: DuckInterval,
    days_since_epoch: i32, micros_since_epoch: i64, micros_since_midnight: i64,
    bytes: Vec<u8>, uuid_bits: u128) {
unsafe { writer.write_i8(row, -8) };
unsafe { writer.write_i16(row, -16) };
unsafe { writer.write_i32(row, -32) };
unsafe { writer.write_i64(row, -64) };
unsafe { writer.write_u8(row, 8) };
unsafe { writer.write_u16(row, 16) };
unsafe { writer.write_u32(row, 32) };
unsafe { writer.write_u64(row, 64) };
unsafe { writer.write_f32(row, 3.5) };
unsafe { writer.write_f64(row, 2.5) };
unsafe { writer.write_bool(row, true) };
unsafe { writer.write_varchar(row, s) };   // &str
unsafe { writer.write_str(row, s) };       // alias for write_varchar
unsafe { writer.write_interval(row, interval) };  // DuckInterval

// Temporal and binary types (v0.10.0+):
unsafe { writer.write_date(row, days_since_epoch) };
unsafe { writer.write_timestamp(row, micros_since_epoch) };
unsafe { writer.write_time(row, micros_since_midnight) };
unsafe { writer.write_blob(row, &bytes) };
unsafe { writer.write_uuid(row, uuid_bits) };        // UUID's textual 128 bits
}
}

write_varchar and write_blob panic for a value longer than vector::string::MAX_STRING_LEN (u32::MAX bytes, DuckDB's string length limit) rather than store a truncated one. Inside scalar_callback! and the typed scalar constructors the panic becomes a SQL error. To handle the error yourself, use try_write_varchar / try_write_blob, which return Result<(), ExtensionError> and write nothing on error.

UUID is not stored as you'd expect

A UUID column is physically a HUGEINT, but the 128 bits in the vector are not the bits you see in the text form: DuckDB flips the top bit so that comparing the signed integers orders UUIDs the same way comparing their strings does.

SELECT '11111111-2222-3333-4444-555555555555'::UUID
  read_i128 (raw storage) : 0x91111111222233334444555555555555
  read_uuid (textual bits): 0x11111111222233334444555555555555

read_uuid / write_uuid apply the flip for you and speak in textual bits (u128) — the same convention as Value::uuid / Value::as_uuid and every Rust Uuid type. Reach for read_i128 / write_i128 only when you want the raw storage, and use quack_rs::vector::{uuid_from_storage, uuid_to_storage} to convert.

Writing NULL

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(writer: &mut VectorWriter, row: usize) {
unsafe { writer.set_null(row) };
}
}

Pitfall L4: set_null calls duckdb_vector_ensure_validity_writable automatically before duckdb_vector_get_validity. A vector with no NULLs yet usually has no validity mask, so without that call get_validity returns NULL and duckdb_validity_set_row_invalid silently does nothing — the row you meant to be NULL reads back as a valid value. VectorWriter::set_null handles this correctly. See Pitfall L4.

NULL rows of STRUCT and ARRAY outputs

For a STRUCT output, set_null(row) (and set_null_range, and DataChunk::propagate_nulls, which uses it) also nulls that row in every field, recursively; for an ARRAY of size n it nulls child rows row * n .. row * n + n. This mirrors DuckDB's internal FlatVector::SetNull, and it matters: struct_extract / s.a reads the field vector without looking at the parent, so a NULL struct row whose fields were left valid returns the stale field value. StructWriter::set_row_null(row) does the same from a StructWriter. LIST / MAP elements are not touched (as in DuckDB).

To reuse such a row, call set_valid(row) first: on a row that is NULL it also marks valid everything set_null nulled below it. Then write the fields, and any field NULLs after that. On a row that is already valid, set_valid leaves the fields alone, so marking a row valid after writing its fields keeps their NULLs.

Clearing NULL (v0.11.0+)

To undo a previous set_null call and mark a row as valid again:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(writer: &mut VectorWriter, row: usize) {
unsafe { writer.set_valid(row) };
}
}

Like set_null, set_valid calls ensure_validity_writable first.


DataChunk

DataChunk wraps a duckdb_data_chunk handle and gives access to its vectors and row count without raw FFI calls:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::data_chunk::DataChunk;

unsafe extern "C" fn my_scan(info: duckdb_function_info, output: duckdb_data_chunk) {
    let chunk = unsafe { DataChunk::from_raw(output) };
    let mut writer = unsafe { chunk.writer(0) };    // VectorWriter for column 0
    unsafe { writer.write_i64(0, 42) };
    unsafe { chunk.set_size(1) };                   // set output row count
}
}

Methods:

  • size() — current row count
  • set_size(n) — set row count (0 = end of stream)
  • column_count() — number of columns
  • vector(col) — raw duckdb_vector handle
  • writer(col) — VectorWriter for a column
  • reader(col) — VectorReader for a column
  • struct_writer(col, field_count) — StructWriter for a STRUCT output column
  • struct_reader(col, field_count) — StructReader for a STRUCT input column
  • struct_field_reader(col, field) — VectorReader for a specific STRUCT field
  • any_null(row) — whether any column is NULL at row
  • propagate_nulls(&mut writer) — mark each output row NULL where any input column is NULL
  • into_chunk_writer() — convert to ChunkWriter, which calls set_size on drop

StructWriter / StructReader

For STRUCT columns, creating a VectorWriter or VectorReader for each field by hand is verbose. StructWriter and StructReader create one per field at construction:

#![allow(unused)]
fn main() {
use quack_rs::data_chunk::DataChunk;
struct Output { success: bool, data: String, count: i64, day: i32, payload: Vec<u8> }
fn demo(chunk: &DataChunk, row: usize, result: &Output) {
// Writing a 5-field STRUCT output:
let mut sw = unsafe { chunk.struct_writer(0, 5) };
unsafe {
    sw.write_bool(row, 0, result.success);
    sw.write_varchar(row, 1, &result.data);
    sw.write_i64(row, 2, result.count);
    sw.write_date(row, 3, result.day);
    sw.write_blob(row, 4, &result.payload);
}

// Reading a 3-field STRUCT input:
let sr = unsafe { chunk.struct_reader(0, 3) };
for row in 0..chunk.size() {
    let name = unsafe { sr.read_str(row, 0) };
    let age = unsafe { sr.read_i32(row, 1) };
    let active = unsafe { sr.read_bool(row, 2) };
}
}
}

ChunkWriter

ChunkWriter wraps an output duckdb_data_chunk and counts the rows handed out by next_row. It calls set_size with that count on drop, so the row count cannot be forgotten or set wrongly. next_row returns None once the chunk holds duckdb_vector_size() rows:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::data_chunk::DataChunk;
struct Item { name: String, value: i64 }
fn demo(output: duckdb_data_chunk, data: &[Item]) {
let mut cw = unsafe { DataChunk::from_raw(output).into_chunk_writer() };
for item in data {
    let Some(row) = cw.next_row() else { break };   // chunk is full
    unsafe { cw.writer(0).write_varchar(row, &item.name) };
    unsafe { cw.writer(1).write_i64(row, item.value) };
}
// set_size called automatically when `cw` is dropped
}
}

ValidityBitmap

For advanced NULL handling beyond VectorWriter::set_null, use ValidityBitmap directly:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
fn demo(some_vector: duckdb_vector, row: usize) {
use quack_rs::vector::ValidityBitmap;

// Writing NULLs:
let mut bitmap = unsafe { ValidityBitmap::ensure_writable(some_vector) };
unsafe { bitmap.set_row_invalid(row as u64) };   // mark as NULL
unsafe { bitmap.set_row_valid(row as u64) };     // mark as non-NULL

// Reading NULLs:
let bitmap = unsafe { ValidityBitmap::get_read_only(some_vector) };
let is_valid = unsafe { bitmap.row_is_valid(row as u64) };
}
}

ValidityBitmap is available in the prelude: use quack_rs::prelude::*.


Utility functions

The quack_rs::vector module provides two utility functions:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
fn demo(some_vector: duckdb_vector) {
use quack_rs::vector::{vector_size, vector_get_column_type};

// Rows per data chunk: 2048 unless DuckDB was built with another STANDARD_VECTOR_SIZE.
let size: u64 = vector_size();

// Returns the LogicalType of a vector (unsafe — requires a valid duckdb_vector).
let lt = unsafe { vector_get_column_type(some_vector) };
}
}

Memory layout details

DuckDB stores vector data as flat arrays. VectorReader and VectorWriter compute element addresses as base_ptr + row * stride:

[value0][value1][value2]...[valueN]   ← typed array
[validity bitmap]                      ← separate bit array, 1 bit per row

The validity bitmap is lazily allocated — it may be null if no NULLs have been written. This is why duckdb_vector_ensure_validity_writable must be called before duckdb_vector_get_validity when writing NULLs; VectorWriter and ValidityBitmap::ensure_writable do so.


Complete scalar function pattern

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::vector::{VectorReader, VectorWriter};
fn transform(v: i64) -> i64 { v }
unsafe extern "C" fn my_scalar(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };

    for row in 0..reader.row_count() {
        if unsafe { !reader.is_valid(row) } {
            unsafe { writer.set_null(row) };
            continue;
        }
        let value = unsafe { reader.read_i64(row) };
        unsafe { writer.write_i64(row, transform(value)) };
    }
}
}

Values & Parameter Extraction

DuckDB hands an extension single values — a table function's bind-time parameters, the options of a COPY statement, a folded constant expression — as duckdb_value handles, which are heap-allocated and must be destroyed after use. quack-rs wraps them in Value, which destroys the handle on drop and reads it through typed getters that return Option instead of aborting on SQL NULL. This page covers reading parameters with Value, building values, and nested (LIST, STRUCT, MAP) values.


The problem

Without Value, every parameter extraction requires three raw FFI calls and careful manual cleanup:

#![allow(unused)]
fn main() {
use libduckdb_sys::*;
unsafe fn demo(info: duckdb_bind_info) {
// Before: raw FFI — easy to leak memory, and `duckdb_get_int64` aborts the
// process if the argument is SQL NULL (DuckDB throws a C++ exception).
let mut param = unsafe { duckdb_bind_get_parameter(info, 0) };
let n = unsafe { duckdb_get_int64(param) };
unsafe { duckdb_destroy_value(&mut param) };  // forget this → memory leak
}
}

The solution: Value

Value wraps a duckdb_value handle and calls duckdb_destroy_value on drop:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
use quack_rs::table::BindInfo;

unsafe extern "C" fn my_bind(info: duckdb_bind_info) {
    let bind_info = unsafe { BindInfo::new(info) };

    // Value is RAII — automatically destroyed when dropped.
    // `as_i64` is `None` for SQL NULL; `as_i64_or` supplies a default.
    let n = unsafe { bind_info.get_parameter_value(0) }.as_i64_or(0);

    // Named parameters work the same way
    let path = unsafe { bind_info.get_named_parameter_value("path") }
        .as_str()
        .unwrap_or_default();
}
}

Typed extraction methods

MethodReads asRust type
as_str()any value, cast to textResult<String, ExtensionError>
as_blob()BLOB onlyResult<Vec<u8>, ExtensionError>
as_i8()TINYINTOption<i8>
as_i16()SMALLINTOption<i16>
as_i32()INTEGEROption<i32>
as_i64()BIGINTOption<i64>
as_i128()HUGEINTOption<i128>
as_u8()UTINYINTOption<u8>
as_u16()USMALLINTOption<u16>
as_u32()UINTEGEROption<u32>
as_u64()UBIGINTOption<u64>
as_u128()UHUGEINTOption<u128>
as_f32()FLOATOption<f32>
as_f64()DOUBLEOption<f64>
as_bool()BOOLEANOption<bool>
as_date(), as_time(), as_time_tz(), as_timestamp(), …the temporal typesOption<i32> / Option<i64> / Option<u64>
as_interval(), as_uuid()INTERVAL, UUIDOption<DuckInterval>, Option<u128>
as_decimal()DECIMAL only (no cast)Option<Decimal>
as_enum_index()ENUM only (no cast)Option<u64>

The scalar getters cast the way SQL's TRY_CAST does: a VARCHAR '42' reads as Some(42) through as_i64(), a DOUBLE 1.5 as Some(2) through as_i32(). They return None when:

  • the value is SQL NULL (for example my_func(n := NULL)),
  • the handle is null (a named parameter the caller did not supply),
  • the type is not a scalar (LIST, STRUCT, MAP, BLOB, ENUM, …), or
  • the cast fails ('abc', or a number out of range for the target), or
  • for the temporal getters, SQL would refuse the conversion: the time of an infinite timestamp (as_time() of 'infinity'::TIMESTAMP), a TIMESTAMP outside TIMESTAMP_NS's 1677–2262 range read with as_timestamp_ns(), or a result outside the target type's range.

No getter calls a duckdb_get_* function in the first three cases — those functions abort the process on a SQL NULL and crash on a null handle — and none modifies the value it reads (DuckDB's getters cast the value in place; quack-rs reads from a copy). The temporal cases are checked before the call too: DuckDB converts those pairs with a cast that throws a C++ exception instead of failing, which aborts the process from Rust.

as_blob() copies the bytes into an owned Vec<u8> without UTF-8 validation. It accepts only a BLOB: DuckDB's conversion of anything else to BLOB can throw. Use as_str() for text.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe fn demo(bind_info: BindInfo) -> Result<(), ExtensionError> {
let bytes = unsafe { bind_info.get_parameter_value(0) }.as_blob()?;
Ok(())
}
}

Defaulting variants

The integer (except as_u128), float, bool and string getters have an _or(default) variant that returns default wherever the plain getter returns None (or Err); as_str_or_default() returns an empty string:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
// Values are DuckDB objects: fill the dispatch table first.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
let val = Value::null_value();
let timeout = val.as_i64_or(30);       // 30 if NULL, absent, or not a number
let host = val.as_str_or("localhost");  // "localhost" if NULL or absent
let port = val.as_u16_or(5432);        // 5432 if NULL, absent, or out of range
assert_eq!((timeout, host.as_str(), port), (30, "localhost", 5432));
}

Checking for NULL

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe fn demo(bind_info: BindInfo) {
let val = unsafe { bind_info.get_named_parameter_value("limit") };
if val.is_null() {
    // the handle is null: the named parameter was not provided
}
if val.is_sql_null() {
    // the parameter was provided as SQL NULL
}
}
}

Escape hatch

If you need the raw handle for an API quack-rs does not wrap, as_raw() borrows it and into_raw() takes ownership:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
use libduckdb_sys::duckdb_value;
unsafe fn demo(bind_info: BindInfo) {
let val = unsafe { bind_info.get_parameter_value(0) };
let raw: duckdb_value = val.into_raw();  // takes ownership, no auto-destroy
// ... use raw handle ...
// caller must call duckdb_destroy_value manually
}
}

Building values

Value::bigint, Value::varchar, Value::date, Value::interval and the other scalar constructors are infallible, with two exceptions that return Result<Value, ExtensionError>:

  • The temporal constructors that take a raw 64-bit payload — time, time_tz, time_ns (with duckdb-1-5), timestamp, timestamp_tz, timestamp_s, timestamp_ms and timestamp_ns. DuckDB stores any payload unchecked, and rendering or casting an out-of-range one aborts, crashes or prints garbage, so quack-rs accepts exactly the range DuckDB's SQL produces (TIME 00:00:00–24:00:00, the TIMESTAMP span 290309-12-22 (BC)–294247-01-10 plus ±infinity, and so on) and returns an error otherwise.
  • Value::decimal(width, scale, unscaled), which checks that width is 1..=38, scale <= width and unscaled has at most width digits.
#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
// Values are DuckDB objects: fill the dispatch table first.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
let noon = Value::time(12 * 3_600 * 1_000_000)?;
assert!(Value::time(-1).is_err());
assert_eq!(noon.as_time(), Some(12 * 3_600 * 1_000_000));
Ok::<(), ExtensionError>(())
}

To build a nested value, Value::list_value and Value::array_value take the element type and the items, Value::struct_value the STRUCT type and one value per field, and Value::enum_value the ENUM type and an index; Value::map and Value::union_value need duckdb-1-5. All return Result<Value, ExtensionError>.

Reading nested values

A parameter of type LIST, STRUCT or MAP is read element by element; each accessor that returns a Value returns an owned one.

MethodReturns
list_len()Number of elements of a LIST (0 for any other type)
list_child(i) / list_items()Element i (Option<Value>) / all elements (Vec<Value>)
struct_field_names()Field names of a STRUCT, in order
struct_child(i)Field i of a STRUCT (Option<Value>); fields are positional
map_len(), map_key(i), map_value(i)Number of pairs; key / value of pair i
#![allow(unused)]
fn main() {
use quack_rs::value::Value;
fn demo(options: &Value) -> Option<String> {
let names = options.struct_field_names();
let idx = names.iter().position(|n| n == "compression")?;
options.struct_child(idx)?.as_str().ok()
}
}

DataChunk, which wraps the chunk a table function's scan callback writes its output to, is described in Reading & Writing Vectors.

Complex Types: STRUCT, LIST, MAP, ARRAY

DuckDB stores its nested types — STRUCT, LIST, MAP and ARRAY — as a parent vector with one or more child vectors. This page shows how a quack-rs extension reads and writes them: the four helper types in vector::complex reach the child vectors, and ListBuilder writes LIST and MAP output without manual offset arithmetic.

Overview

DuckDB typeStoragequack-rs helper
STRUCT{a T, b U, …}Parent vector + N child vectors (one per field)StructVector
LIST<T>Parent vector holds {offset, length} per row; flat child vector holds elementsListVector
MAP<K, V>Stored as LIST<STRUCT{key K, value V}>MapVector
ARRAY<T>[N]Fixed-size array; single child vectorArrayVector

Reading complex types (input vectors)

STRUCT

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
fn demo(parent_vec: duckdb_vector, row_count: usize) {
use quack_rs::vector::{VectorReader, complex::StructVector};

// Inside a scalar function or aggregate update callback:
// parent_vec comes from duckdb_data_chunk_get_vector(chunk, col_idx)
let x_reader = unsafe { StructVector::field_reader(parent_vec, 0, row_count) };
let y_reader = unsafe { StructVector::field_reader(parent_vec, 1, row_count) };

for row in 0..row_count {
    // Each field has its own validity bitmap.
    if unsafe { x_reader.is_valid(row) && y_reader.is_valid(row) } {
        let x: f64 = unsafe { x_reader.read_f64(row) };
        let y: f64 = unsafe { y_reader.read_f64(row) };
        // process (x, y) …
    }
}
}
}

LIST

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
fn demo(list_vec: duckdb_vector, row_count: usize) {
use quack_rs::vector::{VectorReader, complex::ListVector};

let total_elements = unsafe { ListVector::get_size(list_vec) };
let elem_reader = unsafe { ListVector::child_reader(list_vec, total_elements) };

for row in 0..row_count {
    let entry = unsafe { ListVector::get_entry(list_vec, row) };
    for i in 0..entry.length as usize {
        let elem_idx = entry.offset as usize + i;
        if unsafe { elem_reader.is_valid(elem_idx) } {
            let val: i64 = unsafe { elem_reader.read_i64(elem_idx) };
            // process val …
        }
    }
}
}
}

MAP

MAP is LIST<STRUCT{key, value}>. MapVector::key_reader and value_reader read the two fields of the inner struct:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
fn demo(map_vec: duckdb_vector, row_count: usize) {
use quack_rs::vector::complex::MapVector;

let total = unsafe { MapVector::total_entry_count(map_vec) };
let key_reader   = unsafe { MapVector::key_reader(map_vec, total) };
let value_reader = unsafe { MapVector::value_reader(map_vec, total) };

for row in 0..row_count {
    let entry = unsafe { MapVector::get_entry(map_vec, row) };
    for i in 0..entry.length as usize {
        let idx = entry.offset as usize + i;
        let k = unsafe { key_reader.read_str(idx) };   // MAP keys are never NULL
        if unsafe { value_reader.is_valid(idx) } {
            let v: i64 = unsafe { value_reader.read_i64(idx) };
            // process (k, v) …
        }
    }
}
}
}

Writing complex types (output vectors)

STRUCT

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
fn demo(out_vec: duckdb_vector, batch_size: usize, x_values: &[f64], y_values: &[f64]) {
use quack_rs::vector::{VectorWriter, complex::StructVector};

let mut x_writer = unsafe { StructVector::field_writer(out_vec, 0) };
let mut y_writer = unsafe { StructVector::field_writer(out_vec, 1) };

for row in 0..batch_size {
    unsafe { x_writer.write_f64(row, x_values[row]) };
    unsafe { y_writer.write_f64(row, y_values[row]) };
}
}
}

Nested complex types inside STRUCT (v0.11.0+)

When a STRUCT field is itself a LIST, MAP or ARRAY, child_vector(field_idx) on StructWriter or StructReader returns the field's raw vector handle, which the ListVector, MapVector and ArrayVector helpers and ListBuilder accept:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
fn demo(struct_vec: duckdb_vector, row: usize) {
use quack_rs::vector::{ListBuilder, StructWriter};

// STRUCT(name VARCHAR, services VARCHAR[], message VARCHAR)
let mut sw = unsafe { StructWriter::new(struct_vec, 3) };

// Write scalar fields normally
unsafe { sw.write_varchar(row, 0, "hello") };
unsafe { sw.write_varchar(row, 2, "ok") };

// The LIST field at index 1: ListBuilder appends after any elements
// earlier rows already wrote to the child vector.
let services = ["a", "b", "c"];
let mut builder = unsafe { ListBuilder::new(sw.child_vector(1)) };
unsafe {
    builder.push_row(row, services.len(), |writer, base| {
        for (i, s) in services.iter().enumerate() {
            writer.write_varchar(base + i, s);
        }
    });
    builder.finish();
}
}
}

ListBuilder tracks the running offset, writes each parent row's {offset, length} entry, and — importantly — re-fetches the child writer after every reserve:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
fn demo(list_vec: duckdb_vector, rows: &[Vec<i64>]) {
use quack_rs::vector::ListBuilder;

let mut builder = unsafe { ListBuilder::new(list_vec) };
for (row, elements) in rows.iter().enumerate() {
    unsafe {
        builder.push_row(row, elements.len(), |writer, base| {
            for (i, &val) in elements.iter().enumerate() {
                writer.write_i64(base + i, val);
            }
        });
    }
}
unsafe { builder.finish() };
}
}

Why the re-fetch matters. duckdb_list_vector_reserve takes a total capacity, and when it grows it reallocates the child vector's data buffer. A VectorWriter obtained before that call is left holding a dangling pointer. The manual pattern below is safe only because it reserves exactly once, before any writer exists — which requires knowing the total element count up front. ListBuilder has no such requirement. The same applies to writers on anything below the child, such as the fields of a LIST of STRUCTs: fetch them again after every reserve that grows the list (see Pitfall L18).

push_map_row does the same for MAP, handing the closure a writer for the key child and one for the value child.

DuckDB limits a child vector to 2^37 bytes per buffer (MAX_LIST_CHILD_CAPACITY), and a reservation above that — or one the allocator cannot satisfy — throws a C++ exception through the C API, which aborts the process. vector::max_child_capacity(vec) turns the byte limit into an element count for the child's type: 2^34 BIGINTs, 2^33 VARCHARs. ListBuilder applies it by itself. When row lengths come from untrusted input, also set a limit that fits in memory with with_element_limit(n): a row that would exceed the limit is written as NULL instead, and overflowed() reports that it happened.

LIST — manual

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
fn demo(list_vec: duckdb_vector, rows: &[Vec<i64>]) {
use quack_rs::vector::{VectorWriter, complex::ListVector};

let total_elements: usize = rows.iter().map(|r| r.len()).sum();
// Must not exceed quack_rs::vector::max_child_capacity(list_vec); see above.
unsafe { ListVector::reserve(list_vec, total_elements) };

let mut child_writer = unsafe { ListVector::child_writer(list_vec) };
let mut offset = 0usize;
for (row, elements) in rows.iter().enumerate() {
    for (i, &val) in elements.iter().enumerate() {
        unsafe { child_writer.write_i64(offset + i, val) };
    }
    unsafe { ListVector::set_entry(list_vec, row, offset as u64, elements.len() as u64) };
    offset += elements.len();
}
unsafe { ListVector::set_size(list_vec, total_elements) };
}
}

MAP — manual

Writing a MAP follows the LIST pattern, but keys and values go into the two fields of the inner STRUCT vector. Prefer ListBuilder::push_map_row unless you know the total pair count before writing:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
fn demo(map_vec: duckdb_vector, total_pairs: usize, all_pairs: &[Vec<(String, i64)>]) {
use quack_rs::vector::complex::MapVector;

unsafe { MapVector::reserve(map_vec, total_pairs) };

let mut key_writer = unsafe { MapVector::key_writer(map_vec) };
let mut val_writer = unsafe { MapVector::value_writer(map_vec) };
let mut offset = 0usize;
for (row, pairs) in all_pairs.iter().enumerate() {
    for (i, (k, v)) in pairs.iter().enumerate() {
        unsafe { key_writer.write_varchar(offset + i, k) };
        unsafe { val_writer.write_i64(offset + i, *v) };
    }
    unsafe { MapVector::set_entry(map_vec, row, offset as u64, pairs.len() as u64) };
    offset += pairs.len();
}
unsafe { MapVector::set_size(map_vec, total_pairs) };
}
}

Constructing complex logical types

Use LogicalType constructors to define complex column types. Each constructor has a variant that accepts TypeId values (for simple element types) and a _from_logical variant (for nested complex types):

Constructor_from_logical variantCreates
LogicalType::list(TypeId)list_from_logical(&LogicalType)LIST<T>
LogicalType::map(TypeId, TypeId)map_from_logical(&LogicalType, &LogicalType)MAP<K, V>
LogicalType::struct_type(&[(&str, TypeId)])struct_type_from_logical(&[(&str, LogicalType)])STRUCT{...}
LogicalType::union_type(&[(&str, TypeId)])union_type_from_logical(&[(&str, LogicalType)])UNION(...)
LogicalType::array(TypeId, u64)array_from_logical(&LogicalType, u64)ARRAY<T>[N]
LogicalType::enum_type(&[&str])—ENUM(...)
LogicalType::decimal(u8, u8)—DECIMAL(w, s)

Each constructor also has a try_ form (try_list, try_struct_type_from_logical, …) that returns Result<LogicalType, LogicalTypeError>; the plain forms panic where the try_ form returns an error. Errors include a composite TypeId passed where a _from_logical variant is needed, STRUCT field or UNION member names that are equal ignoring ASCII case, and a UNION with more than MAX_UNION_MEMBERS (255) members.

API reference

All helpers are in quack_rs::vector::complex (re-exported from quack_rs::prelude).

StructVector

MethodDescription
get_child(vec, field_idx)Returns the raw child vector for field field_idx
field_reader(vec, field_idx, row_count)Creates a VectorReader for a STRUCT field
field_writer(vec, field_idx)Creates a VectorWriter for a STRUCT field

StructWriter / StructReader complex field access (v0.11.0+)

MethodDescription
StructWriter::child_vector(field_idx)Returns the raw duckdb_vector of a nested field (LIST, MAP, ARRAY)
StructWriter::child_list_vector(field_idx)Alias of child_vector for a LIST field
StructReader::child_vector(field_idx)Same as StructWriter::child_vector, for reading (unsafe)

ListVector

MethodDescription
get_child(vec)Returns the flat element child vector
get_size(vec)Total number of elements across all rows
set_size(vec, n)Sets the number of elements after writing
reserve(vec, capacity)Reserves capacity in the child vector (at most max_child_capacity(vec))
get_entry(vec, row)Returns {offset, length} for a row (reading)
set_entry(vec, row, offset, length)Sets {offset, length} for a row (writing)
child_reader(vec, count)Creates a VectorReader for the element vector
child_writer(vec)Creates a VectorWriter for the element vector

MapVector

MethodDescription
struct_child(vec)Returns the inner STRUCT vector
keys(vec)Returns the key vector (STRUCT field 0)
values(vec)Returns the value vector (STRUCT field 1)
total_entry_count(vec)Total key-value pairs
reserve(vec, n)Reserves capacity for n pairs (at most max_child_capacity(vec))
set_size(vec, n)Sets total entry count after writing
get_entry(vec, row)Returns {offset, length} for a row (reading)
set_entry(vec, row, offset, length)Sets {offset, length} for a row (writing)
key_reader(vec, count) / value_reader(vec, count)Creates a VectorReader for the keys / values
key_writer(vec) / value_writer(vec)Creates a VectorWriter for the keys / values

ArrayVector

MethodDescription
get_child(vec)Returns the child vector of a fixed-size ARRAY vector

NULL Handling & Strings

This page covers checking for NULL before reading a DuckDB vector, writing NULL output, and reading and writing VARCHAR and BLOB values. The two topics belong together: reading a string from a NULL row is undefined behaviour, not just a wrong value.


NULL checks

Every row in a DuckDB vector may be NULL. Always check validity before reading:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, writer: &mut VectorWriter) {
for row in 0..reader.row_count() {
    if unsafe { !reader.is_valid(row) } {
        // Propagate NULL to output
        unsafe { writer.set_null(row) };
        continue;
    }
    // Safe to read
    let value = unsafe { reader.read_str(row) };
}
}
}

Reading a fixed-width value from a NULL row returns garbage. The vector's data buffer is not zeroed at NULL positions, and no error is raised: you get whatever bytes the buffer holds at that position.

For VARCHAR and BLOB it is worse than garbage. A NULL row's 16-byte entry is left as it was, and DuckDB reuses vector buffers between chunks, so the entry can still be the pointer-format record of a string from an earlier chunk whose memory may since have been freed or reused. read_str / read_blob on such a row follow that pointer: undefined behaviour, not merely a wrong answer. Their # Safety sections require the row to be valid.

Writing NULL

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(writer: &mut VectorWriter, row: usize) {
unsafe { writer.set_null(row) };
}
}

Pitfall L4: VectorWriter::set_null calls duckdb_vector_ensure_validity_writable before accessing the validity bitmap. Without it, a vector that has no mask yet makes duckdb_vector_get_validity return NULL, and the NULL you write is silently dropped. Never write NULL manually; always use set_null — which, for a STRUCT or ARRAY output, also nulls the fields / elements of that row the way DuckDB expects. See Pitfall L4.

Clearing NULL (v0.11.0+)

To mark a row as valid after a previous set_null:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(writer: &mut VectorWriter, row: usize) {
unsafe { writer.set_valid(row) };
}
}

VARCHAR reading

Read VARCHAR columns with VectorReader::read_str:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, row: usize) {
let s: &str = unsafe { reader.read_str(row) };
}
}

The returned &str borrows from the DuckDB vector — it must not outlive the callback. Do not store it in a struct; clone it to a String if you need to keep it.

The duckdb_string_t format

Pitfall P7: the Rust bindings do not document the layout of duckdb_string_t. quack-rs decodes it for you; the details below are for reference. See Pitfall P7.

DuckDB stores VARCHAR values in a 16-byte duckdb_string_t struct with two representations, selected at runtime based on string length:

FormatConditionLayout
Inlinelength ≤ 12[len: u32][data: [u8; 12]]
Pointerlength > 12[len: u32][prefix: [u8; 4]][ptr: *const u8]

On a 32-bit target (DuckDB-WASM) the pointer is 4 bytes and the last 4 bytes are unused. The length and the pointer are in the target's byte order.

VectorReader::read_str and the underlying read_duck_string function handle both formats, so you do not need to inspect the raw struct. A value that is not valid UTF-8 is returned as ""; use read_blob to get its bytes.

Empty strings vs NULL

An empty string ("") and NULL are distinct values:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, row: usize) {
// NULL: is_valid returns false
// Empty string: is_valid returns true, read_str returns ""
if unsafe { !reader.is_valid(row) } {
    // This is NULL
} else {
    let s = unsafe { reader.read_str(row) };
    if s.is_empty() {
        // This is an empty string, not NULL
    }
}
}
}

Writing VARCHAR

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(writer: &mut VectorWriter, row: usize, my_str: &str) {
unsafe { writer.write_varchar(row, my_str) };  // &str
}
}

write_varchar copies the string bytes into DuckDB's managed storage, so the &str need not outlive the call. write_blob does the same for &[u8].

Both panic for a value longer than MAX_STRING_LEN (u32::MAX bytes, the most a duckdb_string_t length field can hold) instead of storing a truncated value; inside scalar_callback! and the typed scalar constructors the panic becomes a SQL error. try_write_varchar and try_write_blob return Result<(), ExtensionError> instead and write nothing on error:

#![allow(unused)]
fn main() {
use quack_rs::vector::VectorWriter;
use quack_rs::error::ExtensionError;
fn demo(writer: &mut VectorWriter, row: usize, my_str: &str) -> Result<(), ExtensionError> {
unsafe { writer.try_write_varchar(row, my_str)? };
Ok(())
}
}

Reading BLOB values

BLOB uses the same inline/pointer layout as VARCHAR, but may contain any bytes. Use read_blob so the data is not interpreted as UTF-8:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, row: usize) {
let bytes: &[u8] = unsafe { reader.read_blob(row) };
}
}

Like read_str, the returned slice borrows from the DuckDB vector and must not outlive the callback.


Complete NULL-safe VARCHAR pattern

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::vector::{VectorReader, VectorWriter};
unsafe extern "C" fn my_scalar(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };

    for row in 0..reader.row_count() {
        if unsafe { !reader.is_valid(row) } {
            unsafe { writer.set_null(row) };
            continue;
        }
        let s = unsafe { reader.read_str(row) };
        let upper = s.to_uppercase();
        unsafe { writer.write_varchar(row, &upper) };
    }
}
}

DuckStringView

For advanced use cases where you need access to the raw string bytes or the inline/pointer distinction, quack_rs::vector::string::DuckStringView is available:

#![allow(unused)]
fn main() {
fn demo(data: *const u8, idx: usize) {
use quack_rs::vector::string::{DuckStringView, DUCK_STRING_SIZE};

// From raw 16-byte data (inside a vector callback). `from_raw` is unsafe because it
// follows the pointer of a string longer than 12 bytes; untrusted bytes go through
// `DuckStringView::inline_from_bytes`, which refuses that format instead.
let raw: &[u8; 16] = unsafe { &*data.add(idx * DUCK_STRING_SIZE).cast() };
let view = unsafe { DuckStringView::from_raw(raw) };

println!("length: {}", view.len());
println!("is_empty: {}", view.is_empty());
if let Some(s) = view.as_str() {
    println!("content: {s}");
}
}
}

In practice, prefer reader.read_str(row). DuckStringView is needed only when you have a raw data pointer rather than a VectorReader. Unlike read_str, its as_str returns None, not "", for a value that is not valid UTF-8.


Constants

ConstantValueMeaning
DUCK_STRING_SIZE16Size of one duckdb_string_t in bytes
DUCK_STRING_INLINE_MAX_LEN12Longest value stored inline (no heap pointer), in bytes
MAX_STRING_LENu32::MAX (4,294,967,295)Longest VARCHAR or BLOB value DuckDB can store, in bytes

All three are in quack_rs::vector::string.

INTERVAL Type

DuckDB's INTERVAL type represents a duration with three independent components: months, days and microseconds. The quack_rs::interval module provides the DuckInterval struct, which matches DuckDB's in-memory layout, and overflow-safe conversions to microseconds.


Why a custom struct?

Pitfall P8: the Rust bindings do not document the INTERVAL layout or how DuckDB converts intervals. DuckInterval and the functions below encode both. See Pitfall P8.

DuckDB's C duckdb_interval struct is 16 bytes with this exact layout:

offset 0:  months (i32)  — calendar months
offset 4:  days   (i32)  — calendar days
offset 8:  micros (i64)  — microseconds (not limited to one day)
total:     16 bytes

DuckInterval is #[repr(C)] with the same field order, and a compile-time assertion checks that it is exactly 16 bytes.


Reading INTERVAL values

#![allow(unused)]
fn main() {
use quack_rs::interval::DuckInterval;
use quack_rs::vector::VectorReader;
fn demo(reader: &VectorReader, row: usize) {
let iv: DuckInterval = unsafe { reader.read_interval(row) };
println!("{} months, {} days, {} µs", iv.months, iv.days, iv.micros);
}
}

VectorReader::read_interval handles the raw pointer arithmetic and alignment using read_interval_at internally.


DuckInterval fields

#![allow(unused)]
fn main() {
use quack_rs::interval::DuckInterval;

let iv = DuckInterval {
    months: 1,    // 1 calendar month
    days: 15,     // 15 calendar days
    micros: 3_600_000_000, // 1 hour in microseconds
};
}

The fields are public, so a DuckInterval can be built directly.

The derived PartialEq, Eq and Hash compare the three fields, so { months: 1, .. } and { days: 30, .. } are different values here, while in SQL INTERVAL '1 month' = INTERVAL '30 days' is true.

Zero interval

#![allow(unused)]
fn main() {
use quack_rs::interval::DuckInterval;
let zero = DuckInterval::zero();    // { months: 0, days: 0, micros: 0 }
let zero = DuckInterval::default(); // same
assert_eq!(zero, DuckInterval::zero());
}

Converting to microseconds

Months and days have no fixed length in wall-clock time, so an interval has no single exact length. When you need one number, for ordering or bucketing, convert to microseconds with the approximation DuckDB uses when it compares intervals and in epoch_us(interval): 1 month = 30 days.

This is not date arithmetic: DuckDB adds an interval to a date by calendar months (DATE '2024-01-31' + INTERVAL 1 MONTH is 2024-02-29). It also matches SQL comparison only when the three fields share a sign: INTERVAL '1 month' - INTERVAL '1 day' converts to the same total as INTERVAL '29 days' but compares greater in SQL.

Checked conversion (returns Option)

#![allow(unused)]
fn main() {
use quack_rs::interval::DuckInterval;
use quack_rs::interval::interval_to_micros;

let iv = DuckInterval { months: 0, days: 1, micros: 500_000 };
match interval_to_micros(iv) {
    Some(us) => println!("{us} microseconds"),
    None => println!("overflow"),
}

// Method form:
let us: Option<i64> = iv.to_micros();
assert_eq!(us, Some(86_400_000_000 + 500_000));
}

Returns None if the total does not fit in an i64, which takes extreme values such as months: i32::MAX, days: i32::MAX, micros: i64::MAX. The sum is computed exactly, so large fields of opposite signs that cancel out still convert.

Saturating conversion (returns i64)

#![allow(unused)]
fn main() {
use quack_rs::interval::DuckInterval;
use quack_rs::interval::interval_to_micros_saturating;

let iv = DuckInterval { months: i32::MAX, days: i32::MAX, micros: i64::MAX };
let us: i64 = interval_to_micros_saturating(iv); // i64::MAX

// Method form:
let us: i64 = iv.to_micros_saturating();
assert_eq!(us, i64::MAX);
assert_eq!(interval_to_micros_saturating(iv), i64::MAX);
}

The saturating form clamps an out-of-range total to i64::MAX or i64::MIN. Neither form panics; use the checked form when an overflow must be reported rather than clamped.


Conversion constants

ConstantValueMeaning
MICROS_PER_DAY86_400_000_000Microseconds in 24 hours
MICROS_PER_MONTH2_592_000_000_000Microseconds in 30 days
#![allow(unused)]
fn main() {
use quack_rs::interval::{MICROS_PER_DAY, MICROS_PER_MONTH};

assert_eq!(MICROS_PER_DAY, 86_400 * 1_000_000);
assert_eq!(MICROS_PER_MONTH, 30 * MICROS_PER_DAY);
}

Low-level: read_interval_at

If you have a raw data pointer (e.g., from duckdb_vector_get_data), you can read an interval directly:

#![allow(unused)]
fn main() {
fn demo(data_ptr: *const u8, row_idx: usize) {
use quack_rs::interval::read_interval_at;

// SAFETY: data is a valid DuckDB INTERVAL vector data pointer, idx is in bounds.
let iv = unsafe { read_interval_at(data_ptr, row_idx) };
}
}

In practice, use VectorReader::read_interval(row), which computes the data pointer for you; its remaining unsafe contract is the row index and the column type.


Complete example: aggregate over INTERVAL

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_data_chunk, duckdb_function_info};
use quack_rs::aggregate::{AggregateState, FfiState};
use quack_rs::vector::VectorReader;
#[derive(Default)]
struct TotalDurationState {
    total_micros: i64,
}
impl AggregateState for TotalDurationState {}

unsafe extern "C" fn update(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    states: *mut duckdb_aggregate_state,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    for row in 0..reader.row_count() {
        if unsafe { !reader.is_valid(row) } { continue; }
        let iv = unsafe { reader.read_interval(row) };
        let us = iv.to_micros_saturating();
        let state_ptr = unsafe { *states.add(row) };
        if let Some(st) = unsafe { FfiState::<TotalDurationState>::with_state_mut(state_ptr) } {
            st.total_micros = st.total_micros.saturating_add(us);
        }
    }
}
}

Memory layout verification

A compile-time assertion checks that DuckInterval is 16 bytes with at least 4-byte alignment, the layout of DuckDB's duckdb_interval. If it fails, the crate does not compile, so a layout change is caught at build time rather than at run time.

Dates, Times and Timestamps

VectorReader and VectorWriter read and write DuckDB's DATE, TIME, TIMETZ, TIMESTAMP and INTERVAL values as the raw integers DuckDB stores; quack_rs::datetime converts them to and from calendar fields. This page also covers DECIMAL and HUGEINT, whose conversions live in the same module.

SQL typeStorageAccessor
DATEi32 — days since 1970-01-01read_date / write_date
TIMEi64 — microseconds since midnightread_time / write_time
TIMETZpacked u64read_time_tz / write_time_tz
TIMESTAMPi64 — microseconds since the epochread_timestamp / write_timestamp
TIMESTAMPTZi64 — microseconds since the epoch, UTCread_timestamp_tz / write_timestamp_tz
TIMESTAMP_Si64 — seconds since the epochread_timestamp_s / write_timestamp_s
TIMESTAMP_MSi64 — milliseconds since the epochread_timestamp_ms / write_timestamp_ms
TIMESTAMP_NSi64 — nanoseconds since the epochread_timestamp_ns / write_timestamp_ns
INTERVAL{ months: i32, days: i32, micros: i64 }read_interval / write_interval (see INTERVAL Type)

Turning those integers into year/month/day means implementing the proleptic Gregorian calendar, and getting it to agree with DuckDB's SQL semantics exactly rather than approximately. DuckDB already exposes the conversions, and they are in the stable prefix of the C API, so quack_rs::datetime wraps them instead of reimplementing them. They need no feature flag.

Decomposing and composing

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, writer: &mut VectorWriter, row: usize) {
use quack_rs::datetime;

// DATE -> calendar date
let days = unsafe { reader.read_date(row) };
let date = unsafe { datetime::date_from_days(days) };
println!("{:04}-{:02}-{:02}", date.year, date.month, date.day);

// …and back. `None` means DuckDB cannot represent the date.
match unsafe { datetime::date_to_days(date) } {
    Some(days) => unsafe { writer.write_date(row, days) },
    None => unsafe { writer.set_null(row) },
}
}
}

Time, TimeTz and Timestamp work the same way:

#![allow(unused)]
fn main() {
use quack_rs::datetime;
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, writer: &mut VectorWriter, rows: usize) {
for row in 0..rows {
let Some(ts) = (unsafe { datetime::timestamp_from_micros(reader.read_timestamp(row)) }) else {
    // ±infinity (or the first ~4 hours of the i64 range): no calendar form.
    unsafe { writer.set_null(row) };
    continue;
};
assert!((0..1_000_000).contains(&ts.time.micros));   // ts.date and ts.time are plain structs

let micros = unsafe { datetime::timestamp_to_micros(ts) };   // Option<i64>
}
}
}

Invalid input is None, not an abort

Several of DuckDB's conversions throw a C++ exception on bad input, and the C API does not catch it — so calling them directly with, say, month 13 aborts the whole process ("Rust cannot catch foreign exceptions"). The wrappers apply DuckDB's own conditions first and return None instead:

FunctionReturns None when
date_to_daysmonth not 1–12, day not in that month (leap years included), or the date is outside 5877642-06-25 BC – 5881580-07-10; datetime::is_valid_date is the same check
timestamp_from_microsthe value is ±infinity, or below -106_751_991 * MICROS_PER_DAY (which includes i64::MIN)
timestamp_to_microsthe date is invalid, the result overflows i64, or it lands on ±infinity
time_from_microsthe value is outside 0..=MICROS_PER_DAY (00:00:00–24:00:00)
time_tz_bitsthe time is outside 0..=MICROS_PER_DAY, or the offset beyond ±15:59:59 (TIME_TZ_MAX_OFFSET_SECONDS)
time_tz_from_bitsthe bits decode to a time or offset that time_tz_bits would refuse
decimal_to_f64width > 38 or scale > width

time_from_micros and time_tz_from_bits guard an assertion rather than an exception: a release build of DuckDB decomposes an out-of-range time into out-of-range fields, and a build with assertions enabled aborts.

time_to_micros does no range check, exactly like DuckDB: an hour of 25 simply gives a TIME past midnight.

TIMETZ is a packed 64-bit value, not a plain integer — build and read it through the helpers rather than by hand:

#![allow(unused)]
fn main() {
use quack_rs::datetime;
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, writer: &mut VectorWriter, row: usize) {
let bits = unsafe { datetime::time_tz_bits(12 * 3_600 * 1_000_000, -5 * 3_600) }
    .expect("noon, UTC-5, is in range");
unsafe { writer.write_time_tz(row, bits) };

let decoded = unsafe { datetime::time_tz_from_bits(reader.read_time_tz(row)) }
    .expect("DuckDB wrote a valid TIMETZ");
assert_eq!(decoded.offset_seconds, -5 * 3_600);
}
}

Infinity

DuckDB reserves two values of DATE and of TIMESTAMP for infinity and -infinity. Decomposing one into a calendar date is meaningless, so check first:

#![allow(unused)]
fn main() {
use quack_rs::datetime;
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, writer: &mut VectorWriter, row: usize) {
let days = unsafe { reader.read_date(row) };
if unsafe { datetime::is_finite_date(days) } {
    let date = unsafe { datetime::date_from_days(days) };
    // …
}
}
}

Note the exact values, which are easy to get wrong:

ConstantValue
DATE_INFINITY_DAYSi32::MAX
DATE_NEGATIVE_INFINITY_DAYS-i32::MAX
TIMESTAMP_INFINITY_MICROSi64::MAX
TIMESTAMP_NEGATIVE_INFINITY_MICROS-i64::MAX

Negative infinity is -i32::MAX, not i32::MIN. i32::MIN is an ordinary (if absurd) finite date, and treating it as infinity would silently drop real rows.

DECIMAL

DECIMAL is stored in the narrowest integer that fits its declared width, so the width has to travel with the value:

Declared widthPhysical storage
1 – 4i16
5 – 9i32
10 – 18i64
19 – 38i128

read_decimal / write_decimal take the width and pick the right one. Get it from the column's LogicalType:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_vector;
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(vec: duckdb_vector, reader: &VectorReader, writer: &mut VectorWriter, row: usize) {
let logical = unsafe { quack_rs::vector::vector_get_column_type(vec) };
let width = unsafe { logical.decimal_width() };
let scale = unsafe { logical.decimal_scale() };

let unscaled = unsafe { reader.read_decimal(row, width) };
// The represented number is unscaled / 10^scale. The doubled value must
// still fit in `width` digits; write_decimal does not check.
unsafe { writer.write_decimal(row, width, unscaled * 2) };
}
}

datetime::f64_to_decimal and datetime::decimal_to_f64 convert through DuckDB's own routines when a floating-point view is what you want. decimal_to_f64 returns None for a width above 38 or a scale above the width: DuckDB would index its powers-of-ten table out of bounds.

Wide integers

HUGEINT is { lower: u64, upper: i64 } and UHUGEINT is two u64s. read_i128 / write_i128 and read_u128 / write_u128 handle the halves; datetime::hugeint_to_f64, f64_to_hugeint, uhugeint_to_f64 and f64_to_uhugeint convert through DuckDB's own routines, so they round as DuckDB does.

Running SQL from an Extension

Extensions often need to run SQL against the database that is loading them: checking whether a table exists before registering a replacement scan, creating a helper view, reading a setting, or looking up a credential through duckdb_secrets(). This page covers queries, prepared statements with bound parameters, keeping a connection for use after loading, and cancelling a running query.

The C API has everything for that — duckdb_query, duckdb_prepare, duckdb_bind_*, duckdb_fetch_chunk — and all of it is in the stable prefix, so it needs no feature flag. What it does not have is any help releasing the handles: every one of them has a matching destroy that must run exactly once, including on the error paths, which is where hand-written FFI usually leaks.

quack_rs::query wraps them:

TypeOwnsReleased by
QueryResultduckdb_resultduckdb_destroy_result
OwnedDataChunkduckdb_data_chunkduckdb_destroy_data_chunk
PreparedStatementduckdb_prepared_statementduckdb_destroy_prepare
OwnedConnectionduckdb_connectionduckdb_disconnect

During registration

Connection (from entry_point_v2!) can run SQL directly:

#![allow(unused)]
fn main() {
use quack_rs::connection::Connection;
use quack_rs::error::ExtensionError;

fn register(con: &Connection) -> Result<(), ExtensionError> {
    // Create a helper view the extension's functions rely on.
    unsafe { con.execute("CREATE OR REPLACE VIEW my_ext_config AS SELECT 1 AS version") }?;

    // Read something back. The cast makes the column VARCHAR, which read_str requires.
    let mut result = unsafe { con.query("SELECT current_setting('threads')::VARCHAR") }?;
    if let Some(chunk) = result.next_chunk()? {
        // The reader must outlive the `&str` it hands out, so bind it first.
        let reader = unsafe { chunk.reader(0) };
        let threads = unsafe { reader.read_str(0) };
        eprintln!("DuckDB is using {threads} threads");
    }
    Ok(())
}
}

Results arrive a chunk at a time — at most duckdb_vector_size() rows each — so call next_chunk until it returns Ok(None):

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
use quack_rs::query::OwnedConnection;
// `Connection` (from entry_point_v2!) only exists during an extension load. An
// `OwnedConnection` has the same query / execute / prepare methods.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
let mut db = std::ptr::null_mut();
unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); }
let con = unsafe { OwnedConnection::open(db) }.unwrap();
let run = || -> Result<(), ExtensionError> {
let mut result = unsafe { con.query("SELECT i FROM range(10000) t(i)") }?;
let mut total: i64 = 0;
while let Some(chunk) = result.next_chunk()? {
    let reader = unsafe { chunk.reader(0) };
    for row in 0..chunk.size() {
        total += unsafe { reader.read_i64(row) };
    }
}
assert_eq!(total, 49_995_000);
Ok(()) };
run().unwrap();
}

next_chunk returns Result<Option<OwnedDataChunk>, ExtensionError> because a result can stop early. A streaming result (PreparedStatement::execute_streaming, with the duckdb-1-5 feature) produces rows as it runs, so a runtime error part-way through, an interrupt, or another statement run on the same connection (which invalidates the stream) surfaces at next_chunk. The C API reports that the same way as the end of the rows (a null chunk); quack-rs reads the error DuckDB recorded and returns it, so a partial result cannot pass for a complete one. The ? above is what keeps it from being silently truncated. After Ok(None) or an error, later calls return the same thing.

Inspecting a result

MethodReturns
column_count()Number of columns
column_name(i)Name of column i (Option<String>)
column_type(i)Top-level TypeId of column i
column_logical_type(i)Full LogicalType of column i, keeping STRUCT fields, LIST element type, DECIMAL width and scale
result_kind()ResultKind::Rows, ChangedRows, Nothing or Invalid
rows_changed()Rows changed by an INSERT / UPDATE / DELETE; 0 for other statements
is_streaming()Whether the result is streaming (duckdb-1-5)

Several statements in one string

query and execute accept several ;-separated statements, and DuckDB runs every one, in order. The result you get back is the first statement that produces rows — or, when none does, the last statement's; the results of later row-producing statements are discarded. So "SELECT 1; INSERT …" runs the INSERT but execute reports 0 rows changed. The first failing statement fails the call, after the ones before it have run (and, outside an explicit transaction, committed). An empty string, or just ;, succeeds with an empty result. prepare takes exactly one statement.

Bind values, do not interpolate them

Anything that did not come from your own source text — a table name from a function argument, a path from a config option — goes through a parameter. Parameters are 1-indexed, matching the C API.

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
use quack_rs::query::OwnedConnection;
// `Connection` (from entry_point_v2!) only exists during an extension load. An
// `OwnedConnection` has the same query / execute / prepare methods.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
let mut db = std::ptr::null_mut();
unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); }
let con = unsafe { OwnedConnection::open(db) }.unwrap();
con.execute("CREATE TABLE audit (name VARCHAR, n BIGINT)").unwrap();
let user_supplied_name = "O'Brien'); DROP TABLE audit; --";
let run = || -> Result<(), ExtensionError> {
let stmt = unsafe { con.prepare("INSERT INTO audit VALUES (?, ?)") }?;
stmt.bind_str(1, user_supplied_name)?;   // safe even if it contains quotes
stmt.bind_i64(2, 42)?;
stmt.execute()?;
let mut r = con.query("SELECT count(*) FROM audit WHERE name LIKE 'O''Brien%'")?;
assert_eq!(unsafe { r.next_chunk()?.unwrap().reader(0).read_i64(0) }, 1);
Ok(()) };
run().unwrap();
}

bind_str passes the length explicitly, so embedded NUL bytes are preserved and no CString conversion can fail.

There is a typed bind for every integer width (bind_i8 … bind_u128), bind_f32 / bind_f64, bind_bool, bind_blob, bind_null, bind_decimal, bind_date, bind_time, bind_timestamp, bind_timestamp_tz and bind_interval. bind_value takes any Value, which covers the composite types. Like the Value constructors, bind_decimal validates width, scale and digit count, and bind_time and the timestamp binds refuse a payload outside the range DuckDB's SQL produces; the C API's duckdb_bind_* functions check nothing.

Named parameters resolve by name:

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
use quack_rs::query::OwnedConnection;
// `Connection` (from entry_point_v2!) only exists during an extension load. An
// `OwnedConnection` has the same query / execute / prepare methods.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
let mut db = std::ptr::null_mut();
unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); }
let con = unsafe { OwnedConnection::open(db) }.unwrap();
con.execute("CREATE TABLE t AS SELECT range AS id FROM range(10)").unwrap();
let id = 7;
let run = || -> Result<(), ExtensionError> {
let stmt = unsafe { con.prepare("SELECT * FROM t WHERE id = $needle") }?;
let index = stmt.parameter_index("needle").expect("named parameter");
stmt.bind_i64(index, id)?;
let mut r = stmt.execute()?;
assert_eq!(unsafe { r.next_chunk()?.unwrap().reader(0).read_i64(0) }, 7);
Ok(()) };
run().unwrap();
}

Reuse a statement by clearing its bindings between executions:

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
use quack_rs::query::OwnedConnection;
// `Connection` (from entry_point_v2!) only exists during an extension load. An
// `OwnedConnection` has the same query / execute / prepare methods.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
let mut db = std::ptr::null_mut();
unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); }
let con = unsafe { OwnedConnection::open(db) }.unwrap();
let run = || -> Result<(), ExtensionError> {
let stmt = con.prepare("SELECT ?::BIGINT * 2")?;
let ids = [1_i64, 2, 3];
for id in ids {
    stmt.clear_bindings()?;
    stmt.bind_i64(1, id)?;
    let mut result = stmt.execute()?;
    // …
}
Ok(()) };
run().unwrap();
}

After registration

The connection DuckDB passes to your entry point is borrowed — the entry point disconnects it when your closure returns. Inside a scalar, table or aggregate callback you have no connection at all: the C API gives you a duckdb_client_context, and there is no duckdb_client_context_get_connection.

If a callback or a background thread needs to run SQL, open your own connection during registration and keep it. A duckdb_connection holds its own reference to the database instance, so it stays valid after loading finishes:

#![allow(unused)]
fn main() {
use quack_rs::connection::Connection;
use quack_rs::error::ExtensionError;
use quack_rs::query::OwnedConnection;
use std::sync::{Mutex, OnceLock};

// `OwnedConnection` is `Send` but not `Sync` (see below), so a `static` must
// hold it behind a `Mutex`.
static CONN: OnceLock<Mutex<OwnedConnection>> = OnceLock::new();

fn register(con: &Connection) -> Result<(), ExtensionError> {
    let owned = unsafe { con.open_connection() }?;
    let _ = CONN.set(Mutex::new(owned));
    Ok(())
}
}

OwnedConnection is Send but deliberately not Sync: DuckDB permits moving a connection between threads, not using one concurrently. Open one connection per thread, or guard it with a mutex.

Cancelling a query and reading its progress

OwnedConnection::interrupt_handle returns an InterruptHandle, which is Send + Sync and borrows the connection, so it cannot outlive it. Another thread can call its cancel() to stop the running query, which then fails with an interrupt error at DuckDB's next check, or its progress() to read a QueryProgress (percentage, rows_processed, total_rows_to_process). percentage is -1.0 when DuckDB cannot report progress, for example when the progress bar is disabled (SET enable_progress_bar = true). OwnedConnection::interrupt and progress do the same on the calling thread.

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
use quack_rs::query::{OwnedConnection, QueryResult};
use std::sync::mpsc::{channel, RecvTimeoutError};
use std::time::Duration;

fn query_with_timeout(con: &OwnedConnection, sql: &str) -> Result<QueryResult, ExtensionError> {
    let watchdog = con.interrupt_handle();
    let (done, finished) = channel::<()>();
    std::thread::scope(|scope| {
        scope.spawn(move || {
            // Cancel the query if it has not finished within 30 seconds.
            if let Err(RecvTimeoutError::Timeout) = finished.recv_timeout(Duration::from_secs(30)) {
                watchdog.cancel();
            }
        });
        let result = con.query(sql);
        drop(done); // wakes the watchdog
        result
    })
}
}

Errors

Failures carry DuckDB's own message:

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
use quack_rs::query::OwnedConnection;
// `Connection` (from entry_point_v2!) only exists during an extension load. An
// `OwnedConnection` has the same query / execute / prepare methods.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
let mut db = std::ptr::null_mut();
unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); }
let con = unsafe { OwnedConnection::open(db) }.unwrap();
let err = unsafe { con.query("SELECT * FROM no_such_table") }.unwrap_err();
assert!(err.as_str().contains("no_such_table"));
assert_eq!(con.execute("SELECT 1").unwrap(), 0);
}

The connection stays usable afterwards.

Bulk Appender

Appender is an RAII wrapper around DuckDB's appender, the fastest way for an extension to bulk-insert rows into an existing table: rows are buffered and written in batches instead of going through one INSERT statement each. This page covers appending row by row and chunk by chunk, error handling, and what happens when a row fails half-way.

No feature flag required. DuckDB has kept duckdb_appender_* in the frozen stable prefix of the extension API (slots 281–291 and 330–356) since v1.2.0, so using the appender does not push your extension onto the version-pinned unstable ABI. See ABI Compatibility for why that distinction matters. Three methods are the exception and need duckdb-1-5: error_data, clear, and append_default_to_chunk.

Row at a time

Call one append_* per column, then finish the row with end_row. row calls end_row for you when its closure succeeds, so a forgotten end_row cannot leave the table silently short:

#![allow(unused)]
fn main() {
use quack_rs::appender::{AppendError, Appender};
use libduckdb_sys::duckdb_connection;

unsafe fn demo(con: duckdb_connection) -> Result<(), AppendError> {
// SAFETY: `con` is a valid, open connection (e.g. from an entry point).
let appender = unsafe { Appender::new(con, None, c"measurements") }?;

for (sensor, reading) in [("a", 1.5_f64), ("b", 2.5)] {
    appender.row(|row| {
        row.append_str(sensor)?;
        row.append_f64(reading)
    })?;
}

appender.close()?;
Ok(())
}
fn live_connection() -> libduckdb_sys::duckdb_connection {
    std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
    let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess);
    }
    con
}
/// First column of the first row, as BIGINT; `None` for NULL.
fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    let reader = unsafe { chunk.reader(0) };
    unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) }
}
let con = live_connection();
unsafe { quack_rs::query::execute(con, "CREATE TABLE measurements (sensor VARCHAR, reading DOUBLE)") }.unwrap();
unsafe { demo(con) }.unwrap();
assert_eq!(query_i64(con, "SELECT count(*) FROM measurements"), Some(2));
assert_eq!(query_i64(con, "SELECT (sum(reading) * 10)::BIGINT FROM measurements"), Some(40));
}

append_str uses duckdb_append_varchar_length, so interior NUL bytes survive; the NUL-terminated duckdb_append_varchar would stop at the first one. append_str and append_bytes refuse a value longer than u32::MAX bytes, which DuckDB would silently truncate.

append_time and append_timestamp refuse, before calling DuckDB, a payload outside the range DuckDB's SQL produces (TIME 00:00:00–24:00:00; a TIMESTAMP from 290309-12-22 (BC), plus -infinity). DuckDB stores such a value unchecked, and reading the row back later fails or crashes.

A chunk at a time

append_chunk takes a whole DataChunk whose column types match the appender's active columns. It makes fewer FFI calls than appending row by row and suits data that is already in vectors; it is refused while a row appended by hand is still open.

#![allow(unused)]
fn main() {
use quack_rs::appender::{AppendError, Appender};
use quack_rs::data_chunk::DataChunk;
use libduckdb_sys::duckdb_connection;

unsafe fn load(con: duckdb_connection, chunks: &[DataChunk]) -> Result<(), AppendError> {
// SAFETY: `con` is a valid, open connection.
let appender = unsafe { Appender::new(con, None, c"events") }?;
for chunk in chunks {
    appender.append_chunk(chunk)?;
}
appender.close()?;
Ok(())
}
}

Pass a schema (or a fully-qualified catalog + schema) when the default schema is not what you want:

#![allow(unused)]
fn main() {
use quack_rs::appender::{AppendError, Appender};
use libduckdb_sys::duckdb_connection;
unsafe fn demo(con: duckdb_connection) -> Result<(), AppendError> {
let a = unsafe { Appender::new(con, Some(c"main"), c"events") }?;
let b = unsafe { Appender::with_catalog(con, Some(c"mydb"), Some(c"main"), c"events") }?;
let _ = (a, b);
Ok(())
}
}

Appending a subset of columns

add_column narrows the active column list; the omitted columns take their DEFAULT (or NULL). Both add_column and clear_columns flush every row appended so far.

#![allow(unused)]
fn main() {
use quack_rs::appender::{AppendError, Appender};
unsafe fn demo(appender: &Appender) -> Result<(), AppendError> {
appender.add_column(c"id")?;              // now only `id` is expected
appender.row(|row| row.append_i32(7))?;
appender.clear_columns()?;                // back to every column
Ok(())
}
}

Use TableDescription::column_has_default to find out whether a column has a default before relying on one. append_default() fails for a default that is not a constant, such as nextval('seq') or now(): DuckDB's appender evaluates defaults once, when it is created.

Errors arrive late, and invalidate the batch

Appended rows are buffered. A constraint violation therefore surfaces at flush or close, not at the append_* call that caused it, and it invalidates every buffered row.

With duckdb-1-5, clear discards the offending buffer so you can carry on without re-appending rows that were already committed:

#![allow(unused)]
fn main() {
use quack_rs::appender::Appender;
fn demo(appender: &Appender) {
if let Err(err) = appender.flush() {
    eprintln!("flush failed: {err}");
    let _ = appender.clear(); // drop the offending buffered rows
}
}
}

A row that fails half-way loses the batch

DuckDB counts the values of the current row and cannot take one back. If a row closure fails or panics after its first value went in, the half-written row can be neither finished nor dropped, and DuckDB's close then returns success while writing nothing: every row buffered since the last flush is gone. (DuckDB also flushes by itself each time 204,800 rows accumulate; rows written by such a flush are safe.)

quack-rs does not let that pass silently. The appender becomes poisoned: every later row, append, end_row, flush and close returns an error saying how many buffered rows were not written. With duckdb-1-5, clear() discards them and makes the appender usable again; without it, create a new appender. close() with a row that was started but not ended is an error too (finish the row and close again).

A value that fails as the first of its row loses nothing, and when you append by hand you can retry a rejected value in the same column. If losing a batch is not acceptable, flush() at the points you can afford to go back to.

After a successful close() the appender refuses all further work.

Row order and schema changes

  • Row-at-a-time rows wait in their own buffer until a full chunk (2,048 rows) accumulates, while append_chunk adds its rows to the table-bound buffer directly, so a chunk lands ahead of rows appended before it that are still buffered. Flush first if insertion order matters.
  • Buffered rows are written by column position at flush time. If another connection drops a column and adds one in between, a value lands in the new column with no error.

API

MethodDescription
Appender::new(con, schema, table) (unsafe)Create for table in schema (None = default)
Appender::with_catalog(con, catalog, schema, table) (unsafe)Create fully qualified
column_count() / column_type(i)The active column list
add_column(name) / clear_columns()Narrow / reset the active column list
row(closure)Append one row, calling end_row if the closure succeeds
end_row()Finish the current row explicitly
append_bool/_i8/_i16/_i32/_i64/_i128Signed integers and BOOLEAN
append_u8/_u16/_u32/_u64/_u128Unsigned integers
append_f32/_f64FLOAT / DOUBLE
append_str(&str) / append_bytes(&[u8])VARCHAR (NUL-safe) / BLOB
append_date/_time/_timestamp/_intervalTemporal types
append_value(&Value)Anything else — LIST, STRUCT, MAP, UUID, DECIMAL, ENUM
append_null() / append_default()SQL NULL / the column's DEFAULT
append_chunk(&chunk)Append an entire DataChunk
flush() / close()Flush buffered rows / flush and close
error_message()Message from the last failed operation (Option<String>)
append_default_to_chunk(&chunk, col, row) ¹Write a column's DEFAULT into a chunk cell
clear() ¹Discard buffered, unflushed rows (and un-poison the appender)
error_data() ¹Structured ErrorData from the last failed operation

¹ Requires the duckdb-1-5 feature flag.

Every fallible method returns Result<_, AppendError>. AppendError is ErrorData (message and machine-readable category) when duckdb-1-5 is enabled, and ExtensionError (message only) otherwise — enabling the feature upgrades the error type in place without changing any method's shape.

Safety

new and with_catalog are unsafe: you must pass a valid, open duckdb_connection (such as the one provided to your extension's entry point).

Both return Result for a reason: DuckDB's duckdb_append_* functions do not check whether the appender was created successfully before dereferencing it, so an appender whose creation failed must never be used. A failed create returns Err and no Appender, so such an appender cannot be reached.

Drop

Dropping an Appender closes (and so flushes) it, but the result is ignored — DuckDB's own header notes that after destruction "it is no longer possible to obtain the specific error message". Call close() explicitly whenever the outcome matters.

Table Metadata

TableDescription reads the column metadata of an existing DuckDB table from inside your extension: column names, whether a column has a DEFAULT, and, with duckdb-1-5, the column count and types. Replacement scans, table functions and copy functions use it to inspect a table before deciding what to do.

No feature flag required for creating a description or reading column names and defaults: duckdb_table_description_* has been in the frozen stable prefix of the extension API (slots 292–297) since v1.2.0. Two accessors, column_count and column_type, were added in DuckDB 1.5 in the unstable region and need duckdb-1-5.

#![allow(unused)]
fn main() {
use quack_rs::table_description::TableDescription;
use libduckdb_sys::duckdb_connection;

unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
// SAFETY: `con` is a valid, open connection.
let desc = unsafe { TableDescription::create(con, "main", "events") }?;

assert_eq!(desc.column_name(0).as_deref(), Some("id"));
assert_eq!(desc.column_has_default(0), Some(false));
Ok(())
}
}

with_catalog addresses a table in another catalog, and takes None to mean "the default" for either the catalog or the schema:

#![allow(unused)]
fn main() {
use quack_rs::table_description::TableDescription;
use libduckdb_sys::duckdb_connection;
unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
let desc = unsafe { TableDescription::with_catalog(con, Some("mydb"), None, "events") }?;
let _ = desc;
Ok(())
}
}

API

MethodDescription
TableDescription::create(con, schema, table) (unsafe)Describe schema.table; Err if the table does not exist
TableDescription::with_catalog(con, catalog, schema, table) (unsafe)Describe a fully-qualified table (None = default)
column_name(i)Column name, or None if i is out of range (or the name is not valid UTF-8)
column_has_default(i)Whether the column has a DEFAULT, or None if i is out of range
column_count() ¹Number of columns
column_type(i) ¹Column LogicalType, or None if i is out of range

¹ Requires the duckdb-1-5 feature flag.

Out-of-range indices return None rather than panicking, so a description can be walked without knowing the width up front when duckdb-1-5 is off.

  • Bulk Appender — append_default fills a column without a DEFAULT with NULL, and fails for a DEFAULT that is not a constant (nextval(...), random()), which column_has_default reports as true
  • Type System — what a LogicalType describes

Structured Errors

Requires the duckdb-1-5 feature flag (DuckDB 1.5.0+).

ErrorData is an RAII wrapper around DuckDB's duckdb_error_data handle, the structured error type that several DuckDB 1.5 C API functions return. Unlike a bare error string, an ErrorData carries both a human-readable message and a machine-readable category (DuckDbErrorType), so your extension can branch on the kind of failure (for example, distinguishing Io from OutOfMemory).

Expression::fold, the virtual file system, the Arrow bridge and the appender all report failures as an ErrorData.

Inspecting an error

#![allow(unused)]
fn main() {
use quack_rs::error_data::{DuckDbErrorType, ErrorData};

fn handle(err: ErrorData) {
if err.has_error() {
    match err.error_type() {
        DuckDbErrorType::Io => eprintln!("I/O failure"),
        DuckDbErrorType::OutOfMemory => eprintln!("out of memory"),
        other => eprintln!("{other:?}: {}", err.message().unwrap_or_default()),
    }
}
}
}

Constructing an error

Build a structured error to hand back to DuckDB (for example from a callback):

#![allow(unused)]
fn main() {
// ErrorData is a DuckDB object: fill the dispatch table first.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
use quack_rs::error_data::{DuckDbErrorType, ErrorData};

let err = ErrorData::new(DuckDbErrorType::InvalidInput, "row index out of range");
assert!(err.has_error());
assert_eq!(err.error_type(), DuckDbErrorType::InvalidInput);
}

Propagating with ?

into_extension_error converts an ErrorData into the SDK's ExtensionError, so a structured DuckDB error can flow through the ? operator in your registration or callback logic:

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
use quack_rs::file_system::{FileOpenOptions, FileSystem};
use quack_rs::client_context::ClientContext;
use quack_rs::error_data::ErrorData;

fn read_header(ctx: &ClientContext) -> Result<(), ExtensionError> {
let fs = FileSystem::from_client_context(ctx)
    .ok_or_else(|| ExtensionError::new("no file system"))?;
let handle = fs
    .open(c"data.bin", &FileOpenOptions::read_only())
    .map_err(ErrorData::into_extension_error)?;
let _ = handle;
Ok(())
}
}

API

ItemDescription
ErrorData::new(error_type, message)Construct a structured error
ErrorData::from_raw(raw) (unsafe)Take ownership of a raw duckdb_error_data
has_error()true if the handle represents an actual error
error_type()The DuckDbErrorType category
message()Option<String> — the error text
into_extension_error()Consume into an ExtensionError
is_null() / as_raw() / into_raw()Handle inspection / escape hatches

DuckDbErrorType is a #[non_exhaustive] enum mirroring duckdb_error_type (Io, OutOfMemory, Conversion, Catalog, Constraint, Permission, …). Unknown or future categories map to DuckDbErrorType::Invalid.

UTF-8 validation

The free function check_valid_utf8 exposes DuckDB's own UTF-8 validator. Its rules match Rust's exactly: it accepts every Unicode scalar value and rejects surrogates, overlong forms, code points above U+10FFFF, truncated sequences and stray continuation bytes, just as std::str::from_utf8 does. Use it when you want DuckDB's structured ErrorData for the failure; for a yes/no answer, std::str::from_utf8 is equivalent and needs no database:

#![allow(unused)]
fn main() {
// ErrorData is a DuckDB object: fill the dispatch table first.
std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
use quack_rs::error_data::check_valid_utf8;

fn demo(bytes: &[u8]) {
match check_valid_utf8(bytes) {
    Ok(()) => { /* safe to pass to DuckDB */ }
    Err(err) => eprintln!("invalid UTF-8: {}", err.message().unwrap_or_default()),
}
}
assert!(check_valid_utf8("héllo".as_bytes()).is_ok());
for bad in [&b"\xff"[..], b"\xed\xa0\x80", b"\xc0\xaf", b"\xf4\x90\x80\x80", b"\xe2\x82"] {
    assert!(check_valid_utf8(bad).is_err());
    assert!(std::str::from_utf8(bad).is_err());
}
}

Ownership

ErrorData calls duckdb_destroy_error_data on drop. An ErrorData received from a fallible 1.5 API owns its handle: let it drop, or call into_extension_error() / into_raw() to move the data out.

Bound Expressions

Requires the duckdb-1-5 feature flag (DuckDB 1.5.0+).

Expression is an RAII wrapper around DuckDB's duckdb_expression handle. You obtain one from a scalar function's bind callback via ScalarBindInfo::argument, which lets the bind phase inspect each argument's static type and — when the argument is a constant — fold it to a concrete Value.

Use it for scalar functions whose behaviour depends on a constant argument (a format string, a precision, a regex) that should be validated or pre-computed once at bind time rather than on every row.

Folding a constant argument at bind time

#![allow(unused)]
fn main() {
use quack_rs::scalar::{RawScalarBindInfo, ScalarBindInfo};

unsafe extern "C" fn my_bind(info: RawScalarBindInfo) {
    let bind = unsafe { ScalarBindInfo::new(info) };

    if let Some(arg) = unsafe { bind.argument(0) } {
        // Inspect the argument's static return type.
        let _ty = arg.return_type();

        // If the argument is constant, evaluate it once here instead of
        // recomputing it for every row in the execute callback.
        if arg.is_foldable() {
            let ctx = unsafe { bind.get_client_context() };
            match arg.fold(&ctx) {
                Ok(value) => {
                    // Stash `value` as bind data for the execute phase.
                    let _ = value;
                }
                Err(err) => bind.set_error(&err.message().unwrap_or_default()),
            }
        }
    }
}
}

API

MethodDescription
return_type()Option<LogicalType> — the expression's static type
is_foldable()true if the expression is constant and can be folded
fold(&client_context)Result<Value, ErrorData> — evaluate a constant expression
from_raw(raw) (unsafe) / as_raw() / is_null()Handle inspection / escape hatches

fold only succeeds when is_foldable returns true; otherwise it returns an ErrorData of type InvalidInput. A successful fold can still yield a SQL NULL (for example NULL::INTEGER), whose typed getters return None.

When evaluation itself fails, DuckDB (checked on 1.5.5) hands the C API the exception's JSON form ({"exception_type":"Conversion","exception_message":"...",...}) and always tags it INVALID_INPUT. fold unpacks it: err.message() is the plain message and err.error_type() the type the exception named (Conversion, OutOfRange, ...). Text that is not that JSON is passed through unchanged.

Obtaining an Expression

ScalarBindInfo (the wrapper around a scalar bind callback's duckdb_bind_info) provides two accessors:

MethodReturns
argument(index) (unsafe)Option<Expression> — RAII, the ergonomic path
get_argument(index) (unsafe)raw duckdb_expression — escape hatch

Use argument_count() to bound the index.

Ownership

Expression calls duckdb_destroy_expression on drop. The handle returned by argument() is owned by the caller, so the wrapper cleans it up automatically.

Virtual File System

Requires the duckdb-1-5 feature flag (DuckDB 1.5.0+).

The file_system module exposes DuckDB's virtual file system (VFS) to your extension, so a custom table function, replacement scan, or copy function can read and write files through the same abstraction DuckDB uses internally. Paths then resolve through every registered file system — httpfs (s3://, http://) when it is loaded, in-memory files, and so on — where std::fs only ever sees local disk.

Obtaining a FileSystem

A FileSystem comes from a ClientContext, which most function callbacks can hand you (for example via BindInfo::get_client_context() or ScalarBindInfo::get_client_context()).

Reading a file

#![allow(unused)]
fn main() {
use quack_rs::client_context::ClientContext;
use quack_rs::file_system::{FileOpenOptions, FileSystem};

fn read_all(ctx: &ClientContext) -> Option<Vec<u8>> {
let fs = FileSystem::from_client_context(ctx)?;
let handle = fs.open(c"s3://bucket/data.csv", &FileOpenOptions::read_only()).ok()?;

let mut buf = Vec::new();
handle.read_to_end(&mut buf).ok()?;
Some(buf)
}
}

Writing a file

#![allow(unused)]
fn main() {
use quack_rs::client_context::ClientContext;
use quack_rs::file_system::{FileOpenOptions, FileSystem};

fn write_report(ctx: &ClientContext, bytes: &[u8]) -> Result<(), quack_rs::error_data::ErrorData> {
let fs = FileSystem::from_client_context(ctx).expect("file system");
let handle = fs.open(c"report.bin", &FileOpenOptions::write_create())?;
handle.write_all(bytes)?;
handle.sync()?;
handle.close()?;
Ok(())
}
}

Open options and flags

FileOpenOptions describes how a file is opened. Two convenience constructors cover the common cases; use set_flag for anything else.

Constructor / methodEffect
FileOpenOptions::read_only()Open for reading
FileOpenOptions::write_create()Open for writing, creating if absent
FileOpenOptions::new()Empty; configure with set_flag
set_flag(flag, value)Set an individual FileFlag; returns true on success

FileFlag variants: Read, Write, Create, CreateNew, Append.

  • Create, CreateNew and Append need Write as well.
  • CreateNew ("create, failing if the file exists") also sets Create. DuckDB maps it to FILE_FLAGS_EXCLUSIVE_CREATE, which only has that meaning together with FILE_FLAGS_FILE_CREATE; on its own it neither creates a missing file nor refuses an existing one. On Windows an existing file is not refused: DuckDB's local file system there ignores the exclusive flag and opens with OPEN_ALWAYS (checked in DuckDB 1.5.5), so the existing file is opened, untruncated, without an error.
  • set_flag(flag, false) does not clear a flag: the C API ORs flags in and ignores value. Build a fresh FileOpenOptions instead.

FileHandle operations

MethodReturnsDescription
read(&mut buf)Result<usize, ErrorData>Read up to buf.len() bytes (0 = EOF)
read_exact(&mut buf)Result<(), ErrorData>Read exactly buf.len() bytes, or fail
read_to_end(&mut vec)Result<usize, ErrorData>Append the rest of the file
write(&buf)Result<usize, ErrorData>Write up to buf.len() bytes
write_all(&buf)Result<(), ErrorData>Write all of buf, or fail
seek(position)Result<(), ErrorData>Seek to an absolute byte offset
tell()Result<u64, ErrorData>Current byte offset
size()Result<u64, ErrorData>Total file size in bytes
sync()Result<(), ErrorData>Flush buffered writes to durable storage
close()Result<(), ErrorData>Close the file
error_data()ErrorDataStructured error from the last failed operation

Short reads and short writes are real

duckdb_file_handle_read and duckdb_file_handle_write return "the number of bytes actually read/written" — not a promise that the whole buffer moved. On a local file the two almost always agree; over httpfs a short transfer is common. Prefer read_exact / read_to_end / write_all, which loop for you, and reach for the raw read / write only when a partial transfer is what you want.

tell() and size() are Result for the same reason: the C API reports failure with a negative return value, which is far too easy to clamp to zero and then silently treat as an empty file.

FileSystem exposes open(path, options) and error_data(). Both FileSystem and FileHandle are RAII: they are destroyed (and the handle closed) on drop.

Lifetimes

A FileSystem<'ctx> borrows the ClientContext it came from, and a FileHandle<'fs> borrows the FileSystem that opened it. A handle refers to the database's file system, so using it after the database is closed is a use-after-free (valgrind: invalid read in duckdb::FileHandle::Read); the borrow makes that a compile error instead:

#![allow(unused)]
fn main() {
use quack_rs::client_context::ClientContext;
use quack_rs::file_system::{FileOpenOptions, FileSystem};

fn demo(ctx: &ClientContext) {
    let fs = FileSystem::from_client_context(ctx).unwrap();
    let handle = fs.open(c"data.csv", &FileOpenOptions::read_only()).unwrap();
    drop(fs); // error[E0505]: cannot move out of `fs` because it is borrowed
    let _ = handle.size();
}
}

Selection Vectors

Requires the duckdb-1-5 feature flag (DuckDB 1.5.0+).

A SelectionVector is a list of row indices used to logically reorder or filter a data vector without copying its payload — the building block behind DuckDB's zero-copy filtering. Extensions that implement custom filtering or reordering in vectorized callbacks can allocate one, fill in the indices, and hand its raw handle to the relevant DuckDB vector operations.

This is an advanced, low-level primitive; most extensions never need it.

Allocating and filling

#![allow(unused)]
fn main() {
use quack_rs::selection_vector::SelectionVector;

fn demo() -> Result<(), quack_rs::error::ExtensionError> {
// Select source rows 3, 1, 4, 1, 5 (in that order) — note repeats are allowed.
let mut sel = SelectionVector::new(5)?;
sel.as_mut_slice().copy_from_slice(&[3, 1, 4, 1, 5]);

assert_eq!(sel.len(), 5);
assert_eq!(sel.as_slice(), &[3, 1, 4, 1, 5]);
Ok(())
}
}

The indices are 32-bit (sel_t / u32) and are zeroed by new — DuckDB itself leaves them uninitialised, so the wrapper clears them to keep as_slice() from exposing stale heap contents. Fill them via as_mut_slice().

new returns Err for a length above selection_vector::MAX_LEN: 2^32 (the range of sel_t) on 64-bit targets, and the largest u32 slice that fits in isize::MAX bytes on 32-bit ones. DuckDB does not check the length itself: a large enough length overflows its allocation-size arithmetic into a tiny buffer, and a request of 2^48 bytes or more makes its allocator throw, which would abort the extension.

API

MethodDescription
SelectionVector::new(size)Allocate size zeroed indices; Err above MAX_LEN
len() / is_empty()Number of indices
as_slice()&[u32] — read the indices
as_mut_slice()&mut [u32] — fill the indices
as_raw()The raw duckdb_selection_vector handle for DuckDB vector ops

SelectionVector is RAII: it is destroyed on drop.

Instance Cache

Requires the duckdb-1-5 feature flag (DuckDB 1.5.0+).

An InstanceCache wraps DuckDB's duckdb_instance_cache, which lets several database handles share one underlying DuckDB instance per database path. Opening the same path twice through the cache returns handles backed by the same instance, which avoids the "database is already open in another instance" conflict and saves the cost of re-initialising the database.

This is primarily useful for extensions or host integrations that open secondary databases on behalf of a query.

Opening through the cache

#![allow(unused)]
fn main() {
use quack_rs::instance_cache::InstanceCache;

fn demo() -> Result<(), quack_rs::error::ExtensionError> {
let cache = InstanceCache::new();

// Returns a duckdb_database the caller OWNS and must close with duckdb_close.
let db = cache.get_or_create(c"analytics.db", None)?;
let _ = db;
Ok(())
}
}

Pass a DbConfig to control how a freshly created instance is configured. When an instance already exists for the path, the config must match the one it was created with — a different one is an error ("Can't open a connection to same database file with a different configuration than existing connections"), not silently ignored. None means DuckDB's defaults, so it too conflicts with an instance created with a custom config. An empty path or :memory: is never cached: each call creates a separate in-memory database.

#![allow(unused)]
fn main() {
use quack_rs::instance_cache::InstanceCache;
use quack_rs::config::DbConfig;

fn demo() -> Result<(), quack_rs::error::ExtensionError> {
let cache = InstanceCache::new();
let config = DbConfig::new()?;
// configure `config` as needed...
let db = cache.get_or_create(c"analytics.db", Some(&config))?;
let _ = db;
Ok(())
}
}

API

MethodDescription
InstanceCache::new()Create a new, empty cache
get_or_create(path, config)Open path, creating the instance if needed
as_raw()The raw duckdb_instance_cache handle

get_or_create returns Result<duckdb_database, ExtensionError>; on failure the error carries DuckDB's message.

Threads

InstanceCache is Send + Sync, so one cache can serve several threads: DuckDB's DBInstanceCache guards its map with a mutex and serialises creation of each database. One ordering hazard remains, and it is DuckDB's: while the last handle to a cached file database is being closed, a concurrent get_or_create of the same path can fail with "Unique file handle conflict" because the closing instance still holds the file. Keep one handle open while other threads may open the path, or retry.

Ownership

InstanceCache is RAII and destroys the cache on drop. The duckdb_database returned by get_or_create is, however, owned by the caller: close it with duckdb_close when finished. The cache holds only a weak reference to each instance. An instance stays alive while at least one handle opened through the cache is open, and a handle stays valid after the cache itself is dropped. Once the last handle is closed, the next get_or_create for that path creates a fresh instance.

  • config — DbConfig, the RAII configuration builder accepted here

Arrow Interop

Requires the duckdb-1-5-4 feature flag.

DuckDB's C API has a family of conversion functions (present since 1.4.4) that move data directly between a duckdb_data_chunk and the Arrow C Data Interface, without a query result in between. The quack_rs::arrow module wraps all eight of them.

No arrow crate dependency

The Arrow C Data Interface is an ABI, not a library: ArrowSchema and ArrowArray are plain #[repr(C)] records with a release callback. libduckdb-sys defines them directly — and asserts in its own test suite that they match arrow-rs's FFI_ArrowSchema / FFI_ArrowArray field for field — so quack-rs speaks Arrow without pulling in the arrow crate, and an extension that does use arrow-rs bridges across with a pointer cast.

Why the feature is duckdb-1-5-4 and not duckdb-1-5

All eight C functions are already in duckdb_ext_api_v1 in DuckDB 1.4.4 (extension_api.hpp at v1.4.4, slots 410 to 434; they moved to 411 to 509 in 1.5.0). The floor comes from the bindings: libduckdb-sys declared both records as opaque zero-sized bindgen placeholders (_unused: [u8; 0]) until 1.10504.0, and the caller-allocated structs these functions need cannot be created from a zero-sized type. src/arrow.rs carries a const assertion that fails with exactly that message against an older binding. Because the feature implies duckdb-1-5, an extension built with it also needs a DuckDB 1.5.0+ engine.

The types

TypeWrapsFreed by
ArrowOptionsduckdb_arrow_optionsduckdb_destroy_arrow_options
ArrowSchemathe ArrowSchema ABI recordrelease(schema)
ArrowArraythe ArrowArray ABI recordrelease(array)
ArrowConvertedSchemaduckdb_arrow_converted_schemaduckdb_destroy_arrow_converted_schema

RawArrowSchema and RawArrowArray are the ABI records themselves, re-exported for code that already speaks the raw interface.

Exporting a chunk

ArrowOptions carries the settings DuckDB renders Arrow with — the timezone for TIMESTAMPTZ, the string offset width, registered extension types. Take them from the result whose chunks you are exporting, so schema and data agree:

#![allow(unused)]
fn main() {
use quack_rs::arrow::{data_chunk_to_arrow, to_arrow_schema};
use quack_rs::query::QueryResult;

fn demo(result: &mut QueryResult) -> Result<(), Box<dyn std::error::Error>> {
// SAFETY: the connection that ran this query stays open while `options` is
// in use — the options point at that connection's client context.
let options = unsafe { result.arrow_options() }?;

let columns: Vec<(String, quack_rs::types::LogicalType)> = (0..result.column_count())
    .filter_map(|i| Some((result.column_name(i)?, result.column_logical_type(i)?)))
    .collect();
let pairs: Vec<(&str, &quack_rs::types::LogicalType)> =
    columns.iter().map(|(n, t)| (n.as_str(), t)).collect();

let schema = to_arrow_schema(&options, &pairs)?;
assert_eq!(schema.format(), Some("+s")); // a record batch is a struct

while let Some(chunk) = result.next_chunk()? {
    let array = data_chunk_to_arrow(&options, &chunk)?;
    // hand `array` (plus `schema`) to any Arrow consumer
    let _ = array;
}
Ok(())
}
}

ArrowOptions must not outlive its connection

The options hold a raw pointer to the connection's client context, and the conversion functions dereference it. Used after the connection is closed they read freed memory. ArrowOptions<'conn> carries that lifetime:

  • ArrowOptions::from_connection(&con) is safe. It borrows the OwnedConnection, so the compiler rejects a drop(con) while the options are still in use.
  • ArrowOptions::from_raw_connection(raw) (for a raw duckdb_connection) and QueryResult::arrow_options() / ArrowOptions::from_result are unsafe. A QueryResult does not borrow the connection that ran it, so nothing checks that the connection is still open. The caller must keep it open for as long as the options are used.

Importing an array

Going the other way needs the Arrow schema translated into DuckDB's own type descriptors first. That translation is reusable — do it once, not per batch:

#![allow(unused)]
fn main() {
use quack_rs::arrow::{data_chunk_from_arrow, schema_from_arrow, ArrowArray, ArrowSchema};

fn demo(
    con: libduckdb_sys::duckdb_connection,
    schema: &mut ArrowSchema,
    array: ArrowArray,
) -> Result<(), Box<dyn std::error::Error>> {
// SAFETY: `con` is a live DuckDB connection.
let converted = unsafe { schema_from_arrow(con, schema) }?;
// SAFETY: same connection; `array` was built against `schema`.
let chunk = unsafe { data_chunk_from_arrow(con, array, &converted) }?;
let _ = chunk.size();
Ok(())
}
}

Note the asymmetry, which mirrors what DuckDB actually does:

  • schema_from_arrow borrows the schema. You still own it and it is released when its ArrowSchema drops.
  • data_chunk_from_arrow takes the array by value. DuckDB sets arrow_array->release = nullptr before the conversion loop body — so it claims the array even when the conversion then fails. The by-value binding is still dropped on the way out, which releases the array in the one case where DuckDB does not claim it (a zero-column schema, where the loop never runs).

The resulting chunk keeps the Arrow buffers alive, so most columns share the data rather than copy it. Dictionary-encoded columns are the exception: the wrapper copies them into flat vectors.

What the wrapper refuses that DuckDB would not

duckdb_data_chunk_from_arrow indexes arrow_array->children[i] once per column in the converted schema with no bounds check, dereferences each child without a null check, reads offset + length rows from each child without comparing its length, and dereferences an array without checking whether it was already released. Those are segfaults or out-of-bounds reads, not errors. data_chunk_from_arrow checks them first — which is why ArrowConvertedSchema remembers the column count of the schema it was built from — and returns an InvalidInput error instead.

It also refuses a zero-row array. DuckDB passes arrow_array->length through as the chunk's capacity, and a capacity of zero trips a D_ASSERT(size > 0) that aborts a debug build of DuckDB, while a release build carries on. Skip empty batches, or create the empty chunk directly with duckdb_create_data_chunk.

What it cannot check, and what data_chunk_from_arrow's # Safety section therefore makes the caller's job:

  • The array must conform to the converted schema. Nothing in an Arrow array records its type, so DuckDB reads each child's buffers as the format the schema declares. An int32 child imported under a utf8 schema has its values read as string offsets into a buffer that does not exist. Arrays exported with data_chunk_to_arrow under the schema you converted conform.
  • A fixed-width dictionary whose indices can be NULL needs one element of padding past its values: DuckDB points NULL indices at an entry there, and the copy data_chunk_from_arrow makes of every dictionary column reads it (docs/upstream-duckdb-reports.md, item 29).
  • The buffers must be as long as the lengths say, and length must be the true row count. DuckDB allocates the chunk for length rows before its error handling starts, so an absurd length is an allocation failure that aborts the process.

It also refuses valid Arrow layouts that DuckDB imports wrongly: it walks the array alongside its schema and returns InvalidInput, naming the node, for

  • an offset below the top level that DuckDB applies to the wrong rows: a struct inside an offset struct or a list, a union's members, a run-end-encoded array's value validity (docs/upstream-duckdb-reports.md, item 24);
  • a dictionary with NULLs under a list that starts past element 0, or with more than 2048 rows and NULLs of its own or an enclosing struct's, which DuckDB copies past a 2048-row heap mask (items 9 and 25);
  • a dictionary whose values are themselves dictionary-encoded (item 26);
  • list views that overlap or leave gaps (item 27);
  • a sparse union whose +us: type codes are not 0, 1, … (item 28), or whose null_count is not 0 (item 32);
  • a dictionary whose null_count is -1 ("not computed", item 31);
  • a geoarrow.wkb column read as more than 2048 rows (item 33);
  • a run-end-encoded array where DuckDB reads a plain one: a fixed-size list's child, or another run-end array's values (item 24).

Arrays that data_chunk_to_arrow produced, paired with the schema they were produced with, never take these shapes. Arrays from other producers can: one that slices a nested array without copying it may leave offsets below the top level. Copying the slice before export avoids them. Every error DuckDB reports from the conversion arrives as InvalidInput.

Round trips are not always exact

Two types come back different from an Arrow round trip through DuckDB's own converters:

  • TIMETZ comes back as TIME with the offset dropped: 01:02:03+05:30 returns as 01:02:03.
  • BIT comes back as BLOB.

Check the converted types (ArrowConvertedSchema) when a round trip must be lossless.

DuckDB would export three kinds of value wrongly, with no error (checked on 1.4.4, 1.5.0 and 1.5.5), so data_chunk_to_arrow checks the chunk first and refuses one that holds such a value, at any nesting depth:

  • An INTERVAL whose microseconds exceed about ±106,751 days (2,562,047 hours) would wrap, because Arrow counts nanoseconds in an i64 and DuckDB multiplies by 1000 unchecked: INTERVAL 2562048 HOUR would export as a negative interval.
  • A 39-digit UHUGEINT would export as a decimal128(38, 0) it does not fit; from 2^127 it comes out negative (2^128 - 1 becomes -1).
  • A 39-digit HUGEINT would export as a decimal128(38, 0) it does not fit, unless arrow_lossless_conversion is set (then it exports as a 16-byte fixed-size binary and is not refused).

After the export, the array is also checked against the schema DuckDB declares for the chunk's types. Before 1.5.5, BIGNUM (and, from 1.5.0, GEOMETRY) exported under arrow_output_version = '1.4' is written as binary views while the schema declares plain binary, which a consumer reads as offsets (docs/upstream-duckdb-reports.md, item 34); such an export is refused. Set arrow_output_version = '1.0' on those releases.

Bridging to arrow-rs

This sketch uses the arrow crate's FFI_ArrowArray, which quack-rs does not depend on, so it is not compiled with the book:

// quack-rs -> arrow-rs
let ffi: FFI_ArrowArray = unsafe { std::mem::transmute(array.into_raw()) };

// arrow-rs -> quack-rs, neutralising the source so only one side releases
let array = unsafe { ArrowArray::take_from(std::ptr::from_mut(&mut ffi).cast()) };

take_from moves the record out and writes a released placeholder back, so the foreign wrapper's own Drop becomes a no-op instead of a double free.

Thread safety

None of these types are Send or Sync. The Arrow C Data Interface says nothing about which thread may call release, and duckdb_arrow_options wraps a ClientProperties tied to the connection's client context.

TLS Configuration

Extensions that make outbound HTTPS connections (e.g., fetching remote data, calling REST APIs) need a way to inject TLS configuration — client certificates for mTLS, custom CA bundles, or restricted cipher suites.

The tls module provides the TlsConfigProvider trait so that extensions can supply their TLS setup through a uniform interface, regardless of which TLS library they use (rustls, native-tls, etc.).

Design

The trait is type-erased via Arc<dyn Any + Send + Sync> so that quack-rs does not depend on any specific TLS library. The code that consumes the provider downcasts the returned Arc to the concrete config type, after checking config_type_name(). TlsConfigProvider requires Send + Sync.

Implementing a TLS Provider

#![allow(unused)]
fn main() {
use quack_rs::tls::{TlsConfigProvider, TlsVersion};
use quack_rs::error::ExtensionError;
use std::any::Any;
use std::sync::Arc;

struct MyTlsProvider {
    // In practice: Arc<rustls::ClientConfig>
    config: Arc<String>,
    mtls_enabled: bool,
}

impl TlsConfigProvider for MyTlsProvider {
    fn client_config(&self) -> Result<Arc<dyn Any + Send + Sync>, ExtensionError> {
        Ok(self.config.clone())
    }

    fn provider_name(&self) -> &str { "my-extension-tls" }
    fn config_type_name(&self) -> &str { "String" } // in practice: "rustls::ClientConfig"

    fn min_tls_version(&self) -> TlsVersion {
        TlsVersion::Tls12  // Minimum recommended
    }

    fn supports_mtls(&self) -> bool { self.mtls_enabled }

    fn accepts_invalid_certs(&self) -> bool {
        false  // MUST default to false
    }
}
}

Security Requirements

Implementations must:

  • Return false from accepts_invalid_certs() unless explicitly configured otherwise by the user. Certificate validation bypass (CWE-295) should never be the default.
  • Return TlsVersion::Tls12 or higher from min_tls_version(). TLS 1.0 and 1.1 are deprecated per RFC 8996.
  • Emit an ExtensionWarning via WarningCollector when certificate validation is disabled or when using a TLS version below 1.2.

Auditing a Provider

audit_tls_provider() checks a provider for two common misconfigurations and returns one ExtensionWarning per problem found:

  • Certificate verification bypass (CWE-295): code TLS_NO_VERIFY, severity High
  • A deprecated minimum version, TLS 1.0 or 1.1 (CWE-327): code TLS_DEPRECATED_VERSION, severity Medium

Feed the result into a WarningCollector:

#![allow(unused)]
fn main() {
use quack_rs::error::ExtensionError;
use quack_rs::tls::{audit_tls_provider, TlsConfigProvider, TlsVersion};
use quack_rs::warning::WarningCollector;
use std::any::Any;
use std::sync::Arc;

struct InsecureProvider;

impl TlsConfigProvider for InsecureProvider {
    fn client_config(&self) -> Result<Arc<dyn Any + Send + Sync>, ExtensionError> {
        Ok(Arc::new(()))
    }
    fn provider_name(&self) -> &str { "insecure" }
    fn config_type_name(&self) -> &str { "()" }
    fn min_tls_version(&self) -> TlsVersion { TlsVersion::Tls10 }
    fn supports_mtls(&self) -> bool { false }
    fn accepts_invalid_certs(&self) -> bool { true }
}

let collector = WarningCollector::new();
for w in audit_tls_provider(&InsecureProvider) {
    collector.emit(w);
}

let codes: Vec<&str> = collector.snapshot().iter().map(|w| w.code).collect();
assert_eq!(codes, ["TLS_NO_VERIFY", "TLS_DEPRECATED_VERSION"]);
}

Downcasting Safely

Never use .unwrap() or .expect() when downcasting in FFI callback contexts (see Pitfall L3). Handle a failed downcast as an error:

#![allow(unused)]
fn main() {
use std::any::Any;
use std::sync::Arc;
use quack_rs::error::ExtensionError;

// With rustls this would be `downcast_ref::<rustls::ClientConfig>()`.
fn use_config(config: &Arc<dyn Any + Send + Sync>) -> Result<(), ExtensionError> {
    let text = config
        .downcast_ref::<String>()
        .ok_or_else(|| ExtensionError::new("expected a String TLS config"))?;
    let _ = text;
    Ok(())
}

let config: Arc<dyn Any + Send + Sync> = Arc::new(String::from("pem bundle"));
assert!(use_config(&config).is_ok());
let wrong: Arc<dyn Any + Send + Sync> = Arc::new(42_u32);
assert!(use_config(&wrong).is_err());
}

Structured Warnings

Extensions that access external resources (network, files, credentials) should emit structured warnings when potentially unsafe operations occur. The warning module provides a consistent, thread-safe API for collecting and surfacing security warnings.

Core Types

ExtensionWarning

A structured warning with:

FieldTypeDescription
code&'static strMachine-readable code (e.g., "TLS_NO_VERIFY")
severityWarningSeverityInfo / Low / Medium / High / Critical
messageStringHuman-readable description
cweOption<u32>Optional CWE identifier

WarningSeverity

Five levels mirroring common security advisory severity. WarningSeverity implements Ord in this order, so severity >= WarningSeverity::High selects the warnings that need attention:

  • Info — no security impact, but worth noting
  • Low — minimal security impact
  • Medium — potential security concern
  • High — significant security risk
  • Critical — immediate action recommended

WarningCollector

A thread-safe collector backed by Mutex<Vec<ExtensionWarning>>. Share it across threads via Arc<WarningCollector>. A panic on another thread while it held the lock does not lose warnings: the collector recovers the list from the poisoned lock and keeps using it.

Usage

#![allow(unused)]
fn main() {
use quack_rs::warning::{ExtensionWarning, WarningSeverity, WarningCollector};

let collector = WarningCollector::new();

// Emit a warning when detecting an insecure configuration
collector.emit(ExtensionWarning {
    code: "TLS_NO_VERIFY",
    severity: WarningSeverity::High,
    message: "TLS certificate verification is disabled".into(),
    cwe: Some(295),
});

// Check warnings
assert_eq!(collector.len(), 1);
assert!(!collector.is_empty());

// Read without clearing
let snapshot = collector.snapshot();
assert_eq!(snapshot.len(), 1);
assert_eq!(collector.len(), 1);  // still there

// Consume all warnings
let warnings = collector.drain();
assert_eq!(warnings.len(), 1);
assert!(collector.is_empty());  // now empty
}

Display Format

ExtensionWarning implements Display as [SEVERITY] CODE: message (CWE-nnn); the CWE suffix is omitted when cwe is None:

[HIGH] TLS_NO_VERIFY: TLS certificate verification is disabled (CWE-295)
[MEDIUM] TLS_DEPRECATED_VERSION: TLS provider "my-tls" allows deprecated TLS 1.0 (RFC 8996) (CWE-327)

Integration with TLS Auditing

tls::audit_tls_provider() returns a Vec<ExtensionWarning> that can be fed straight into a WarningCollector; see Auditing a Provider for a complete example.

Best Practices

  • Create a single WarningCollector per extension (typically in global or bind-data state)
  • Use snapshot() for read-only diagnostics; use drain() when consuming warnings for output
  • Include a CWE identifier whenever one applies
  • Surface collected warnings through a table function of your own (for example SELECT * FROM __extension_warnings())

Secrets Management

Extensions that access external services (HTTP APIs, databases, cloud storage) commonly need credentials. DuckDB has a native secrets system (CREATE SECRET), but the extension C API exposes none of it: duckdb_ext_api_v1 has no duckdb_secret_* functions, so an extension cannot ask DuckDB for a credential.

The secrets module therefore offers two separate things:

  • list_duckdb_secrets reads the metadata of the secrets the user has configured, with sensitive fields redacted by DuckDB.
  • SecretsManager and SecretEntry are a trait and a type for the credential source an extension has to provide itself (an environment variable, a config option, a file, its own key store), with leak-resistant defaults already in place.

Listing DuckDB's secrets

list_duckdb_secrets queries the duckdb_secrets() table function and returns one DuckDbSecretInfo per secret: name, secret_type, provider, persistent, storage, scope (the URI prefixes it applies to) and secret_string. DuckDB redacts sensitive fields in that table, so a secret created with SECRET 'super-secret-value' comes back as ...;secret=redacted. Use it to find out which secrets exist and what they cover, not to authenticate:

#![allow(unused)]
fn main() {
use quack_rs::secrets::list_duckdb_secrets;
use libduckdb_sys::duckdb_connection;
unsafe fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
// SAFETY: `con` is a valid, open connection.
let secrets = unsafe { list_duckdb_secrets(con) }?;
if !secrets.iter().any(|s| s.secret_type == "s3") {
    eprintln!("no S3 secret configured; run CREATE SECRET (TYPE s3, ...)");
}
Ok(())
}
}

Core Types

SecretEntry

A single secret entry with metadata and key-value fields. Designed to minimize accidental credential leakage:

  • Debug redacts field values and the scope — field keys are shown, and every value and a non-empty scope are replaced with "[REDACTED]"
  • Drop zeroizes sensitive data — every field key and value, the provider and the scope are overwritten with zeros using std::ptr::write_volatile before deallocation. This covers the buffers a SecretEntry owns, not a String the caller passed in and still holds
  • No PartialEq — prevents accidental non-constant-time comparisons of secret material
  • Clone duplicates the secret — it is supported, but each clone is another copy of the credential in memory

SecretsManager

The trait an extension implements over its own credential source. It requires Send + Sync:

#![allow(unused)]
fn main() {
use quack_rs::secrets::{SecretEntry, SecretsManager};

struct MySecrets {
    entries: Vec<SecretEntry>,
}

impl SecretsManager for MySecrets {
    fn get_secret(&self, name: &str, secret_type: &str) -> Option<SecretEntry> {
        self.entries.iter()
            .find(|e| e.name() == name && e.secret_type() == secret_type)
            .cloned()
    }

    fn list_secrets(&self, secret_type: Option<&str>) -> Vec<SecretEntry> {
        self.entries.iter()
            .filter(|e| secret_type.is_none() || secret_type == Some(e.secret_type()))
            .cloned()
            .collect()
    }

    fn remove_secret(&self, _name: &str, _secret_type: &str) -> bool {
        false // read-only example
    }
}
}

Building Secret Entries

Use the builder pattern:

#![allow(unused)]
fn main() {
use quack_rs::secrets::SecretEntry;

let entry = SecretEntry::new("my_api_key", "bearer")
    .with_provider("config")
    .with_scope("https://api.example.com")
    .with_field("token", "sk-abc123")
    .with_field("refresh_token", "xyz789");

assert_eq!(entry.name(), "my_api_key");
assert_eq!(entry.secret_type(), "bearer");
assert_eq!(entry.get_field("token"), Some("sk-abc123"));
}

Safe Diagnostics

Use field_keys() for logging without leaking secrets:

#![allow(unused)]
fn main() {
use quack_rs::secrets::SecretEntry;

let entry = SecretEntry::new("key", "s3")
    .with_field("access_key", "AKIA...")
    .with_field("secret_key", "wJalr...");

// Safe for logging — returns keys only, no values
let mut keys = entry.field_keys();
keys.sort_unstable(); // the fields are a HashMap, so the order is unspecified
assert_eq!(keys, ["access_key", "secret_key"]);
}

Debug Output

The Debug implementation redacts field values and the scope. For an entry with a scope and one field, {:#?} prints:

SecretEntry {
    name: "api_key",
    secret_type: "bearer",
    provider: "config",
    scope: "[REDACTED]",
    fields: {
        "token": "[REDACTED]",
    },
}

Security Best Practices

  1. Never log secret field values — use field_keys() for diagnostics
  2. Drop clones promptly — minimize the window during which sensitive data resides in memory
  3. Implement remove_secret with zeroization — don't just remove the reference; zeroize the data before deallocation
  4. Thread safety — SecretsManager requires Send + Sync, because DuckDB may run the callbacks that consult it on several threads at once

Testing Guide

This page covers how to test a DuckDB extension written with quack-rs. The strategy has two tiers: pure-Rust unit tests for business logic (no DuckDB required), and SQLLogicTest end-to-end (E2E) tests that load the packaged extension into a real DuckDB process. Between the two, the InMemoryDb helper lets cargo test register and call your real callbacks against a bundled DuckDB.


Architectural limitation: the loadable-extension dispatch wall

This is the most important thing to understand before writing tests.

DuckDB loadable extensions use libduckdb-sys with features = ["loadable-extension"]. This intentionally does not link the DuckDB runtime into the extension binary. Instead, every DuckDB C API call (duckdb_vector_get_data, duckdb_create_logical_type, etc.) goes through a dispatch table: one global AtomicPtr per C API function, filled in only when the extension's entry point calls duckdb_rs_extension_api_init as DuckDB loads it.

In cargo test, no DuckDB process loads your extension. The dispatch table is never initialized, and the first call to any DuckDB C API function panics:

DuckDB API not initialized or DuckDB feature omitted

Opening an InMemoryDb fills the table for the whole test process, after which the APIs below work.

What this breaks

APIWhy it fails
VectorReader::newcalls duckdb_vector_get_data
VectorWriter::newcalls duckdb_vector_get_data
Connection::register_*calls the DuckDB registration C API
LogicalType::newcalls duckdb_create_logical_type
LogicalType::dropcalls duckdb_destroy_logical_type
BindInfo::add_result_columncalls duckdb_bind_add_result_column

What still works in cargo test

APIWhy it works
AggregateTestHarnesspure Rust, zero DuckDB dependency
MockVectorWriter / MockVectorReaderin-memory buffers, zero DuckDB dependency
MockRegistrarrecords registrations without calling the C API
SqlMacro::to_sql()generates SQL strings, no DuckDB needed
interval_to_microspure arithmetic
validate / scaffoldpure Rust
InMemoryDblinks a real DuckDB via the duckdb crate (bundled-test or bundled-test-prebuilt feature)

Mock types for callback logic

Keep the per-row computation in a plain Rust function and test that function directly — it needs no vectors at all. The FFI callback is then a thin loop around it, and for the common shapes the typed constructors (ScalarFunctionBuilder::map1, map1_str, …) write that loop for you, NULL handling included.

#![allow(unused)]
fn main() {
// The logic: plain Rust, tested with plain `#[test]`s.
fn shout(s: &str) -> String {
    s.to_uppercase()
}

#[test]
fn shout_uppercases() {
    assert_eq!(shout("hello"), "HELLO");
}
}
#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::vector::{VectorReader, VectorWriter};
fn shout(s: &str) -> String { s.to_uppercase() }

// The callback: a thin loop over the real vectors.
unsafe extern "C" fn shout_callback(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let rows = usize::try_from(unsafe { libduckdb_sys::duckdb_data_chunk_get_size(input) })
        .unwrap_or(0);
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };
    for row in 0..rows {
        if unsafe { reader.is_valid(row) } {
            let out = shout(unsafe { reader.read_str(row) });
            unsafe { writer.write_varchar(row, &out) };
        } else {
            unsafe { writer.set_null(row) };
        }
    }
}
}

To test the loop itself against real vectors, use InMemoryDb (below): it runs the callback inside a real DuckDB.

MockVectorReader and MockVectorWriter are in-memory stand-ins with the same method names as VectorReader and VectorWriter. They are separate types, so a function written against the mocks cannot be handed the real reader and writer; they are for prototyping and checking row-loop logic without a database. They do reproduce the behaviour of a real vector that a more forgiving mock would hide:

  • set_null clears a validity bit and a later write_* does not set it again (a real vector keeps returning NULL for that row);
  • a row that is never written is valid, not NULL — use is_written to check a loop wrote every row;
  • writing past the capacity given to MockVectorWriter::new panics.
#![allow(unused)]
fn main() {
use quack_rs::testing::{MockVectorReader, MockVectorWriter};

let reader = MockVectorReader::from_strs([Some("hello"), None, Some("world")]);
let mut writer = MockVectorWriter::new(3);
for i in 0..reader.row_count() {
    match reader.try_get_str(i) {
        Some(s) => writer.write_varchar(i, &s.to_uppercase()),
        None => writer.set_null(i),
    }
}
assert_eq!(writer.try_get_str(0), Some("HELLO"));
assert!(writer.is_null(1));
assert_eq!(writer.try_get_str(2), Some("WORLD"));
assert!((0..3).all(|i| writer.is_written(i) || writer.is_null(i)));
}

Testing registration with MockRegistrar

MockRegistrar implements the Registrar trait without calling any DuckDB C API. Use it to verify your registration function registers the right set of functions:

#![allow(unused)]
fn main() {
use quack_rs::connection::Registrar;
use quack_rs::testing::MockRegistrar;
use quack_rs::scalar::ScalarFunctionBuilder;
use quack_rs::types::TypeId;
use quack_rs::error::ExtensionError;
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};

unsafe extern "C" fn upper(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
unsafe extern "C" fn lower(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}

fn register_all(reg: &impl Registrar) -> Result<(), ExtensionError> {
    let upper = ScalarFunctionBuilder::new("upper_ext")
        .param(TypeId::Varchar)
        .returns(TypeId::Varchar)
        .function(upper);
    let lower = ScalarFunctionBuilder::new("lower_ext")
        .param(TypeId::Varchar)
        .returns(TypeId::Varchar)
        .function(lower);
    unsafe {
        reg.register_scalar(upper)?;
        reg.register_scalar(lower)?;
    }
    Ok(())
}

#[test]
fn test_register_all() {
    let mock = MockRegistrar::new();
    register_all(&mock).unwrap();
    assert_eq!(mock.total_registrations(), 2);
    assert!(mock.has_scalar("upper_ext"));
    assert!(mock.has_scalar("lower_ext"));
}
}

MockRegistrar refuses, with the same error, what the real registration refuses before it calls DuckDB: a missing return type or callback, an empty function set, a copy function with neither direction, a config option without a type or default, a composite or literal TypeId in any slot, an ANY return type. Checks that need DuckDB (a name or signature already taken, a type the running DuckDB lacks, a config default that does not convert) are not run.

Limitation: MockRegistrar cannot be used with builders that hold LogicalType values (created via .returns_logical() or .param_logical()), because LogicalType::drop calls duckdb_destroy_logical_type, which panics while the dispatch table is uninitialised. Use TypeId parameters with MockRegistrar.


SQL-level testing with InMemoryDb (bundled-test feature)

For SQL-level assertions — verifying that a SQL macro produces the correct output, or that a CREATE TABLE + INSERT + SELECT pipeline works — enable the bundled-test Cargo feature. This provides InMemoryDb, which wraps the duckdb crate's bundled DuckDB and automatically initialises the loadable-extension dispatch table before opening a connection (see Pitfall P9).

Two features expose InMemoryDb; pick the one that fits your build-time budget:

# Zero-config but slow: compile libduckdb from C++ source (~5–10 min cold).
[dev-dependencies]
quack-rs = { version = "0.18", features = ["bundled-test"] }
# Fast: link against a pre-built libduckdb. Set DUCKDB_DOWNLOAD_LIB=1 at build
# time and libduckdb-sys downloads the upstream release zip (~40 MB, cached
# under target/); or set DUCKDB_LIB_DIR=/path/to/libduckdb if you already have
# one extracted. Header discovery needs libduckdb-sys >= 1.10503; with an older
# one, also set DUCKDB_INCLUDE_DIR.
[dev-dependencies]
quack-rs = { version = "0.18", features = ["bundled-test-prebuilt"] }

Both keep duckdb out of a plain cargo test and out of your published crate's dependency tree — it is pulled in only when one of these features is on.

If your tests need to LOAD your own locally built .duckdb_extension file, use InMemoryDb::open_unsigned instead of open(): the allow_unsigned_extensions option can only be set at startup, not with SET after the database is running.

#![allow(unused)]
fn main() {
use quack_rs::testing::InMemoryDb;
use quack_rs::sql_macro::SqlMacro;

#[test]
fn test_clamp_macro_sql() {
    let db = InMemoryDb::open().unwrap();

    // Generate and execute the CREATE MACRO SQL
    let m = SqlMacro::scalar("clamp", &["x", "lo", "hi"], "greatest(lo, least(hi, x))").unwrap();
    db.execute_batch(&m.to_sql()).unwrap();

    // Verify correct output
    let result: i64 = db.query_one("SELECT clamp(5, 1, 10)").unwrap();
    assert_eq!(result, 5);

    let clamped: i64 = db.query_one("SELECT clamp(15, 1, 10)").unwrap();
    assert_eq!(clamped, 10);
}
}

Testing your FFI callbacks for real

Opening an InMemoryDb also populates the loadable-extension dispatch table — for the whole process, not just that handle. After that the entire C API works, so you can register a real function and call it from SQL inside cargo test:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_connection, DuckDBSuccess};
use quack_rs::data_chunk::DataChunk;
use quack_rs::query::query;
use quack_rs::scalar::ScalarFunctionBuilder;
use quack_rs::testing::InMemoryDb;
use quack_rs::types::TypeId;
use quack_rs::vector::VectorWriter;

quack_rs::scalar_callback!(triple_it, |_info, input, output| {
    let chunk = unsafe { DataChunk::from_raw(input) };
    let reader = unsafe { chunk.reader(0) };
    let mut writer = unsafe { VectorWriter::from_vector(output) };
    for row in 0..chunk.size() {
        unsafe { writer.write_i64(row, reader.read_i64(row) * 3) };
    }
});

#[test]
fn triple_it_works() {
    // 1. Initialise the dispatch table.
    let _dispatch = InMemoryDb::open().unwrap();

    // 2. Open a raw connection.
    let mut db = std::ptr::null_mut();
    let mut con: duckdb_connection = std::ptr::null_mut();
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), DuckDBSuccess);
    }

    // 3. Register with the usual builder.
    unsafe {
        ScalarFunctionBuilder::try_new("triple_it").unwrap()
            .param(TypeId::BigInt)
            .returns(TypeId::BigInt)
            .function(triple_it)
            .register(con)
            .unwrap();
    }

    // 4. Run SQL and assert on the answer.
    let mut result = unsafe { query(con, "SELECT triple_it(14)") }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    assert_eq!(unsafe { chunk.reader(0).read_i64(0) }, 42);
}
}

This is the most valuable coverage available inside cargo test: it exercises the builder, DuckDB's planner, your extern "C" callback, and the vector accessors' pointer arithmetic in one go. tests/ffi_roundtrip.rs in the quack-rs repository does this for every vector type.

The mocks are still the right tool for unit-testing callback logic without a database, and are the only option when the bundled-test features are off.


Why two tiers?

Pitfall P3 — Unit tests are insufficient. 435 unit tests passed in duckdb-behavioral while the extension had three critical bugs: a SEGFAULT on load, 6 of 7 functions not registering, and wrong results from a combine bug. E2E tests caught all three.

Test tierWhat it catchesWhat it misses
Unit testsLogic bugs in state structsFFI wiring, registration failures, SEGFAULT
E2E testsFFI wiring, registration, load-time crashes, wrong resultsOnly the inputs you did not write a test for

Both tiers are required. Unit tests give fast, deterministic feedback. E2E tests prove the extension actually works inside DuckDB.


Unit tests with AggregateTestHarness

AggregateTestHarness<S> simulates the DuckDB aggregate lifecycle in pure Rust without any DuckDB dependency:

flowchart LR
    N["new()"] --> U["update() × N"]
    U --> C["combine() <i>(optional)</i>"]
    C --> F["finalize()"]

Basic usage

#![allow(unused)]
fn main() {
use quack_rs::testing::AggregateTestHarness;
use quack_rs::aggregate::AggregateState;

#[derive(Default, Debug, PartialEq)]
struct SumState { total: i64 }
impl AggregateState for SumState {}

#[test]
fn test_sum() {
    let mut h = AggregateTestHarness::<SumState>::new();
    h.update(|s| s.total += 10);
    h.update(|s| s.total += 20);
    h.update(|s| s.total += 5);
    assert_eq!(h.finalize().total, 35);
}
}

Convenience: aggregate

For testing over a collection of inputs:

#![allow(unused)]
fn main() {
use quack_rs::aggregate::AggregateState;
use quack_rs::testing::AggregateTestHarness;
#[derive(Default)] struct WordCountState { count: usize }
impl AggregateState for WordCountState {}
fn count_words(s: &str) -> usize { s.split_whitespace().count() }
#[test]
fn test_word_count() {
    let result = AggregateTestHarness::<WordCountState>::aggregate(
        ["hello world", "one", "two three four", ""],
        |s, text| s.count += count_words(text),
    );
    assert_eq!(result.count, 6);  // 2 + 1 + 3 + 0
}
}

Testing combine (Pitfall L1)

DuckDB creates fresh target states — set up by state_init, which with FfiState<T> means T::default() — and calls combine to merge into them. combine must propagate all fields, including configuration fields, not just accumulated data. Test this explicitly:

#![allow(unused)]
fn main() {
use quack_rs::aggregate::AggregateState;
use quack_rs::testing::AggregateTestHarness;
#[derive(Default)] struct MyState { window_size: i64, count: i64 }
impl AggregateState for MyState {}
#[test]
fn combine_propagates_config() {
    let mut h1 = AggregateTestHarness::<MyState>::new();
    h1.update(|s| {
        s.window_size = 3600;  // config field
        s.count += 5;          // data field
    });

    // h2 simulates a fresh target state: `state_init` gave it `MyState::default()`
    let mut h2 = AggregateTestHarness::<MyState>::new();

    h2.combine(&h1, |src, tgt| {
        tgt.window_size = src.window_size;  // MUST propagate config
        tgt.count += src.count;
    });

    let result = h2.finalize();
    assert_eq!(result.window_size, 3600);  // Would be 0 if forgotten
    assert_eq!(result.count, 5);
}
}

Inspecting intermediate state

#![allow(unused)]
fn main() {
use quack_rs::aggregate::AggregateState;
use quack_rs::testing::AggregateTestHarness;
#[derive(Default)] struct SumState { total: i64 }
impl AggregateState for SumState {}
let mut h = AggregateTestHarness::<SumState>::new();
h.update(|s| s.total += 5);
assert_eq!(h.state().total, 5);   // borrow without consuming
h.update(|s| s.total += 3);
assert_eq!(h.state().total, 8);
}

Resetting

#![allow(unused)]
fn main() {
use quack_rs::aggregate::AggregateState;
use quack_rs::testing::AggregateTestHarness;
#[derive(Default)] struct SumState { total: i64 }
impl AggregateState for SumState {}
let mut h = AggregateTestHarness::<SumState>::new();
h.update(|s| s.total = 999);
h.reset();
assert_eq!(h.state().total, 0);  // back to S::default()
}

Pre-populating state

#![allow(unused)]
fn main() {
use quack_rs::aggregate::AggregateState;
use quack_rs::testing::AggregateTestHarness;
#[derive(Default)] struct MyState { window_size: i64, count: i64 }
impl AggregateState for MyState {}
let initial = MyState { window_size: 3600, count: 0 };
let h = AggregateTestHarness::with_state(initial);
assert_eq!(h.finalize().window_size, 3600);
}

Unit tests for scalar functions

Scalar logic is pure Rust — test it directly:

#![allow(unused)]
fn main() {
// From examples/hello-ext/src/lib.rs — scalar function logic
pub fn first_word(s: &str) -> &str {
    s.split_whitespace().next().unwrap_or("")
}

#[test]
fn first_word_basic() {
    assert_eq!(first_word("hello world"), "hello");
    assert_eq!(first_word("  padded  "), "padded");
    assert_eq!(first_word(""), "");
    assert_eq!(first_word("   "), "");
}
}

Unit tests for SQL macros

SqlMacro::to_sql() is pure Rust — no DuckDB connection needed:

#![allow(unused)]
fn main() {
use quack_rs::sql_macro::SqlMacro;

#[test]
fn scalar_macro_sql() {
    let m = SqlMacro::scalar("double_it", &["x"], "x * 2").unwrap();
    assert_eq!(m.to_sql(),
        r#"CREATE OR REPLACE MACRO "double_it"("x") AS (x * 2)"#);
}

#[test]
fn table_macro_sql() {
    let m = SqlMacro::table("recent", &["n"], "SELECT * FROM events LIMIT n").unwrap();
    assert_eq!(m.to_sql(),
        r#"CREATE OR REPLACE MACRO "recent"("n") AS TABLE SELECT * FROM events LIMIT n"#);
}
}

E2E testing with SQLLogicTest

Community extensions are tested using DuckDB's SQLLogicTest format, which runs SQL directly in DuckDB and compares the output line by line.

File location

test/sql/my_extension.test

Format

# my_extension tests

require my_extension

statement ok
LOAD my_extension;

query I
SELECT my_function('hello world');
----
2

Directives:

DirectiveMeaning
requireLoad the extension; skip the file if it is not available
statement okSQL must succeed
statement errorSQL must fail
query IQuery returning one INTEGER column
query IIQuery returning two INTEGER columns
query TQuery returning one TEXT column
----Expected output follows

Installing DuckDB (1.4.x or 1.5.x)

E2E testing needs the DuckDB CLI. Download it with curl; no system package manager is needed. Every 1.4.x and 1.5.x release loads an extension stamped with C API version v1.2.0 (1.4.4 through 1.5.5 declare that version; 1.5.6 declares v1.5.6 and accepts every earlier one). The quack-rs CI extension-load job loads its example extension into 1.4.4, 1.5.0, 1.5.5 and the latest release (currently 1.5.6). Develop against the current release, 1.5.6:

# DuckDB 1.5.6 (current release)
curl -fsSL https://github.com/duckdb/duckdb/releases/download/v1.5.6/duckdb_cli-linux-amd64.zip \
    -o /tmp/duckdb.zip \
    && unzip -o /tmp/duckdb.zip -d /tmp/ \
    && chmod +x /tmp/duckdb \
    && /tmp/duckdb --version
# → v1.5.6

To match a pinned CI engine instead, substitute v1.4.4, v1.5.0 or v1.5.5 in the URL.

For macOS, replace linux-amd64 with osx-universal. For Windows, use windows-amd64 and unzip to a directory on %PATH%.

Running E2E tests

# Build the extension
cargo build --release

# Append the metadata footer DuckDB's loader requires. append_metadata ships
# with quack-rs: cargo install quack-rs --bin append_metadata
append_metadata \
    target/release/libmy_extension.so \
    /tmp/my_extension.duckdb_extension \
    --abi-type C_STRUCT \
    --extension-version v0.1.0 \
    --duckdb-version v1.2.0 \
    --platform linux_amd64

# Load it in the DuckDB CLI (-unsigned allows an unsigned extension)
/tmp/duckdb -unsigned -c "
LOAD '/tmp/my_extension.duckdb_extension';
SELECT my_function('hello world');
"

The community extension CI runs these SQLLogicTest files automatically. Give each function at least one test, covering NULL, empty and typical input:

# Test NULL handling
query I
SELECT my_function(NULL);
----
NULL

# Test empty input
query I
SELECT my_function('');
----
0

# Test normal case
query I
SELECT my_function('hello world');
----
2

Pitfall P5 — SQLLogicTest does exact string matching. Copy expected values directly from DuckDB CLI output. NULL is represented as NULL (uppercase). Floats must match to the number of decimal places DuckDB outputs.


Property-based testing with proptest

The proptest crate checks a property over arbitrary inputs, which suits arithmetic and aggregate logic:

#![allow(unused)]
fn main() {
use quack_rs::interval::{interval_to_micros_saturating, DuckInterval};
use proptest::prelude::*;

proptest! {
    #[test]
    fn saturating_never_panics(months: i32, days: i32, micros: i64) {
        let iv = DuckInterval { months, days, micros };
        // Must not panic for any input
        let _ = interval_to_micros_saturating(iv);
    }
}
}

quack-rs's own test suite uses proptest for interval conversion and AggregateTestHarness properties.


What to test

ScenarioUnitE2E
NULL input → NULL output✓
Empty string✓✓
Unicode strings✓
Numeric edge cases (0, MAX, MIN)✓
Combine propagates config✓
Multi-group aggregation✓
Function registration success✓
Extension loads without crash✓
SQL macro produces correct output✓ (to_sql)✓

Dev dependencies

[dependencies]
quack-rs = "0.18"

[dev-dependencies]
proptest = "1"
# Only for InMemoryDb; enables the feature for test builds alone.
quack-rs = { version = "0.18", features = ["bundled-test"] }

The testing module is compiled unconditionally (not #[cfg(test)]), so crates that depend on quack-rs can use it in their own tests. InMemoryDb additionally needs the bundled-test or bundled-test-prebuilt feature.

Community Extensions

DuckDB's community extensions repository lets anyone publish a loadable extension that DuckDB users install with INSTALL … FROM community. This page covers scaffolding, description.yml, naming, versioning, platforms, build settings and submission for a community extension built with quack-rs.


Prerequisites

  • A working extension that passes local E2E tests
  • A GitHub repository (the community build runs from it)
  • All functions tested with SQLLogicTest format
  • A globally unique extension name

Scaffolding a new project

quack_rs::scaffold::generate_scaffold generates the project files in one call. It returns paths relative to the project root and writes nothing itself, so join them under the project directory — never write them relative to the current directory, which would overwrite its Cargo.toml and src/lib.rs:

#![allow(unused)]
fn main() {
use quack_rs::scaffold::{ScaffoldConfig, generate_scaffold};
use std::path::Path;

let config = ScaffoldConfig {
    name: "my_extension".to_string(),
    description: "Does something useful".to_string(),
    version: "0.1.0".to_string(),
    license: "MIT".to_string(),
    maintainer: "Your Name".to_string(),
    github_repo: "yourorg/duckdb-my-extension".to_string(),
    excluded_platforms: vec![],
    // `target_duckdb_version`, `use_unstable_c_api` and `git_ref` default to the
    // stable-ABI settings; see `concepts/abi.md` for when to change them.
    ..ScaffoldConfig::default()
};

let files = generate_scaffold(&config).expect("scaffold failed");
let root = Path::new(&config.name); // ./my_extension/
for file in &files {
    let path = root.join(&file.path);
    std::fs::create_dir_all(path.parent().unwrap()).unwrap();
    std::fs::write(&path, &file.content).unwrap();
}
}

In a new repository, add the build tooling submodule once with git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools (git submodule update --init does nothing until then — Pitfall P4).

This generates:

my_extension/
├── Cargo.toml
├── Makefile
├── extension_config.cmake
├── src/lib.rs
├── src/wasm_lib.rs
├── description.yml
├── test/sql/my_extension.test
├── .github/workflows/extension-ci.yml
├── .gitmodules
├── .gitignore
└── .cargo/config.toml

description.yml

The file the scaffold generates, with the fields a community submission uses:

extension:
  name: my_extension
  description: One-line description of what your extension does
  version: 0.1.0
  language: Rust
  build: cargo
  license: MIT
  requires_toolchains: rust;python3
  # excluded_platforms: "wasm_mvp;wasm_eh;wasm_threads"   # optional
  maintainers:
    - Your Name

repo:
  github: yourorg/duckdb-my-extension
  # Must be a commit hash, not a branch: the community repository builds
  # exactly this revision and signs the result, so a moving reference would
  # make the build unreproducible. DuckDB's docs: "Provide the hash of the
  # latest commit on the branch targeting stable as `ref`".
  ref: 0a11ddc058beb2d480ccbfa83e16a68400c5d076
  # ref_next: <hash>   # optional: a revision compatible with DuckDB main,
  #                    # used while a new DuckDB release is being prepared.

docs:
  hello_world: |
    SELECT my_extension_hello('world');
  extended_description: |
    A longer description, rendered on the community-extensions site.

Use quack_rs::validate to pre-validate fields before submission:

use quack_rs::validate::{
    validate_extension_name,
    validate_extension_version,
    validate_spdx_license,
    validate_excluded_platforms_str,
};

fn main() -> Result<(), quack_rs::error::ExtensionError> {
validate_extension_name("my_extension")?;
validate_extension_version("0.1.0")?;
validate_spdx_license("MIT")?;
validate_excluded_platforms_str("wasm_mvp;wasm_eh")?;
Ok(())
}

Naming rules

Extension names must satisfy all of the following:

  • Match ^[a-z][a-z0-9_]*$ (lowercase, digits, underscores — no hyphens, because DuckDB looks up the entry point as <name>_init_c_api)
  • Not exceed 64 characters
  • Be globally unique across the entire DuckDB community extensions ecosystem

Check existing names at community-extensions.duckdb.org before choosing. Use vendor-prefixed names to avoid collisions:

myorg_analytics   ✓
analytics         ✗  (likely taken or too generic)

Pitfall P1 — The [lib] name in Cargo.toml MUST exactly match the extension name. If your crate name is duckdb-my-ext (producing libduckdb_my_ext.so) but description.yml says name: my_ext, the community build fails with FileNotFoundError.


Versioning

FormatExampleMeaning
7–40 lowercase hex chars (a git hash)690bfc5Unstable — no guarantees
0.y.z0.1.0Pre-release — working toward stability
x.y.z (x > 0)1.0.0Stable — full semver guarantees

validate_extension_version is deliberately permissive: it accepts these three formats and anything else made of [A-Za-z0-9._+-], such as the date-based build ids some published extensions use. classify_extension_version accepts only the three formats above and returns the stability tier:

use quack_rs::validate::semver::{classify_extension_version, ExtensionStability};

fn main() -> Result<(), quack_rs::error::ExtensionError> {
// Returns the tier and the version string it classified.
let (stability, _version) = classify_extension_version("0.1.0")?;
match stability {
    ExtensionStability::Unstable => println!("git hash"),
    ExtensionStability::PreRelease => println!("0.y.z"),
    ExtensionStability::Stable => println!("x.y.z, x>0"),
}
Ok(())
}

Platform targets

Community extensions are built for:

PlatformDescriptionOpt-in?
linux_amd64Linux x86_64 (glibc)
linux_amd64_muslLinux x86_64 (musl)yes
linux_arm64Linux AArch64 (glibc)
linux_arm64_muslLinux AArch64 (musl)yes
osx_amd64macOS x86_64
osx_arm64macOS Apple Silicon
windows_amd64Windows x86_64
windows_amd64_mingwWindows x86_64 (MinGW)
windows_arm64Windows AArch64yes
wasm_mvpWebAssembly (MVP)
wasm_ehWebAssembly (exception handling)
wasm_threadsWebAssembly (threads)

An opt-in platform is not built unless an extension asks for it, so listing one in excluded_platforms has no effect. validate::platform::is_opt_in_platform reports which these are.

linux_amd64_gcc4 used to appear in this table and no longer exists: DuckDB retired the legacy CXX ABI target, and DuckDBPlatform() now raises a compile error rather than emitting a _gcc4 suffix. validate_platform rejects it with that explanation.

This table is derived from config/distribution_matrix.json in duckdb/extension-ci-tools, and scripts/check-platform-table.py fails CI when quack-rs's copy drifts from it.

If your extension cannot be built for a platform (e.g., it uses a platform-specific system library), add it to excluded_platforms:

#![allow(unused)]
fn main() {
use quack_rs::scaffold::ScaffoldConfig;

let config = ScaffoldConfig {
    excluded_platforms: vec![
        "wasm_mvp".to_string(),
        "wasm_eh".to_string(),
        "wasm_threads".to_string(),
    ],
    ..ScaffoldConfig::default()
};
}

Validate individual platform names with validate_platform:

#![allow(unused)]
fn main() {
use quack_rs::validate::validate_platform;
assert!(validate_platform("linux_amd64").is_ok());
assert!(validate_platform("invalid").is_err());
}

Cargo.toml requirements

[package]
name = "my_extension"
version = "0.1.0"
edition = "2021"

[lib]
name = "my_extension"       # Must match description.yml `name`
crate-type = ["cdylib"]

[dependencies]
quack-rs = "0.18"
libduckdb-sys = { version = ">=1.4.4, <2", features = ["loadable-extension"] }

[profile.release]
panic = "unwind"             # Required — "abort" disables quack-rs's panic guards
opt-level = 3
lto = true
codegen-units = 1
strip = true

This matches the scaffold's Cargo.toml (which also declares a staticlib example target for WebAssembly builds).

ADR-4 (in LESSONS.md) — Do NOT use the duckdb crate's bundled feature. A loadable extension must call into the DuckDB that loads it, not bundle its own copy. libduckdb-sys with loadable-extension provides function pointers that are filled in at load time from the API struct DuckDB passes in.


Release profile check

validate_release_profile checks the four release-profile settings. Only panic = "unwind" is required; lto = true, opt-level = 3 and codegen-units = 1 are recommended, and the returned ReleaseProfileCheck reports each one:

#![allow(unused)]
fn main() {
use quack_rs::validate::validate_release_profile;

// Pass all four release profile settings from your Cargo.toml
assert!(validate_release_profile("unwind", "true", "3", "1").is_ok());
// Err — see below
assert!(validate_release_profile("abort", "true", "3", "1").is_err());
}

panic must be "unwind". quack-rs wraps every extern "C" entry point in catch_unwind so a panic in your code becomes a DuckDB error rather than a crash, and catch_unwind cannot catch anything under panic = "abort": the runtime aborts before unwinding starts, killing the user's DuckDB session.

The older advice to set abort came from panics escaping an extern "C" boundary once being undefined behavior. They no longer are — Rust defines that as an abort — and quack-rs catches them before the boundary anyway.


CI workflow

The scaffold generates .github/workflows/extension-ci.yml, which:

  1. Runs on pushes and pull requests to main
  2. Runs cargo fmt --check and cargo clippy on Linux, and cargo test on Linux, macOS and Windows
  3. Runs make configure and make release, which use extension-ci-tools to build the .duckdb_extension file
  4. Runs the SQLLogicTests in test/sql with make test

After scaffolding:

cd my_extension
git init
git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools
git submodule update --init --recursive
make configure
make release

Pitfall P4 — The extension-ci-tools submodule must be initialized. make configure fails if the submodule is missing.


Submitting to the community registry

  1. Create a pull request against the community-extensions repository
  2. Add your description.yml under extensions/my_extension/description.yml
  3. CI runs automatically to verify the build
  4. Once approved, users can install your extension:
INSTALL my_extension FROM community;
LOAD my_extension;

Binary compatibility

An extension that uses only the stable C API (the default quack-rs features) is stamped C_STRUCT with C API version v1.2.0, and one binary loads into every DuckDB 1.4.x and 1.5.x release for its platform. The libduckdb-sys = ">=1.4.4, <2" range above is correct for such an extension.

An extension that enables the duckdb-1-5* features should be stamped C_STRUCT_UNSTABLE with the exact DuckDB release it was built against, so that DuckDB loads it only into that release. (Stamped C_STRUCT instead, it loads anywhere and quack-rs's runtime layout check refuses a mismatched release at LOAD.) Pin libduckdb-sys to that release's bindings (~1.10505.0 for DuckDB 1.5.5) and rebuild for each new DuckDB release. See ABI Compatibility.

The community build pipeline rebuilds extensions for each DuckDB release.

Pitfall P2 — For C_STRUCT, the -dv flag to append_extension_metadata.py must be the C API version (v1.2.0), not the DuckDB release version (v1.4.4). Use quack_rs::DUCKDB_API_VERSION to avoid hardcoding this.


Security considerations

The DuckDB team does not audit community extensions for security, so the responsibility is yours:

  • Never let a panic escape an FFI boundary: quack-rs's callbacks catch panics and report them as SQL errors, which requires panic = "unwind"
  • Validate user inputs at system boundaries (extension entry point is the boundary)
  • Do not include secrets, API keys, or credentials in your binary
  • Dynamic SQL in SQL macros must not construct queries from unsanitized user data

Pitfall Catalog

The 31 known pitfalls of writing a DuckDB extension in Rust against the C extension API, with the symptom, root cause and fix for each. The first were found while building duckdb-behavioral, a production DuckDB community extension; the rest while building and auditing quack-rs. Most of them affect any Rust extension that calls the C API directly, and quack-rs prevents most of them. The summary at the end lists each one with its status.


L1: COMBINE must propagate ALL config fields

Status: Testable with AggregateTestHarness.

Symptom: Aggregate function returns wrong results. No error, no crash.

Root cause: DuckDB's segment tree creates fresh target states, initialised by state_init (with FfiState<T>, a T::default()), then calls combine to merge source states into them. If your combine only propagates data fields (count, sum) but omits configuration fields (window_size, mode), the configuration is still its state_init default at finalize time, silently corrupting results.

This bug passed 435 unit tests before being caught by E2E tests.

Fix:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { window_size: i64, mode: u8, count: i64 }
impl AggregateState for MyState {}
unsafe extern "C" fn combine(
    _info: duckdb_function_info,
    source: *mut duckdb_aggregate_state,
    target: *mut duckdb_aggregate_state,
    count: idx_t,
) {
    for i in 0..count as usize {
        let src_ptr = unsafe { *source.add(i) };
        let tgt_ptr = unsafe { *target.add(i) };
        if let (Some(src), Some(tgt)) = (
            FfiState::<MyState>::with_state(src_ptr),
            FfiState::<MyState>::with_state_mut(tgt_ptr),
        ) {
            tgt.window_size = src.window_size;  // config — MUST copy
            tgt.mode = src.mode;                // config — MUST copy
            tgt.count += src.count;             // data — accumulate
        }
    }
}
}

Test this with AggregateTestHarness::combine — see Testing Guide.


L2: State destroy double-free

Status: Made impossible by FfiState<T>.

Symptom: Crash or memory corruption on extension unload.

Root cause: If state_destroy frees the inner Box but does not null the pointer, a second state_destroy call (common in error paths) frees already-freed memory → undefined behavior.

Fix: FfiState<T>::destroy_callback clears the slot's tag before dropping the T, and drops only a slot whose tag matches, so a second call is a no-op. Use it instead of writing your own destructor:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { window_size: i64, mode: u8, count: i64 }
impl AggregateState for MyState {}
unsafe extern "C" fn state_destroy(states: *mut duckdb_aggregate_state, count: idx_t) {
    unsafe { FfiState::<MyState>::destroy_callback(states, count) };
}
}

You rarely need even this wrapper: .ffi_state::<MyState>() on the aggregate builder installs destroy_callback together with FfiState<MyState>'s size and init callbacks.


L3: No panic across FFI boundaries

Status: Made impossible by init_extension and the callback guards (which require panic = "unwind").

Symptom: The whole DuckDB process aborts when an extension callback panics.

Root cause: a panic cannot unwind out of an extern "C" function. Since Rust 1.81 the runtime aborts the process when one tries (before 1.81 it was undefined behaviour), so an uncaught panic!() or .unwrap() in a callback takes down the user's whole DuckDB session.

Fix: Use Result and ? inside init_extension. Never use unwrap() in FFI callbacks. FfiState::with_state_mut returns Option, not Result, so callers use if let:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default)] struct MyState { window_size: i64, mode: u8, count: i64 }
impl AggregateState for MyState {}
unsafe fn demo(state_ptr: duckdb_aggregate_state) {
// Safe pattern — no unwrap in FFI callback
if let Some(st) = unsafe { FfiState::<MyState>::with_state_mut(state_ptr) } {
    st.count += 1;
}

// Dangerous — never do this in an FFI callback
let st = unsafe { FfiState::<MyState>::with_state_mut(state_ptr) }.unwrap(); // panics if None
}
}

quack-rs's callback macros and typed builders catch a panic and report it as a SQL error. That requires panic = "unwind" in the release profile, which is what the scaffold generates and what validate_release_profile insists on: under panic = "abort" nothing can be caught.


L4: ensure_validity_writable is required before NULL output

Status: Made impossible by VectorWriter::set_null.

Symptom: NULLs you write are silently lost — the row reads back as a valid value (whatever is in the data buffer).

Root cause: a vector that has never held a NULL usually has no validity mask at all, and duckdb_vector_get_validity then returns NULL (as duckdb.h documents). duckdb_validity_set_row_invalid returns early on a NULL mask, so nothing is written and nothing crashes. duckdb_vector_ensure_validity_writable allocates the mask, after which get_validity returns it. (Dereferencing the NULL pointer yourself, instead of going through the C API helpers, would crash.)

Fix: Always call duckdb_vector_ensure_validity_writable before accessing the validity bitmap on the write path. VectorWriter::set_null does this automatically:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe fn demo(writer: &mut VectorWriter, row: usize) {
// Correct — handled by set_null
unsafe { writer.set_null(row) };

// Wrong — validity bitmap may not be allocated yet
// let validity = duckdb_vector_get_validity(output);          // NULL
// duckdb_validity_set_row_invalid(validity, row);            // silently ignored
}
}

For STRUCT and ARRAY outputs set_null also nulls the children at that row, as DuckDB's internal FlatVector::SetNull does; a bare duckdb_validity_set_row_invalid on the parent leaves the fields valid, and struct_extract on the NULL row returns their stale values.


L5: Boolean reading must use u8 != 0, not *const bool

Status: Made impossible by VectorReader::read_bool.

Symptom: Undefined behavior; Rust requires bool to be exactly 0 or 1.

Root cause: DuckDB's C API does not guarantee that boolean values in vectors are exactly 0 or 1. Values of 2, 255, etc. cast to Rust bool is undefined behavior.

Fix: Read as u8 and compare with != 0. VectorReader::read_bool always does this:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe fn demo(reader: &VectorReader, row: usize) {
let b: bool = unsafe { reader.read_bool(row) };  // safe: uses u8 != 0 internally
}
}

L6: Function set name must be set on EACH member

Status: Made impossible by AggregateFunctionSetBuilder.

Symptom: Functions are silently not registered. No error returned.

Root cause: When using duckdb_register_aggregate_function_set, the function name must be set on EACH individual duckdb_aggregate_function using duckdb_aggregate_function_set_name, not just on the set.

This is completely undocumented. Discovered by reading DuckDB's C++ test code at test/api/capi/test_capi_aggregate_functions.cpp.

In duckdb-behavioral, 6 of 7 functions failed to register silently due to this bug.

Fix: AggregateFunctionSetBuilder calls duckdb_aggregate_function_set_name on every individual function before adding it to the set. Use it instead of managing the set manually.


L7: LogicalType memory leak

Status: Made impossible by LogicalType RAII wrapper.

Symptom: Memory leak proportional to number of registered functions.

Root cause: duckdb_create_logical_type allocates memory that must be freed with duckdb_destroy_logical_type. Forgetting leaks memory.

Fix: LogicalType implements Drop and calls duckdb_destroy_logical_type automatically when it goes out of scope.


L8: DEFAULT_NULL_HANDLING does not propagate NULLs for scalar functions

Status: Made impossible by ScalarFunctionBuilder::map1 / map2 / map1_str / map2_str. DataChunk::propagate_nulls fixes it in one line for hand-written callbacks.

Symptom: A scalar function returns a value where SQL requires NULL — but only for arguments that come from a column. SELECT f(NULL) looks correct, because a literal NULL is constant-folded before the function is reached, so the bug survives review and ships.

Root cause: The name suggests DuckDB returns NULL on your behalf. For a scalar function registered through the C API it does not, at run time: CAPIScalarFunction calls the callback for every row including NULL ones and checks only the error flag, and the one NULL check in ExpressionExecutor — VerifyNullHandling — has its entire body inside #ifdef DEBUG. Every DuckDB a user installs is a release build.

Fix: use the typed closure constructors, which skip NULL rows and write NULL for them, or call DataChunk::propagate_nulls(&mut writer) at the end of a hand-written callback. map1_opt / map2_opt and NullHandling::SpecialNullHandling are for functions that genuinely mean to see NULLs. Aggregates are no different: update receives NULL rows under either setting too, so check is_valid before reading (see L12).


L9: duckdb_data_chunk_from_arrow takes the array even when it fails

Status: Made impossible by arrow::data_chunk_from_arrow, which takes the ArrowArray by value.

Symptom: One of two opposite bugs, depending on which way you guessed. Treat the array as still yours after a failed conversion and you double-release it. Treat it as gone in every case and a zero-column conversion leaks the whole Arrow buffer tree.

Root cause: duckdb.h says "Data ownership is passed on to DuckDB's DataChunk", which reads like a success-path statement. arrow-c.cpp sets arrow_array->release = nullptr inside the per-column loop, before the work that can throw — so the array is claimed on the error path too, but only if the loop runs at all. A zero-column converted schema leaves release intact and the array still belongs to the caller.

Fix: own the record in a wrapper whose Drop releases only if release survived, and consume it by value. The by-value binding drops on the way out: a no-op when DuckDB nulled release, a correct release when it did not. The mirror case is handled by the same rule — ToArrowSchema / ToArrowArray install release last, so a failed export leaves nothing to free.


L10: Scalar bind data is dropped when DuckDB copies the expression

Status: Fixable only from the extension, and now possible: ScalarBindInfo::set_bind_data_copy.

Symptom: A scalar function that allocates per-query state in its bind callback reads null from duckdb_scalar_function_get_bind_data during execution, for some queries and not others. Nothing crashes and nothing is reported: the callback simply runs without the state it bound, so the answer is quietly wrong.

Root cause: duckdb_scalar_function_set_bind_data registers the pointer and its destructor, but not how to duplicate it. DuckDB copies a bound expression whenever it duplicates a plan, and CScalarFunctionBindData::Copy() in src/main/capi/scalar_function-c.cpp (read at v1.5.5; byte-identical in v1.5.4) only fills the copy in when a copy callback exists:

unique_ptr<FunctionData> Copy() const override {
    auto copy = make_uniq<CScalarFunctionBindData>(info);
    if (copy_callback) {
        copy->bind_data = copy_callback(bind_data);
        copy->delete_callback = delete_callback;
        copy->copy_callback = copy_callback;
    }
    return std::move(copy);   // bind_data stays null without a callback
}

With no callback the copy carries bind_data = nullptr, and the original is untouched — which is why the failure is intermittent rather than total, and why it survives a test suite that only ever executes the first-bound expression.

Fix: use ScalarBindData::set, which registers a generated, panic-safe copy callback (it requires T: Clone + Send + Sync). With the raw API, register a copy callback alongside the bind data, in the same bind callback and after set_bind_data:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
use std::os::raw::c_void;
use quack_rs::scalar::ScalarBindInfo;
#[derive(Clone)] struct MyBindData;
unsafe extern "C" fn destroy(p: *mut c_void) { drop(unsafe { Box::from_raw(p.cast::<MyBindData>()) }); }
unsafe fn demo(bind_info: ScalarBindInfo, boxed: Box<MyBindData>) {
unsafe extern "C" fn copy(data: *mut c_void) -> *mut c_void {
    if data.is_null() {
        return std::ptr::null_mut();
    }
    let src = unsafe { &*data.cast::<MyBindData>() };
    Box::into_raw(Box::new(src.clone())).cast()
}

unsafe {
    bind_info.set_bind_data(Box::into_raw(boxed).cast(), Some(destroy));
    bind_info.set_bind_data_copy(Some(copy));
}
}
}

The duplicate is freed with the same destructor as the original, so copy must return an independently owned allocation — returning the pointer it was given is a double free.

Related: the copy callback runs across the FFI boundary like any other, so it must not unwind. Wrap anything that can panic in callback::catch_ffi_panic and return null.


L11: C API aggregates crash under agg(x) OVER () and agg(x ORDER BY y)

Status: A DuckDB defect, reported upstream as duckdb/duckdb#26109. Cannot be prevented or detected from an extension; documented on AggregateFunctionBuilder, AggregateFunctionSetBuilder and FfiState.

Symptom: An aggregate that works under SELECT agg(x) FROM t and GROUP BY segfaults (or corrupts memory, or returns a wrong answer) when used as a window over a whole-partition frame — agg(x) OVER (), OVER (PARTITION BY p) — or as an ordered aggregate, agg(x ORDER BY y).

Root cause: CAPIAggregateUpdate (src/main/capi/aggregate_function-c.cpp) flattens the input vectors but not the state vector, then hands the callback FlatVector::GetDataUnsafe(state). The C API registers no simple_update, so two executors fall back to calling update with a constant state vector and count > 1: WindowConstantAggregatorLocalState (statep(Value::POINTER(0))) and SortedAggregateFunction (agg_state_vec.SetVectorType(CONSTANT_VECTOR)). The callback reads states[i] for every row, as the C API contract says it may; only states[0] exists. Reproduced with a plain C aggregate (no quack-rs) against DuckDB 1.4.4, 1.5.0 and 1.5.5; AddressSanitizer places the fault in the callback, called from CAPIAggregateUpdate.

Fix: none on the extension side — the callback receives a raw duckdb_aggregate_state * and cannot tell a constant vector from a flat one, and reading states[1] to find out is itself the out-of-bounds read. Until DuckDB fixes it, document for your users that the aggregate must not be used in those two query shapes. Frames that are not whole-partition (ROWS BETWEEN 5 PRECEDING AND CURRENT ROW, segment-tree windows) and DISTINCT windows were checked and work.


L12: Aggregate update receives NULL rows under DEFAULT_NULL_HANDLING

Status: Documented on NullHandling, UpdateFn and both aggregate builders' null_handling. Pinned by aggregate_update_receives_null_rows_under_either_null_handling in tests/ffi_roundtrip/lifecycle.rs.

Symptom: An aggregate that reads every row — state.sum += reader.read_i64(row) — returns a wrong answer, with no error, as soon as its input column contains a NULL. The value read for a NULL row is whatever the data buffer happens to hold.

Root cause: quack-rs used to document (in NullHandling, the builders and this book) that DuckDB's aggregate executor filters NULL rows out before update unless SpecialNullHandling is set. It does not. CAPIAggregateUpdate (src/main/capi/aggregate_function-c.cpp) flattens each input vector and passes the whole chunk, validity and all. For an aggregate the setting is read in one place, BoundAggregateExpression::PropagatesNullValues, which only the correlated-subquery decorrelator (flatten_dependent_join.cpp) consults to pick an INNER or LEFT join; the aggregate VerifyNullHandling check is compiled only under #ifdef DEBUG. Checked against DuckDB 1.5.5: update saw every NULL row, ungrouped and under GROUP BY, under both settings, and no correlated subquery tried answered differently under the two.

Fix: in update, skip rows where VectorReader::is_valid(row) is false, whatever the null handling. Use SpecialNullHandling to declare that the aggregate returns non-NULL for NULL input; it does not change which rows arrive.

A related trap under either setting: in a correlated subquery, (SELECT my_count(x) FROM t2 WHERE t2.k = t1.k) is NULL, not my_count of an empty input, for an outer row with no match. DuckDB rewrites that NULL to 0 only for its own count. Wrap the subquery in coalesce(..., 0) if it matters.


L13: A C API aggregate without a destructor is wrong in a running window

Status: Fixed in quack-rs: every aggregate builder registers a destructor, a no-op when none is given. Pinned by an_aggregate_without_a_destructor_is_right_in_a_running_window in tests/ffi_roundtrip/agg_window.rs. Reported in docs/upstream-duckdb-reports.md, item 7.

Symptom: agg(x) OVER (ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW) with no PARTITION BY or ORDER BY returns the wrong running value, with no error: a sum over 1..5 reads 1 2 3 4 5. The same aggregate is right with ORDER BY, ungrouped, grouped, and in DuckDB's own sum.

Root cause: DuckDB streams such a window only for an aggregate with no destructor (PhysicalStreamingWindow::IsStreamingFunction). Streaming calls update once per row, count 1, on a one-row dictionary slice it moves along, and CAPIAggregateUpdate flattens that input vector in place, so the slice becomes row 0's value for the rest of the chunk. Checked against 1.4.4, 1.5.0 and 1.5.5.

Fix: register a destructor, even an empty one. quack-rs does this for you; with the raw C API, call duckdb_aggregate_function_set_destructor.


L14: The C API behaves differently across the releases one build loads into

Status: Four cases fixed in quack-rs (ScalarBindInfo::argument, LogicalType::try_decimal, LogicalType::try_new, the scalar collision check); CI job test-older-engines runs the suite against DuckDB 1.4.4 and 1.4.5 (default features), 1.5.0 (duckdb-1-5), 1.5.3 (duckdb-1-5-3) and 1.5.4 (duckdb-1-5-4).

Symptom: code tested against the release Cargo.lock pins (1.5.5) aborts or misbehaves in an older release the same binary loads into. A default-feature extension loads into every release from 1.4.4; a duckdb-1-5 one built against the 1.5.4 bindings has the 546-slot layout of 1.5.2 to 1.5.6, so the ABI guard rightly lets it load into all five.

Root cause: a C function's contract can change in a release while its slot stays put. duckdb_scalar_function_bind_get_argument gained its try only in 1.5.5 (before it, a subquery argument throws through the extension's callback: an abort in Rust); duckdb_create_decimal_type its width/scale check only in 1.5.4 (before it, DECIMAL(0, 0) comes back as a type); duckdb_register_scalar_function its ALTER_ON_CONFLICT only in 1.5.0 (before it, no existing name can take another overload); and 1.4.x's C API reports TIME_NS, which its SQL has, as INVALID. All four were found only by running the whole suite against the oldest release of each range.

Fix: when a wrapper relies on C API behaviour, find the release that introduced it (the source at each tag answers that), and either check the engine version at run time (abi::engine_version) or make the check in Rust. Test against the oldest release a build can load into, not only the pinned one.


L15: combine must leave its source states unchanged

Status: Documented on CombineFn and in the aggregate chapter; pinned by combine_must_leave_its_source_unchanged (tests/ffi_roundtrip/agg_window.rs). Not preventable by the SDK: combine receives raw state pointers.

Symptom: an aggregate is right in GROUP BY and wrong in a sliding window (ROWS BETWEEN n PRECEDING AND CURRENT ROW): in the regression test a sum whose combine moved its value out of the source gave 4985 of 5000 rows wrong.

Root cause: a window's segment tree keeps one state per tree node and combines each node's state into every frame that covers it, from several threads at once (WindowSegmentTreePart::WindowSegmentValue, window_segment_tree.cpp). A combine that consumes its source (mem::take, zeroing a counter) is right for the first frame and wrong for the rest, and a write to a source another thread reads is a data race.

Fix: read the source; copy or clone what the target needs. AggregateState requires Sync for the same reason.


L16: A valid Arrow array is not always one DuckDB imports correctly

Status: Fixed in quack-rs: data_chunk_from_arrow walks the array with its schema and refuses the layouts DuckDB 1.4.4 to 1.5.5 mishandles (src/arrow/import_layout.rs, tests/ffi_roundtrip/arrow_layout.rs; docs/upstream-duckdb-reports.md, items 9, 24 to 29 and 31 to 33).

Symptom: an array that arrow-rs or another producer built, valid by the Arrow specification, imports with values from the wrong rows, reads past a buffer, or corrupts the heap. Arrays DuckDB exported itself never show it, which is why round-trip tests pass.

Root cause: DuckDB's importer tracks where a node's rows start with two parameters, parent_offset and nested_offset, and some paths pass the wrong one: a struct gives its children only its own offset, union members start at row 0, a dictionary's validity ignores the list's offset. List views, recoded union type ids and nested dictionaries are mishandled too.

Fix: never assume a producer's layout matches the one DuckDB writes. Test an importer with hand-built arrays that put offsets at every level, and refuse what the engine cannot import rather than return wrong values.


L17: A COPY … FROM reader must not declare result columns

Status: Refused for typed table functions (their bind fails with a message); documented for raw ones on CopyFunctionBuilder::copy_from and BindInfo::add_result_column. Pinned by tests/ffi_roundtrip/copy_from_columns.rs; docs/upstream-duckdb-reports.md, item 37.

Symptom: on a DuckDB built with assertions, COPY t FROM … fails with chunk.ColumnCount() == types.size() and the database is invalidated. A release build silently drops the extra column, so the bug hides in testing.

Root cause: CCopyFromBind hands the reader's bind the INSERT's own list of expected types as its result types, and duckdb_bind_add_result_column appends to that list, so every chunk the INSERT receives is wider than the table. duckdb.h says the reader "should not" declare columns; nothing enforces it.

Fix: in a COPY … FROM reader's bind, read the target's columns with BindInfo::result_column_count and its siblings, and declare none.


L18: A LIST reserve moves every buffer below its child

Status: Documented in the # Safety sections of VectorWriter::from_vector, StructWriter::new, StructVector::field_writer and ValidityBitmap::ensure_writable; measured by tests/ffi_roundtrip/nested_reserve.rs.

Symptom: a writer on a STRUCT field of a list's elements writes into freed memory after the list is grown, although it was never a direct child of the list.

Root cause: duckdb_list_vector_reserve resizes the child with Vector::Resize, which reallocates the data and validity buffers of the child and of every STRUCT field and ARRAY element vector below it, down to the next LIST (whose child has its own buffer). Writers cache both pointers.

Fix: fetch every writer and bitmap below a list's child again after each reserve on that list (a ListBuilder row that grows it counts).


L19: Addresses of constants are not identities

Status: Fixed in FfiState (0.18.0): its per-type tag salt is a hash of TypeId::of::<T>(), not the address of type_name::<T>(). AUDIT.md 10.2 (High) and 10.8.

Symptom: a per-type tag compared across callbacks mismatches only in release builds of the user's crate — FfiState::with_state returns None, destroy_callback skips states — while debug and fat-LTO builds hide it.

Root cause: core::any::type_name::<T>().as_ptr() (and a function pointer, and the address of any &'static constant) can differ between codegen units: rustc emits a private copy of a constant in each codegen unit that uses it, so with codegen-units > 1 and no fat LTO (Cargo's default release profile) two uses of the same constant can have different addresses. Rust makes no address-identity guarantee for functions or constants.

Fix: derive identity from a value, e.g. hash TypeId::of::<T>().

Evidence: a standalone crate taking type_name::<T>().as_ptr() for one T in 8 modules: release profile (codegen-units = 16, lto = false) → 6 distinct addresses; dev profile → 1; codegen-units = 1, lto = true → 1. In quack-rs itself, cargo test --release --lib aggregate:: with CARGO_PROFILE_RELEASE_CODEGEN_UNITS=16 CARGO_PROFILE_RELEASE_LTO=false failed 5 tests on the address salt and passes on the TypeId one; the end-to-end suite under the same profile failed 10 of 279 (every aggregate: NULL or garbage results) and passes all 279. Every other CI build is debug or uses the repository's codegen-units = 1, fat-LTO release profile, which is why the bug got past them; CI's test job now runs this build.


P1: Library name must match extension name

Status: Must be configured in Cargo.toml. Scaffold handles this.

Symptom: Community build fails with FileNotFoundError.

Root cause: The community build expects lib{extension_name}.so. If the Cargo crate name produces a different .so filename, the build fails.

Fix: Set name explicitly in [lib]:

[lib]
name = "my_extension"   # Must match description.yml `name: my_extension`
crate-type = ["cdylib", "rlib"]

P2: Metadata version is C API version, not DuckDB version

Status: The DUCKDB_API_VERSION constant holds the correct value.

Symptom: The metadata script succeeds, and LOAD then refuses the file: "The file was built for DuckDB C API version 'v1.5.5', but we can only load extensions built for DuckDB C API 'v1.2.0' and lower" (verified on DuckDB 1.4.4 and 1.5.5 with a file stamped -dv v1.5.5).

Root cause: The -dv flag to append_extension_metadata.py must be the C API version (v1.2.0), not the DuckDB release version (v1.4.4). These are different strings. DuckDB 1.4.x and 1.5.0 – 1.5.5 declare C API version v1.2.0; 1.5.6 declares v1.5.6 and still loads v1.2.0 extensions.

Fix: Use quack_rs::DUCKDB_API_VERSION ("v1.2.0") in init_extension, and use the same version with append_extension_metadata.py -dv v1.2.0.

This holds only for the C_STRUCT ABI type. For C_STRUCT_UNSTABLE and CPP, -dv is the exact DuckDB release: with USE_UNSTABLE_C_API=1 (required when you use quack-rs's duckdb-1-5 features; see P10), TARGET_DUCKDB_VERSION must be a real release such as v1.5.6, and v1.2.0 would pin the binary to DuckDB v1.2.0. ScaffoldConfig validates this pairing.


P3: E2E testing is mandatory

Status: Documented. See Testing Guide.

Symptom: All unit tests pass but the extension is completely broken.

Root cause: Unit tests cannot detect SEGFAULTs on load, silent registration failures, or wrong results from combine bugs.

Fix: Always run E2E tests using an actual DuckDB binary. The scaffold generates a complete SQLLogicTest skeleton.


P4: extension-ci-tools submodule must be initialized

Status: Build-time check.

Symptom: make configure or make release fails.

Fix: In a new project (for example one fresh from the scaffold) the submodule has never been added: the scaffold writes .gitmodules, but a file cannot create the gitlink git needs, so git submodule update --init finds nothing to do and exits 0 without cloning anything. Add it once:

git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools

In a clone of a repository that already has the submodule:

git submodule update --init --recursive

The generated Makefile checks for the checkout before it includes anything from it and prints both commands if it is missing.


P5: SQLLogicTest expected values must match exactly

Status: Test-authoring care required.

Symptom: Tests fail in CI but pass locally (or vice versa).

Root cause: SQLLogicTest does exact string matching. Output format (decimal places, NULL representation, column separators) must match character-for-character.

Fix: Generate expected values by running the SQL in DuckDB CLI and copying the output. NULL is NULL (uppercase). Integers have no decimal places.


P6: duckdb_register_aggregate_function_set silently fails

Status: Builder returns Err. Also see L6.

Symptom: Function appears registered but is not found in SQL.

Root cause: The return value of duckdb_register_aggregate_function_set is often ignored. When it returns DuckDBError, the function set is not registered.

Fix: The builder checks the return value and propagates it as Err.


P7: duckdb_string_t format is undocumented

Status: Handled by VectorReader::read_str, VectorReader::read_blob, and DuckStringView.

Symptom: VARCHAR reading produces garbage, empty strings, or crashes; BLOB reading silently drops bytes that are not valid UTF-8.

Root cause: DuckDB stores strings in a 16-byte struct with two formats (inline ≤ 12 bytes, pointer > 12 bytes) that are not documented in libduckdb-sys. The length and the pointer are in the target's own byte order, so a decoder that reads them as little-endian misreads every string on a big-endian target.

Fix: Use VectorReader::read_str(row) for UTF-8 text and VectorReader::read_blob(row) for arbitrary binary data. See NULL Handling & Strings.


P8: INTERVAL struct layout is undocumented

Status: Handled by DuckInterval and read_interval_at.

Symptom: Interval calculations produce wrong results or crashes.

Root cause: DuckDB's INTERVAL is { months: i32, days: i32, micros: i64 } (16 bytes total). This is not documented in libduckdb-sys. Month conversion uses 1 month = 30 days (DuckDB's approximation).

Fix: Use VectorReader::read_interval(row) and DuckInterval. See INTERVAL Type.


P9: loadable-extension dispatch table uninitialised in cargo test

Status: Fixed. InMemoryDb::open() initialises the dispatch table automatically.

Symptom: All three InMemoryDb unit tests panic at runtime:

thread 'testing::in_memory_db::tests::in_memory_db_opens' panicked at
'DuckDB API not initialized or DuckDB feature omitted'

This failure appears only when running cargo test --features bundled-test. Regular cargo test (no feature) does not exercise this code path, so CI can miss it entirely.

Root cause: Cargo's feature-unification merges loadable-extension (from the main libduckdb-sys dependency) and bundled (pulled in by the duckdb crate's features = ["bundled"]) into a single libduckdb-sys build with both features active. In loadable-extension mode every DuckDB C API call is routed through a dispatch table of one AtomicPtr per function, which is normally populated at load time, when DuckDB calls the extension's entry point and the entry point calls duckdb_rs_extension_api_init. In cargo test, no DuckDB host process loads the extension, so the table stays uninitialised and every call panics.

Discovery: This was triggered by the crates.io release workflow (which runs cargo test --all-targets --all-features) failing on macOS. Regular CI at the time (cargo test --all-targets, no --all-features) never compiled the bundled-test path, so the bug was hidden during development and code review.

Fix (implemented in quack-rs 0.6.0):

  1. src/testing/bundled_api_init.cpp — a thin C++ shim that wraps DuckDB's internal CreateAPIv1() (from duckdb/main/capi/extension_api.hpp) as a C-linkage symbol:

    #include "duckdb/main/capi/extension_api.hpp"
    extern "C" duckdb_ext_api_v1 quack_rs_create_api_v1() {
        return CreateAPIv1();
    }
    
  2. build.rs — compiles the shim (via the cc crate) only when the bundled-test or bundled-test-prebuilt feature is active. It finds the DuckDB headers through DEP_DUCKDB_INCLUDE (published by libduckdb-sys >= 1.10503), falling back to the libduckdb-sys build output directory or, for a prebuilt library, DUCKDB_INCLUDE_DIR.

  3. InMemoryDb::open() — calls init_dispatch_table_once() before opening the connection. That function calls quack_rs_create_api_v1() once and feeds the result through duckdb_rs_extension_api_init, populating every AtomicPtr slot in the dispatch table, one per field of duckdb_ext_api_v1 (546 with the 1.5.2 – 1.5.6 bindings, 459 with 1.4.x; see the table in ABI Compatibility). A std::sync::Once guard makes it safe to call from any number of threads and test cases.

  4. CI test-bundled job — runs cargo test --all-targets --features bundled-test and then the release workflow's own cargo test --all-targets --all-features on Linux, macOS and Windows on every PR. The second step was added after the v0.18.0 tag failed on Windows while PR CI was green: until then no PR job ran the duckdb-1-5* tests on macOS or Windows.

ABI compatibility note: DuckDB's duckdb_ext_api_v1 struct is defined identically in both the public duckdb_extension.h (used by libduckdb-sys bindgen) and the internal extension_api.hpp (used by CreateAPIv1()). Both include the DUCKDB_EXTENSION_API_VERSION_UNSTABLE fields. CreateAPIv1() sets every field. The Rust and C++ structs are produced from the same DuckDB release and therefore stay in sync.

Risk table (using DuckDB's internal C++ API):

RiskMitigation
extension_api.hpp is renamed or movedbuild.rs fails with a clear compile error
CreateAPIv1() is renamedSame — C++ compile error
duckdb_ext_api_v1 gains new fieldsCreateAPIv1() fills new fields too
duckdb_ext_api_v1 field order changesBoth structs from same DuckDB release, stay in sync
libduckdb-sys drops loadable-extension dispatchProblem disappears; Once guard becomes cheap no-op

P10: The C API struct has a stable prefix and an unstable tail

Status: Detected at load time by quack_rs::abi (default AbiPolicy::Strict).

Symptom: An extension loads without complaint and then corrupts memory — double free or corruption, a segfault, or silently wrong results — on a DuckDB release other than the one it was built against. Nothing in the build or the load warns you.

Root cause: DuckDB hands a loadable extension a pointer to a duckdb_ext_api_v1 struct of function pointers, and the extension calls through it at compiled-in offsets. The struct has two regions:

RegionSlotsGuarantee
Stable0–356Frozen since v1.2.0 — same slots, order and signatures in every release through v1.5.6 (two slots, 114 and 138, were renamed varint → bignum in v1.4.0 with an identical struct layout)
Unstable357+DuckDB inserts entries in the middle, shifting every later slot

duckdb_appender_clear landed at slot 410 in v1.5.0 and duckdb_geometry_type_get_crs in the middle of v1.5.2's tail; each insertion moves everything after it. An extension compiled against one layout and loaded by another calls the wrong function through the right offset.

Your action: use init_extension, which verifies the layout and refuses a mismatch. If you enable a duckdb-1-5* feature, also stamp the binary C_STRUCT_UNSTABLE with the exact DuckDB release (USE_UNSTABLE_C_API=1 and a real TARGET_DUCKDB_VERSION, or append_metadata --abi-type C_STRUCT_UNSTABLE --duckdb-version vX.Y.Z), so that DuckDB itself refuses to load it into any other release. Two knobs matter:

  • QUACK_RS_TARGET_DUCKDB_VERSION at build time stamps the release you built against, so a DuckDB newer than quack-rs's table is still accepted when your build genuinely targeted it. The community-extension CI rebuilds per release, so this is the normal path.
  • AbiPolicy::Warn or Trust if you would rather load anyway. Trust is the old behaviour, and the failure mode above is the reason it is no longer the default.

If your extension enables no duckdb-1-5* feature it only calls into the stable prefix: the check reports AbiCheck::StableOnly, and the extension loads into every release from v1.2.0 on.


P11: const char * returns are borrowed — freeing one corrupts the heap

Status: Fixed in quack-rs; documented here because extension authors calling the C API directly hit the same trap.

Symptom: corrupted size vs. prev_size in fastbins, free(): invalid pointer, or a SIGABRT at an unrelated later allocation. Nothing points at the call that caused it.

Root cause: The C API returns strings two ways, and only one transfers ownership.

Return typeTypical implementationCaller must
char *strdup(...) or duckdb_malloc + memcpyduckdb_free it
const char *some_std_string.c_str()not free it

duckdb_copy_function_global_init_get_file_path is the second kind: it returns info_ref.file_path.c_str(), the interior pointer of a C++ std::string DuckDB still owns and destroys itself. Calling duckdb_free on it hands the allocator a pointer it never issued.

The trap in the rule: the signature alone is not enough. duckdb_parameter_name is declared const char * and yet returns strdup(identifier.c_str()) — it is owned, and not freeing it leaks. The only reliable check is reading the implementation in DuckDB's src/main/capi/.

Your action: before calling duckdb_free on anything the C API returned, read the implementation. const char * is a strong hint that it is borrowed, but duckdb_parameter_name proves it is only a hint. Every duckdb_free site in quack-rs was audited this way; see LESSONS.md P11 for the full table.

How it was found: by writing the first live test for copy functions. The module had 16 unit tests and none of them registered a copy function against a real DuckDB, so the corruption had never had a chance to happen. Unit tests over an FFI wrapper test the wrapper's arithmetic, not its contract with the library.


P12: duckdb_client_context_get_config_option aborts on a missing setting

Status: A DuckDB defect, not a quack-rs one. Documented on ClientContext::config_option, with an abort-free alternative.

Symptom: Assertion 'scope != SettingScope::INVALID' failed and a SIGABRT when asking for a configuration option that does not exist — but only against a DuckDB built with debug assertions. Release builds return NULL exactly as documented, so this never reproduces for end users and always reproduces in a test suite that links a debug DuckDB.

Root cause (DuckDB 1.5.5):

// src/main/capi/config_options-c.cpp
switch (ctx.TryGetCurrentSetting(option_name, result).GetScope()) {
  ...
  default:                                    // <- INVALID is handled here
    res_scope = DUCKDB_CONFIG_OPTION_SCOPE_INVALID;

// src/include/duckdb/main/setting_info.hpp
SettingScope GetScope() {
    D_ASSERT(scope != SettingScope::INVALID); // <- but never reached in debug
    return scope;
}

The default: arm shows the not-found case is meant to be tolerated; the code just calls GetScope() before checking operator bool().

Your action: use ClientContext::config_option for settings you registered or know exist. To ask whether a setting exists, use SQL — it has no assertion on this path:

SELECT count(*) FROM duckdb_settings() WHERE name = 'my_setting';

Summary

PitfallSDK statusYour action
L1: combine config fieldsTestableTest with AggregateTestHarness::combine
L2: state double-freePreventedUse FfiState::destroy_callback
L3: panic across FFIPreventedUse init_extension, no unwrap in callbacks
L4: NULL silently dropped (no validity mask)PreventedUse VectorWriter::set_null
L5: bool UBPreventedUse VectorReader::read_bool
L6: function set namePreventedUse AggregateFunctionSetBuilder
L7: LogicalType leakPreventedUse LogicalType (RAII)
L8: NULLs reach the callback anywayPreventedUse map1/map2, or DataChunk::propagate_nulls
L9: Arrow array taken on failurePreventedUse arrow::data_chunk_from_arrow (takes by value)
L10: bind data lost on expression copyPreventedUse ScalarBindData::set (or pair set_bind_data with set_bind_data_copy)
L11: aggregate crash under OVER () / ORDER BYDuckDB defectDo not use C API aggregates in those query shapes
L12: aggregate update sees NULL rowsDocumentedSkip rows where is_valid is false
L13: running window without a destructorPreventedEvery aggregate builder registers a destructor
L14: C API differs across releasesPreventedWrappers that rely on newer behaviour check the engine version or check in Rust
L15: combine consumes its sourceDocumentedRead the source states; copy what the target needs
L16: Arrow layouts DuckDB misimportsPreventeddata_chunk_from_arrow refuses them
L17: COPY … FROM reader declares columnsPrevented (typed) / DocumentedRead the target's columns; declare none
L18: LIST reserve moves nested buffersDocumentedFetch writers again after each reserve
L19: constant addresses as identitiesFixedFfiState salts its tag with a hash of TypeId
P1: lib name mismatchScaffoldSet [lib] name in Cargo.toml
P2: API version stringConstantUse DUCKDB_API_VERSION
P3: unit tests insufficientDocumentedWrite SQLLogicTest E2E tests
P4: submodule not initializedBuild-timeNew project: git submodule add …; clone: git submodule update --init
P5: SQLLogicTest exact matchDocumentedCopy output from DuckDB CLI
P6: register set silent failPreventedBuilder returns Err
P7: VARCHAR format undocumentedPreventedUse VectorReader::read_str
P8: INTERVAL layout undocumentedPreventedUse DuckInterval
P9: dispatch table uninitialisedFixedInMemoryDb::open() initialises it via C++ shim
P10: unstable ABI tail shiftsPreventedUse init_extension; set QUACK_RS_TARGET_DUCKDB_VERSION when building
P11: freeing a borrowed const char *FixedRead the C++ impl before duckdb_free; prefer quack-rs wrappers
P12: config-option probe aborts (debug)DocumentedAsk duckdb_settings() in SQL instead

TypeId Reference

quack_rs::types::TypeId is the enum of DuckDB column types that the quack-rs builder APIs accept. Each variant names one of the C API's DUCKDB_TYPE_* integer constants, which libduckdb-sys exposes as DUCKDB_TYPE_DUCKDB_TYPE_* (for example libduckdb_sys::DUCKDB_TYPE_DUCKDB_TYPE_BIGINT).


Full variant table

VariantSQL nameC API constantNotes
TypeId::BooleanBOOLEANDUCKDB_TYPE_BOOLEANtrue/false, stored as a u8
TypeId::TinyIntTINYINTDUCKDB_TYPE_TINYINT8-bit signed
TypeId::SmallIntSMALLINTDUCKDB_TYPE_SMALLINT16-bit signed
TypeId::IntegerINTEGERDUCKDB_TYPE_INTEGER32-bit signed
TypeId::BigIntBIGINTDUCKDB_TYPE_BIGINT64-bit signed
TypeId::UTinyIntUTINYINTDUCKDB_TYPE_UTINYINT8-bit unsigned
TypeId::USmallIntUSMALLINTDUCKDB_TYPE_USMALLINT16-bit unsigned
TypeId::UIntegerUINTEGERDUCKDB_TYPE_UINTEGER32-bit unsigned
TypeId::UBigIntUBIGINTDUCKDB_TYPE_UBIGINT64-bit unsigned
TypeId::HugeIntHUGEINTDUCKDB_TYPE_HUGEINT128-bit signed
TypeId::FloatFLOATDUCKDB_TYPE_FLOAT32-bit IEEE 754
TypeId::DoubleDOUBLEDUCKDB_TYPE_DOUBLE64-bit IEEE 754
TypeId::TimestampTIMESTAMPDUCKDB_TYPE_TIMESTAMPµs since Unix epoch
TypeId::TimestampTzTIMESTAMPTZDUCKDB_TYPE_TIMESTAMP_TZtimezone-aware timestamp
TypeId::DateDATEDUCKDB_TYPE_DATEdays since epoch
TypeId::TimeTIMEDUCKDB_TYPE_TIMEµs since midnight
TypeId::IntervalINTERVALDUCKDB_TYPE_INTERVALmonths + days + µs
TypeId::VarcharVARCHARDUCKDB_TYPE_VARCHARUTF-8 string
TypeId::BlobBLOBDUCKDB_TYPE_BLOBbinary data
TypeId::DecimalDECIMALDUCKDB_TYPE_DECIMALfixed-point decimal
TypeId::TimestampSTIMESTAMP_SDUCKDB_TYPE_TIMESTAMP_Sseconds since epoch
TypeId::TimestampMsTIMESTAMP_MSDUCKDB_TYPE_TIMESTAMP_MSmilliseconds since epoch
TypeId::TimestampNsTIMESTAMP_NSDUCKDB_TYPE_TIMESTAMP_NSnanoseconds since epoch
TypeId::EnumENUMDUCKDB_TYPE_ENUMenumeration type
TypeId::ListLISTDUCKDB_TYPE_LISTvariable-length list
TypeId::StructSTRUCTDUCKDB_TYPE_STRUCTnamed fields (row type)
TypeId::MapMAPDUCKDB_TYPE_MAPkey-value pairs
TypeId::UuidUUIDDUCKDB_TYPE_UUID128-bit UUID
TypeId::UnionUNIONDUCKDB_TYPE_UNIONtagged union of types
TypeId::BitBITDUCKDB_TYPE_BITbitstring
TypeId::TimeTzTIMETZDUCKDB_TYPE_TIME_TZtimezone-aware time
TypeId::UHugeIntUHUGEINTDUCKDB_TYPE_UHUGEINT128-bit unsigned
TypeId::ArrayARRAYDUCKDB_TYPE_ARRAYfixed-length array
TypeId::TimeNsTIME_NSDUCKDB_TYPE_TIME_NSnanosecond-precision time
TypeId::AnyANYDUCKDB_TYPE_ANYwildcard for function signatures
TypeId::VarintBIGNUMDUCKDB_TYPE_BIGNUMarbitrary-precision integer (VARINT before DuckDB 1.4)
TypeId::SqlNullSQLNULLDUCKDB_TYPE_SQLNULLexplicit SQL NULL type
TypeId::IntegerLiteralINTEGER_LITERALDUCKDB_TYPE_INTEGER_LITERALunresolved integer literal
TypeId::StringLiteralSTRING_LITERALDUCKDB_TYPE_STRING_LITERALunresolved string literal
TypeId::GeometryGEOMETRYDUCKDB_TYPE_GEOMETRYspatial geometry value (duckdb-1-5-3)
TypeId::VariantVARIANTDUCKDB_TYPE_VARIANTself-describing nested value, e.g. Iceberg v3 (duckdb-1-5-3)

Feature gate for Geometry / Variant: DUCKDB_TYPE_GEOMETRY (40) and DUCKDB_TYPE_VARIANT (41) require the duckdb-1-5-3 feature, which layers on top of duckdb-1-5 and needs libduckdb-sys >= 1.10503.0 (DuckDB 1.5.3). VARIANT entered the C type enum in DuckDB 1.5.3, after the duckdb-1-5 feature's 1.5.0 floor, and GEOMETRY is gated with it so that one feature covers both. Gating them separately avoids breaking consumers pinned to libduckdb-sys 1.10500–1.10502 (DuckDB 1.5.0–1.5.2). See Known Limitations.


Methods

to_duckdb_type() → DUCKDB_TYPE

Converts to the raw C API integer constant. Used internally by the builder APIs.

#![allow(unused)]
fn main() {
use quack_rs::types::TypeId;

let raw: libduckdb_sys::DUCKDB_TYPE = TypeId::BigInt.to_duckdb_type();
}

from_duckdb_type(raw) → TypeId

Converts a raw DUCKDB_TYPE constant back into a TypeId. Recognizes every variant available in the active feature set, including TIME_NS, ANY, BIGNUM, SQLNULL, INTEGER_LITERAL and STRING_LITERAL (no feature needed: all six exist in every DuckDB this crate supports) and the duckdb-1-5-3 values (GEOMETRY, VARIANT) when that feature is enabled. Panics if the value does not correspond to any variant available in the current feature configuration; try_from_duckdb_type returns None instead.

#![allow(unused)]
fn main() {
use quack_rs::types::TypeId;

let type_id = TypeId::from_duckdb_type(libduckdb_sys::DUCKDB_TYPE_DUCKDB_TYPE_BIGINT);
assert_eq!(type_id, TypeId::BigInt);
assert_eq!(TypeId::try_from_duckdb_type(9_999), None);
}

is_composite() and composite_constructor_hint()

DECIMAL, ENUM, LIST, STRUCT, MAP, ARRAY and UNION carry parameters that a bare type id cannot express, so duckdb_create_logical_type cannot build them. is_composite() returns true for these seven, and composite_constructor_hint() names the LogicalType constructor to use instead:

#![allow(unused)]
fn main() {
use quack_rs::types::TypeId;

assert!(TypeId::List.is_composite());
assert_eq!(
    TypeId::List.composite_constructor_hint(),
    Some("LogicalType::list(element_type)")
);
assert_eq!(TypeId::BigInt.composite_constructor_hint(), None);
}

sql_name() → &'static str

Returns the SQL type name as a static string.

#![allow(unused)]
fn main() {
use quack_rs::types::TypeId;
assert_eq!(TypeId::BigInt.sql_name(), "BIGINT");
assert_eq!(TypeId::Varchar.sql_name(), "VARCHAR");
assert_eq!(TypeId::TimestampTz.sql_name(), "TIMESTAMPTZ");
}

Display

TypeId implements Display, which outputs the SQL name:

#![allow(unused)]
fn main() {
use quack_rs::types::TypeId;
println!("{}", TypeId::Interval);  // prints: INTERVAL
let s = format!("{}", TypeId::UBigInt); // "UBIGINT"
assert_eq!(s, "UBIGINT");
assert_eq!(TypeId::Interval.to_string(), "INTERVAL");
}

VectorReader/VectorWriter mapping

The read and write methods on VectorReader/VectorWriter map to TypeId variants as follows:

TypeIdRead methodWrite methodRust type
Booleanread_boolwrite_boolbool
TinyIntread_i8write_i8i8
SmallIntread_i16write_i16i16
Integerread_i32write_i32i32
BigIntread_i64write_i64i64
UTinyIntread_u8write_u8u8
USmallIntread_u16write_u16u16
UIntegerread_u32write_u32u32
UBigIntread_u64write_u64u64
Floatread_f32write_f32f32
Doubleread_f64write_f64f64
Varcharread_strwrite_varchar&str
Intervalread_intervalwrite_intervalDuckInterval
HugeIntread_i128write_i128i128
UHugeIntread_u128write_u128u128
Blobread_blobwrite_blob&[u8]
Uuidread_uuidwrite_uuidu128 (textual bits)
Dateread_datewrite_datei32 (days since epoch)
Timeread_timewrite_timei64 (µs since midnight)
TimeTzread_time_tzwrite_time_tzu64 (packed)
Timestampread_timestampwrite_timestampi64 (µs since epoch)
TimestampTzread_timestamp_tzwrite_timestamp_tzi64 (µs since epoch)
TimestampSread_timestamp_swrite_timestamp_si64 (s since epoch)
TimestampMsread_timestamp_mswrite_timestamp_msi64 (ms since epoch)
TimestampNsread_timestamp_nswrite_timestamp_nsi64 (ns since epoch)
Decimalread_decimal(row, width)write_decimal(row, width, v)i128 (unscaled)

List, Map, Struct and Array are nested vectors: use the helpers in Complex Types (ListVector, MapVector, StructVector, ArrayVector, StructReader / StructWriter, ListBuilder).

Enum, Union, Bit, TimeNs, Any, Varint, SqlNull, IntegerLiteral, StringLiteral, Geometry and Variant do not yet have dedicated read/write helpers. Access these via the raw data pointer from duckdb_vector_get_data.


Properties

TypeId implements Debug, Clone, Copy, PartialEq, Eq, and Hash, so it can be used as a map key or set element and compared in match expressions:

#![allow(unused)]
fn main() {
use std::collections::HashMap;
use quack_rs::types::TypeId;

let mut type_names: HashMap<TypeId, &str> = HashMap::new();
type_names.insert(TypeId::BigInt, "count");
type_names.insert(TypeId::Varchar, "label");
}

#[non_exhaustive]

TypeId is marked #[non_exhaustive], so quack-rs can add variants for new DuckDB types without a breaking change. A match on TypeId outside quack-rs needs a wildcard arm:

#![allow(unused)]
fn main() {
use quack_rs::types::TypeId;
fn demo(type_id: TypeId) {
match type_id {
    TypeId::BigInt => { /* ... */ }
    TypeId::Varchar => { /* ... */ }
    _ => { /* handle future types */ }
}
}
}

LogicalType

For types that require parameters (such as DECIMAL(p, s) or LIST(INTEGER)), use quack_rs::types::LogicalType:

#![allow(unused)]
fn main() {
use quack_rs::types::{LogicalType, TypeId};

let lt = LogicalType::new(TypeId::BigInt);
// or use the From impl:
let lt: LogicalType = TypeId::BigInt.into();
// LogicalType implements Drop → calls duckdb_destroy_logical_type automatically
}

LogicalType wraps duckdb_logical_type with RAII cleanup, preventing the memory leak described in Pitfall L7.

Constructors

ConstructorCreates
new(type_id)Simple type from a TypeId
from_raw(ptr)Takes ownership of a raw handle (unsafe)
decimal(width, scale)DECIMAL(width, scale)
list(element_type)LIST<T> from a TypeId
list_from_logical(element)LIST<T> from an existing LogicalType
map(key, value)MAP<K, V> from TypeIds
map_from_logical(key, value)MAP<K, V> from existing LogicalTypes
struct_type(fields)STRUCT from &[(&str, TypeId)]
struct_type_from_logical(fields)STRUCT from &[(&str, LogicalType)]
union_type(members)UNION from &[(&str, TypeId)]
union_type_from_logical(members)UNION from &[(&str, LogicalType)]
enum_type(members)ENUM from &[&str]
array(element_type, size)ARRAY<T>[size] from a TypeId
array_from_logical(element, size)ARRAY<T>[size] from an existing LogicalType

Each constructor except from_raw has a try_ form (try_new, try_decimal, try_list, …) that returns an error instead of panicking on invalid input, such as a composite TypeId passed to new, a DECIMAL width above 38, or a UNION with more than MAX_UNION_MEMBERS (255) members.

Introspection methods

The introspection methods are all unsafe, because they call into a loaded DuckDB:

get_type_id, get_alias, set_alias, decimal_width, decimal_scale, decimal_internal_type, enum_internal_type, enum_dictionary_size, enum_dictionary_value, list_child_type, map_key_type, map_value_type, struct_child_count, struct_child_name, struct_child_type, union_member_count, union_member_name, union_member_type, array_size, array_child_type.

See Type System for the full introspection table.

Known Limitations

What quack-rs cannot do because the DuckDB C extension API does not allow it, DuckDB behaviour an extension should plan around, and former limitations that have since been resolved.

Window functions are not available

DuckDB window functions (OVER (...) clauses) are implemented entirely in DuckDB's C++ layer and have no counterpart in the public C extension API.

This is not a gap in quack-rs or in libduckdb-sys — the relevant symbol (duckdb_create_window_function) simply does not exist in the C API:

SymbolC API (1.4.x)?C API (1.5.0+)?C++ API?
duckdb_create_window_functionNoNoYes
duckdb_create_copy_functionNoYesYes
duckdb_create_scalar_functionYesYesYes
duckdb_create_aggregate_functionYesYesYes
duckdb_create_table_functionYesYesYes
duckdb_create_cast_functionYesYesYes

What this means for your extension:

A custom window operator requires a C++ extension. An ordinary aggregate can be used as a window (agg(x) OVER (...)), but see the next section before recommending that to your users: two of those shapes crash every aggregate registered through the C API.

If DuckDB exposes window registration in a future C API version, quack-rs will add wrappers in the corresponding release.

Aggregates crash under OVER () and ORDER BY (DuckDB defect, Pitfall L11)

Every aggregate registered through the C API — quack-rs's or anyone else's — reads out of bounds when DuckDB runs it as a whole-partition window (agg(x) OVER (), agg(x) OVER (PARTITION BY p)) or as an ordered aggregate (agg(x ORDER BY y)). The usual result is a segmentation fault that takes the host process down. The book's own example aggregate does it:

-- with hello-ext loaded: segfaults on DuckDB 1.4.4 and 1.5.5
SELECT max(w) FROM (SELECT word_count(s) OVER () AS w
                    FROM (SELECT 'a b c' AS s FROM range(5000)));

The cause is in DuckDB (CAPIAggregateUpdate hands the callback a constant state vector), there is no way for an extension to detect or prevent it, and it is reported upstream as duckdb/duckdb#26109. Until it is fixed, tell your users not to use your aggregates in those two shapes. Frames that are not whole-partition (ROWS BETWEEN 5 PRECEDING AND CURRENT ROW) and DISTINCT windows work. See Pitfall L11.

Aggregate states leak when finalize reports an error (DuckDB behaviour)

When an aggregate's finalize callback reports an error (AggregateFunctionInfo::set_error), DuckDB 1.5.5 does not call the destructor for every state the query created: an ungrouped query initialised 2 states and destroyed 1, a grouped one 4 and 2. When finalize succeeds, every state is destroyed. The extension cannot tell which states were abandoned, so whatever they own is leaked: with FfiState<T>, whatever T owns on the heap, and the box of a T too large to store inline (see the next section). The query still fails with your message. If that leak matters (a long-lived process whose queries often fail this way), keep what a state owns small. Only finalize was measured; errors reported from other callbacks were not. The behaviour is pinned by aggregate_states_are_not_all_destroyed_when_finalize_fails in tests/ffi_roundtrip/lifecycle.rs.

Grouped-aggregate states the scan never reaches are never destroyed (DuckDB defect)

DuckDB 1.4.4 to 1.5.5 destroys a grouped aggregate's states as its result scan passes them. When the scan stops early, the states it has not reached are never destroyed — not after the query, not when the connection or the database closes. That happens under a LIMIT above the aggregate, an error raised above it, or an interrupt (InterruptHandle::cancel). Measured on one thread with 300,000 groups: under LIMIT 10, 2,048 of 300,000 states were destroyed (4,096 on 1.4.x); with an error raised half-way through the result, 151,552. DuckDB's own aggregates are affected the same way: mode() under LIMIT 10 leaked about 100 MB per query. See docs/upstream-duckdb-reports.md, item 20.

The query's answer is right; the cost is memory. FfiState<T> stores a T of at most 256 bytes, aligned no more strictly than usize, inside DuckDB's own state bytes, which DuckDB frees with the hash table, so such a state leaks nothing unless T itself owns heap memory (a Vec, String or HashMap). A larger T is boxed, and the box leaks with it. Until DuckDB fixes this, keep aggregate states small and free of heap allocations where you can. Pinned by states_a_grouped_scan_never_reaches_leak_no_rust_heap in tests/aggregate_leaks.rs.

Window frames with EXCLUDE never destroy one state per row (DuckDB defect)

A C API aggregate in a window whose frame has EXCLUDE CURRENT ROW, GROUP or TIES is evaluated by DuckDB's segment tree in two parts, and the second part initialises one state per row that is never destroyed: over a 5000-row window, 5000 states on every release from 1.4.4 to 1.5.5 (none without EXCLUDE). The answer is right; as above, a small FfiState<T> leaks only what T owns on the heap. See docs/upstream-duckdb-reports.md, item 35; pinned by a_window_frame_with_exclude_leaves_states_undestroyed in tests/ffi_roundtrip/agg_states.rs.

An abandoned stream keeps its table-function state (DuckDB behaviour)

Dropping a streaming QueryResult part-way through, and then its PreparedStatement, does not free the query's operator states: DuckDB keeps the active query on the connection until the next statement runs there (or the connection closes). A table function's state — for a typed table function, the with_state value and its per-scan clone — therefore lives until then. Nothing leaks, but a state that holds a file, a lock or a large buffer holds it for that long; run any statement (SELECT 1) on the connection to release it. Pinned by an_abandoned_stream_keeps_its_table_state_until_the_next_statement in tests/ffi_roundtrip/query_stream.rs.

Running out of memory inside a callback aborts the process

An allocation failure is not an error quack-rs can report. On the Rust side, the default allocation-error handler aborts. On the DuckDB side, duckdb_list_vector_reserve, duckdb_vector_copy_sel and duckdb_vector_assign_string_element_len allocate without catching, so their std::bad_alloc crosses the Rust callback frame and aborts the process ("Rust cannot catch foreign exceptions"). quack-rs allocates inside callbacks only on error paths and in data_chunk_to_arrow's pre-export check, which copies each column it checks. Bound what your callbacks allocate, and set DuckDB's memory_limit so its own operators fail cleanly before the process runs out.

One allocation failure is worse than an abort. When duckdb_prepare fails to allocate while it records a statement's parameter names, it frees the statement it has already handed back and reports an error; prepare (and everything built on it) then reads and frees that statement again. This is undefined behaviour inside DuckDB's C API that no caller can detect; see docs/upstream-duckdb-reports.md, item 36.

COPY functions (resolved in DuckDB 1.5.0; both directions since)

DuckDB 1.5.0 added duckdb_create_copy_function and related symbols to the public C extension API. quack-rs wraps these in the copy_function module behind the duckdb-1-5 feature flag. See CopyFunctionBuilder for usage.

This was previously listed as a known limitation (no C API counterpart prior to 1.5.0).

COPY … FROM was a second, narrower gap: quack-rs wrapped the writing half only. CopyFunctionBuilder::copy_from now attaches a quack-rs table function as a format's reader, and a copy function may implement either direction or both — a read-only format leaves the writing callbacks unset entirely. See the Copy Functions chapter.

Arrow interop (resolved behind duckdb-1-5-4)

DuckDB's C API has a family of conversion functions (present since 1.4.4) that move data directly between a duckdb_data_chunk and the Arrow C Data Interface. quack-rs wraps all eight non-deprecated entries in the arrow module, with no arrow crate dependency — see the Arrow Interop chapter.

The remaining fourteen Arrow entries in the C API struct are the older duckdb_query_arrow result API, which lives inside #ifndef DUCKDB_API_NO_DEPRECATED; they are deliberately not wrapped.

The feature is duckdb-1-5-4 rather than duckdb-1-5 because libduckdb-sys declared the two Arrow ABI records as opaque zero-sized placeholders until 1.10504.0. The DuckDB functions themselves are present in every release quack-rs supports, but because the feature implies duckdb-1-5, an extension built with it needs a DuckDB 1.5.0+ engine.

Callback accessor wrappers (resolved)

quack-rs wraps the callback accessor functions — the C API functions used inside your callbacks to retrieve arguments, set errors, access bind data, and so on:

CategoryWrapper typeAvailable
Scalar function executionScalarFunctionInfoAlways
Scalar function bindScalarBindInfoduckdb-1-5
Scalar function initScalarInitInfoduckdb-1-5
Aggregate function callbacksAggregateFunctionInfoAlways
Table function bindBindInfoAlways
Table function initInitInfoAlways
Table function scanFunctionInfoAlways
Cast function callbacksCastFunctionInfoAlways
Copy function bindCopyBindInfoduckdb-1-5
Copy function global initCopyGlobalInitInfoduckdb-1-5
Copy function sinkCopySinkInfoduckdb-1-5
Copy function finalizeCopyFinalizeInfoduckdb-1-5

Where the C API provides a client context for a callback — scalar bind and init, table function bind, and the four copy-function callbacks — the wrapper exposes it as get_client_context, which returns a ClientContext (see the client_context module).

Complex type creation (resolved)

LogicalType provides constructors for all complex parameterized types:

MethodType created
LogicalType::decimal(width, scale)DECIMAL(p, s)
LogicalType::enum_type(members)ENUM('a', 'b', ...)
LogicalType::array(child, size)type[N]
LogicalType::union_type(members)UNION(a INT, b VARCHAR)
LogicalType::list(child)LIST(type)
LogicalType::struct_type(fields)STRUCT(...)
LogicalType::map(key, value)MAP(K, V)

The constructors that take child types (list, array, map, struct_type, union_type) have _from_logical variants for nested complex types, and each constructor has a try_ form that returns an error instead of panicking. Introspection methods (get_type_id, list_child_type, struct_child_count, decimal_width, etc.) are also available.

VARIANT and GEOMETRY types (resolved — exposed behind duckdb-1-5-3)

The VARIANT type (a self-describing nested value, used for example by Iceberg v3) entered the C type enum as DUCKDB_TYPE_VARIANT (41) in DuckDB 1.5.3. GEOMETRY (DUCKDB_TYPE_GEOMETRY, 40) was already present earlier in the 1.5.x line.

quack-rs exposes these as TypeId::Variant and TypeId::Geometry, gated behind the duckdb-1-5-3 feature. That feature layers on top of duckdb-1-5 and requires libduckdb-sys >= 1.10503.0 (DuckDB 1.5.3). The separate gate exists because VARIANT postdates the duckdb-1-5 feature's 1.5.0 floor; GEOMETRY is gated with it so that one feature covers both values. Keeping them out of duckdb-1-5 preserves compatibility for consumers pinned to libduckdb-sys 1.10500–1.10502 (DuckDB 1.5.0–1.5.2).

[dependencies]
quack-rs = { version = "0.18", features = ["duckdb-1-5-3"] }

Neither type yet has dedicated VectorReader/VectorWriter helpers; access their data via the raw pointer from duckdb_vector_get_data when needed.

Changelog

All notable changes to quack-rs, mirrored from CHANGELOG.md.

The format follows Keep a Changelog. quack-rs adheres to Semantic Versioning.

Unreleased

0.18.0 — 2026-09-30

This release comes out of a second production-readiness audit (AUDIT.md section 7) and the three passes that followed it (sections 8 to 10). Each code defect below was either reproduced against a real DuckDB before it was fixed or, where nothing could trigger it, derived from DuckDB's source; AUDIT.md records which (VALIDATED or PROVEN). Each code fix has a regression test where one could be written, and the trait-bound fixes are pinned by compile_fail doctests. The exceptions: the fixes AUDIT.md marks PROVEN only, among them the entry points' NULL duckdb_database*, AbiPolicy::Warn's eprintln! and catalog lookups in a catalog that is not "duckdb". The FileHandle drop fix is tested by swapping failing C API stubs into the dispatch table, since no file system DuckDB ships throws from Close(). The 32-bit size fixes are unit-tested on wasm32 itself (CI's wasm job runs the unit tests under node); what DuckDB does with an oversized allocation there is derived from its source, as no wasm32 build of DuckDB is available to test against. Several fixes close holes in the safe API — places where safe code could cause undefined behaviour, a data race or a process abort — and those needed signature or trait-bound changes, so this is a breaking release; it follows 0.16.0. Each such entry is marked Breaking:.

0.17.0 was prepared but never published; its changes are included here, and the entries below describe the change from 0.16.0.

A third pass followed before anything was published (AUDIT.md section 8). It fixed further defects, reproducing them against a real DuckDB wherever that was possible, added CI gates that compile the Rust examples in the book and the README and check the book's links, and corrected the documentation's claim that an aggregate's update never sees NULL rows (see Fixed).

A fourth pass followed that (AUDIT.md section 9). It fixed process aborts, out-of-bounds reads and writes, and wrong answers. Each was reproduced against a real DuckDB before it was fixed, and its regression test was shown failing without the fix; AUDIT.md names the DuckDB versions per finding. The pass also documented thirteen DuckDB defects in docs/upstream-duckdb-reports.md, each with a plain-C reproducer.

A fifth pass followed (AUDIT.md section 10). Each code fix has a regression test shown failing without the fix, except as noted above, and each defect that involves DuckDB was reproduced against a real DuckDB first, except the catalog lookup and the 32-bit size fixes, which are PROVEN. It documented eighteen DuckDB defects (items 20 to 37 of docs/upstream-duckdb-reports.md), each with a plain-C reproducer run on the releases it names.

Within Added, Changed, Fixed and Security, entries are grouped by the pass that produced them: Fifth audit, Fourth audit, and Earlier passes (the 0.17.0 work, the second audit and the third). Security has no Fifth audit group; that pass's memory-safety fixes are under Fixed. Dependencies is not grouped.

Added

Fifth audit

  • tests/aggregate_leaks.rs, a test binary with a counting global allocator, and tests/handle_leaks.rs, which bounds the C heap (glibc mallinfo2) across many create-and-drop rounds of 15 handles whose Drop no functional test observed (DbConfig, Value, Appender, TableDescription, OwnedDataChunk and PreparedStatement::parameter_name; with duckdb-1-5 also Expression, ClientContext, FileOpenOptions, FileSystem, FileHandle, SelectionVector, InstanceCache, ErrorData and Catalog). It runs only on Linux with glibc.
  • vector::max_child_capacity: the most elements DuckDB can hold in a list vector's child buffer, which ListBuilder now respects.
  • A value_render fuzz target (fuzz/, feature live) that renders arbitrary temporal payloads, alone and nested in lists, through a real DuckDB.
  • CI: test-older-engines also runs the suite against 1.4.5 (the last 1.4 release) and against 1.5.3 and 1.5.4 with the duckdb-1-5-3 and duckdb-1-5-4 features: the first releases whose bindings those features compile against (the job pins the bindings to the engine's release, which the test shim requires; a duckdb-1-5-4 build needs only a 1.5.0+ engine at run time). tests/append_metadata_cli.rs runs the append_metadata binary end to end.
  • CI: the miri job also runs the library tests on a big-endian target (s390x-unknown-linux-gnu, interpreted), which is where the string decoder's byte-order defect shows (see Fixed).

Fourth audit

  • scalar_bind_callback! / scalar_init_callback! (duckdb-1-5).
  • validate_parameter_name, DUCKDB_UNCALLABLE_KEYWORDS, DUCKDB_UNREFERENCEABLE_PARAMETER_KEYWORDS.
  • value::UNRENDERABLE, callback::EMPTY_PANIC_PLACEHOLDER, callback::panic_c_message, CopyGlobalInitInfo::get_file_path_bytes, scalar::info::EMPTY_ERROR_PLACEHOLDER, aggregate::info::EMPTY_ERROR_PLACEHOLDER.
  • CI: test-older-engines runs the suite against DuckDB 1.4.4 (default features) and 1.5.0 (duckdb-1-5), the oldest release each can load into; every engine-specific defect below went unnoticed without it. A weekly scheduled run; the scaffold job runs the generated project's make configure release test; check-abi-table.py fingerprints whole signatures, not names.

Earlier passes (0.17.0, second and third audits)

  • Scalar functions as safe Rust closures. ScalarFunctionBuilder::map1 / map2 / map1_str / map2_str / map1_opt / map2_opt take an ordinary closure; parameter and return types come from its signature, NULLs propagate correctly, and a panic becomes a SQL error. VARCHAR gets its own constructors so the closure can borrow a &str straight out of the vector. One indirect call per chunk, not per row. Each returns Result<TypedScalarFunctionBuilder, _>, which offers name(), volatile() and register(con) but cannot change the signature (code that passes it to a Registrar calls Registrar::register_typed_scalar), and the trampoline re-checks each chunk's vector types, so the declared result width always matches the closure's. Building one does not call DuckDB, so it works in a unit test with MockRegistrar; register checks the types.

  • Value gained constructors for the remaining integer and float widths, the rest of the temporal family, INTERVAL, BLOB, DECIMAL, and the composites (STRUCT, LIST, ARRAY, ENUM; MAP and UNION behind duckdb-1-5); there is still none for BIT or BIGNUM. Also is_sql_null (distinct from is_null, which asks about the handle) and as_enum_index. struct_value checks the field count first, because duckdb_create_struct_value takes no count and reads one value per field of the type. list_value / array_value take the element type: duckdb.h contradicts itself here, and the implementation settles it. The element type may itself be a LIST or ARRAY (lists of lists, arrays of arrays); the error explaining the element-type rule appears only when DuckDB itself refuses the value. decimal checks width (1..=38), scale and the unscaled value's digit count before calling DuckDB, which would otherwise abort the process or store a different number. The temporal constructors other than date and interval return Result and refuse a payload DuckDB cannot render (see the Value::time_ns / Value::timestamp entry under Changed).

  • PreparedStatement gained 16 more typed binds and bind_value, the escape hatch for every composite type. bind_decimal validates like Value::decimal: duckdb_bind_decimal checks nothing, and for width <= 18 keeps only the low 64 bits of the unscaled value.

  • Cancellation and progress. OwnedConnection::interrupt_handle returns a Send + Sync InterruptHandle, lifetime-tied to the connection, whose cancel a watchdog thread can use to stop a running query and whose progress reads its QueryProgress. OwnedConnection::interrupt / progress do the same through the connection itself, and the unsafe query::interrupt / query::query_progress take a raw connection.

  • Streaming results. PreparedStatement::execute_streaming and QueryResult::is_streaming (duckdb-1-5).

  • QueryResult::column_logical_type keeps the nested structure that column_type collapses, and result_kind separates rows from row counts (the new query::ResultKind).

  • LogicalType::register — CREATE TYPE from the C API, so an extension can ship a named ENUM or STRUCT. Stable-prefix; no feature needed.

  • ScalarBindData<T> / ScalarLocalState<T> (duckdb-1-5) — typed bind data and per-thread local state for scalar functions, with the duckdb_delete_callback_t generated and panic-safe. The raw set_bind_data / set_state route makes the extension author write their own unsafe extern "C" fn around Box::from_raw, which is the abort hazard under Security relocated into user code. DuckDB reads scalar bind data from every executing thread at once, so ScalarBindData<T> requires T: Send + Sync + 'static and ScalarLocalState<T> T: Send + 'static. ScalarBindData::set also requires T: Clone (wrap other data in Arc<T>): it registers a generated, panic-safe copy callback, without which the bind data is lost whenever the optimizer copies the bound expression — a wrong answer, not an error (Pitfall L10; see ScalarBindInfo::set_bind_data_copy under Fixed). A second set leaks, rather than drops, the first value.

  • vector::ops (duckdb-1-5) makes SelectionVector usable: copy_selected, slice, reference_value, reference_vector and OwnedVector. Documents that slice produces a dictionary vector, after which every reader in this crate reads the wrong rows. OwnedVector::new walks the type, child vectors included, with checked arithmetic and refuses a capacity above vector::ops::MAX_CAPACITY before DuckDB allocates: DuckDB computes the buffer size with an unchecked multiply, so (HUGEINT, 2^60 + 2) would get a 32-byte buffer.

  • Arrow C Data Interface bridge — the new arrow module behind a new duckdb-1-5-4 feature, wrapping the eight-function conversion family already in DuckDB 1.4.4's C API: ArrowOptions<'conn>, which borrows its connection so that a conversion cannot read freed memory after the connection closes (the safe from_connection(&'conn OwnedConnection), and the unsafe from_raw_connection, from_result and QueryResult::arrow_options), owning ArrowSchema / ArrowArray / ArrowConvertedSchema RAII types, and to_arrow_schema / data_chunk_to_arrow / schema_from_arrow / data_chunk_from_arrow. No arrow crate dependency: the module works on the ABI records libduckdb-sys defines, which have arrow-rs's FFI_ArrowSchema / FFI_ArrowArray layout, so bridging is a pointer cast — ArrowArray::take_from moves a record out of a foreign wrapper and leaves a released placeholder, so only one side ever calls release.

    The ownership rules were read out of arrow-c.cpp and arrow_converter.cpp rather than inferred: duckdb_data_chunk_from_arrow sets arrow_array->release = nullptr before the conversion loop body, so it claims the array on the error path too — hence data_chunk_from_arrow takes the array by value, and the by-value binding still releases it in the one case (a zero-column schema) where the loop never runs. to_arrow_schema and data_chunk_to_arrow install release last, after everything that can throw, so a failed conversion leaves nothing to free.

    Two crashes DuckDB does not guard are refused here instead: duckdb_data_chunk_from_arrow indexes arrow_array->children[i] once per schema column with no bounds check and dereferences an already-released array, so ArrowConvertedSchema remembers its column count and both are checked first. data_chunk_from_arrow also returns InvalidInput for a negative length, a nonzero top-level offset (DuckDB ignores it and imports the wrong rows), a null children pointer, a null child, and a child shorter than the array's length, each of which DuckDB would dereference, read out of bounds or misread, and for a zero-row array, which DuckDB passes on as a zero-byte allocation that a debug build asserts against. It refuses a length, or a nested row count, above vector::ops::MAX_CAPACITY, and the valid Arrow layouts DuckDB imports from the wrong rows or out of bounds (see src/arrow/import_layout.rs and the Fifth audit entries under Fixed). Its Safety section requires that the array conform to the schema — nothing in an Arrow array records its type, so a child whose buffers do not match the type its schema declares cannot be checked — and that the length be the array's true row count: DuckDB allocates the chunk before its try block, so a length below that ceiling that the system cannot allocate still aborts the process. Dictionary-encoded and null-type columns are copied into flat vectors before it returns; run-end-encoded children are expanded by DuckDB. TIMETZ comes back as TIME without its offset and BIT as BLOB.

    The duckdb-1-5-4 feature's floor is set by the bindings, not by DuckDB: the eight functions predate 1.5, and the feature needs only the 1.5.0+ engine that duckdb-1-5 does at run time, but libduckdb-sys declared ArrowSchema / ArrowArray as opaque zero-sized bindgen placeholders until 1.10504.0. src/arrow.rs carries a const assertion that says so.

  • COPY … FROM — CopyFunctionBuilder::copy_from attaches a quack-rs table function as a format's reader, so an extension can implement loading as well as writing. Supporting pieces:

    • TableFunctionBuilder::build_handle returns a configured, unregistered TableFunctionHandle — register is now that plus duckdb_register_table_function — because a COPY … FROM reader is attached to a copy function rather than registered on its own.
    • BindInfo::result_column_count / result_column_name / result_column_type (duckdb-1-5) read the target table's schema, which COPY … FROM fixes before the bind callback runs. duckdb.h is explicit that such a bind "should not define its own result columns".
    • CopyBindInfo::options exposes the COPY … TO options as the STRUCT value DuckDB builds, and Value::struct_field_names walks the field names off the value's borrowed logical type without exposing the handle (pitfall P11); for a UNION it returns an empty name for the tag, then the member names.
    • CopyFunctionBuilder::extra_info, with the same ownership-until-transfer guarantee as the other builders.

    duckdb_copy_function_set_copy_from_function reports every rejection by doing nothing at all, so copy_from checks duckdb.h's stated precondition — "the table function must take a single VARCHAR parameter (the file path)" — which DuckDB never enforces: CCopyFromBind builds the argument list itself and never consults tf.arguments, so a mismatch surfaces much later inside the reader's own bind callback.

  • TableFunctionBuilder::with_bind_init(bind, init): immutable bind data B: Send + Sync and a fresh scan state S: Send per execution, no Clone needed.

  • TypedScalarFunctionBuilder; Registrar::register_typed_scalar (with a default implementation, so existing Registrars compile).

  • StructWriter::set_row_null.

  • callback::drop_panic_payload, callback::take_panic_message and callback::MAX_NESTED_PAYLOAD_DROPS.

  • selection_vector::MAX_LEN; datetime::is_valid_date, MICROS_PER_DAY, TIME_TZ_MAX_OFFSET_SECONDS, DECIMAL_MAX_WIDTH.

  • CatalogEntryType::is_lookup_supported; ReplacementScanInfo::EMPTY_ERROR_PLACEHOLDER.

  • Aggregate function sets support a different return type per overload (#121). DuckDB resolves an aggregate overload from its parameter types and arity only — the return type takes no part in resolution — so members of one set are free to return different types, which is how DuckDB's own arg_max(ANY, ANY) -> ANY and arg_max(ANY, ANY, ANY) -> ANY[] coexist. quack-rs had the return type on AggregateFunctionSetBuilder alone, which made that impossible to express.

    AggregateOverloadBuilder now carries returns / returns_logical, and AggregateFunctionSetBuilder::overload(..) takes one fully-configured overload — mirroring ScalarFunctionSetBuilder::overload / ScalarOverloadBuilder, which already worked this way:

    #![allow(unused)]
    fn main() {
    AggregateFunctionSetBuilder::new("my_agg")
        .overload(
            AggregateOverloadBuilder::new()
                .param(TypeId::Integer)
                .returns(TypeId::Integer)
                // ... callbacks
        )
        .overload(
            AggregateOverloadBuilder::new()
                .param(TypeId::Varchar)
                .returns(TypeId::Varchar)
                // ... callbacks
        )
        .register(con)?;
    }

    This is additive. returns / returns_logical on the set now act as a default for every overload that does not set its own, so existing returns(..).overloads(range, ..) code keeps working unchanged. As for return types, registration fails, naming the overload index, only when an overload has neither its own nor a set-level default; a missing callback, a duplicate signature or a composite TypeId fails it too.

  • Unit tests asserting that every AggregateOverloadBuilder callback setter stores into its own field. These setters are const fn, which makes their cargo-mutants mutants unviable rather than caught — Default::default() cannot be called in a const context, so the replacement fails to compile and the mutation gate is structurally silent about them. That is a property of the gate, not evidence the setters work.

  • An AddressSanitizer job (ci.yml; blocking — it was informational until its first green run). leak-check answers "did we forget a destructor"; ASAN answers "did we write outside an allocation, or use one after free" — the class behind the two heap-corruption defects fixed in v0.16.0, and the one a crate doing raw pointer arithmetic into DuckDB's memory is most exposed to. Miri cannot reach those paths: they call foreign functions.

  • scripts/duckdb-version-from-lock.sh — derives the DuckDB release tag from the libduckdb-sys pin in Cargo.lock (1.10505.0 → v1.5.5). Two CI jobs hard-coded v1.5.4 next to a comment asking the next person to keep it in sync with the dependency; bumping the lockfile would have left both linking a libduckdb one release older than the bindings being generated against it. test-older-engines checks its output against the engine each of its five matrix entries names, two of them in the pre-1.5 scheme.

  • scripts/sync-book-changelog.py — the book's changelog page is now generated from CHANGELOG.md, with a --check mode wired into the doc job. Kept by hand, the mirror had fallen ~19 KB behind and the published 0.16.0 entry was missing its entire "Portability and feature-combination breakage" subsection. The deliberate differences — the book page's own preamble, an em dash in release headings, and links to repository files rewritten as GitHub URLs — are applied by the script.

  • timeout-minutes on every ci.yml job (36 at this release). The default is 360 per job, so a hang in test-bundled (which compiles DuckDB from C++ source) or leak-check (-Zbuild-std) burned six hours of runner time.

  • First end-to-end coverage of the aggregate function-set registration path (tests/ffi_roundtrip.rs). Neither the aggregate nor the scalar set builder had an E2E test, despite Pitfall L6 — a set member whose name is unset is dropped silently. Four new tests register a real three-overload set (BIGINT -> BIGINT, VARCHAR -> VARCHAR, (BIGINT, BIGINT) -> DECIMAL(18,2) via returns_logical), assert typeof(..) per overload, check the computed values across multiple chunks and under GROUP BY, and cover the set-level default and both rejection paths.

  • LogicalType::try_get_type_id, which returns None for a type id this crate does not know; get_type_id panics there, and now documents that it does.

  • A fallible form of every LogicalType constructor: the new try_decimal, try_array, try_array_from_logical, try_list_from_logical and try_map_from_logical join the existing try_* functions. Also types::logical_type::MAX_UNION_MEMBERS (255: DuckDB 1.5.6 lowered its limit from 256, and asserts it when building the type).

  • VectorWriter::try_write_varchar / try_write_blob, which return an error and write nothing for a value longer than the new vector::string::MAX_STRING_LEN (u32::MAX bytes, DuckDB's string length limit); also vector::string::check_string_len.

  • ListBuilder::with_element_limit, to cap list lengths that come from input at what fits in memory. Staying under the builder's own ceiling is not enough — MAX_LIST_CHILD_CAPACITY bytes per child buffer, which max_child_capacity turns into an element limit for the child's type (see the Fifth audit entries under Fixed): duckdb_list_vector_reserve has no try/catch, so a failed allocation below that ceiling also aborts the process. ListVector::reserve now documents this.

  • vector::ops::MAX_CAPACITY, the largest capacity OwnedVector::new accepts (2^37 elements, DuckDB's MAX_VECTOR_SIZE, on a 64-bit target; 2^28 - 1 on a 32-bit one).

  • MockVectorWriter::set_valid and MockVectorWriter::is_written (see the MockVectorWriter entry under Changed).

  • CatalogEntryType::may_autoload_extension, the name check behind the new catalog-lookup refusal (see Changed).

  • An EMPTY_ERROR_PLACEHOLDER for table functions (table::info::EMPTY_ERROR_PLACEHOLDER), copy functions (copy_function::info::EMPTY_ERROR_PLACEHOLDER) and casts (CastFunctionInfo::EMPTY_ERROR_PLACEHOLDER), and callback::CAST_FAILED_WITHOUT_MESSAGE (see Fixed).

  • InstanceCache is Send + Sync. DuckDB's instance cache guards its own state with a mutex, so the wrapper was !Send only because it holds a raw pointer.

  • WarningSeverity implements PartialOrd and Ord in the order Info < Low < Medium < High < Critical, so severity >= WarningSeverity::High selects the warnings that need attention.

  • ScalarOverloadBuilder gains volatile, varargs, varargs_logical and, with duckdb-1-5, bind / init; AggregateOverloadBuilder gains extra_info. Each behaves as it does on the single-function builder, including who frees extra_info when registration fails.

  • The prelude exports ScalarBindData and ScalarLocalState (with duckdb-1-5, beside ScalarBindInfo / ScalarInitInfo). Its documentation now lists every re-export, and a unit test fails when one is missing.

  • validate_spdx_license accepts <license> WITH <exception> (for example Apache-2.0 WITH LLVM-exception), which it used to reject. The exception must be on the SPDX exception list, now public as validate::spdx::SPDX_LICENSE_EXCEPTIONS, or an AdditionRef- id.

  • classify_extension_version accepts one leading v on the semantic-version forms (v1.0.0): the spelling DuckDB's versioning documentation uses, and what extension-ci-tools stamps when HEAD carries a vX.Y.Z tag.

  • Pitfall L12 (LESSONS.md, book/src/reference/pitfalls.md): an aggregate's update receives NULL rows under the default NULL handling too. See Fixed.

  • Projects generated by generate_scaffold test something. The generated project's cargo test ran zero tests and its test/sql/<name>.test held only require and commented-out examples, so both CI steps passed whatever the extension did. The project now has a unit test and a SQLLogicTest that queries <name>_hello and checks its output. The generated Makefile checks that extension-ci-tools is checked out and names both commands: git submodule add in a new repository (where the documented git submodule update --init clones nothing) and update --init in a clone (Pitfall P4).

  • The book and the README are compiled as doctests. book/doctest is a standalone crate that turns every page under book/src (except the changelog) and README.md into rustdoc input, so every Rust block not marked ignore compiles against the working copy, and runs unless it is marked no_run; test-bundled-prebuilt runs it. Each remaining ignore block says why. mdbook test cannot do this: it passes no --extern quack_rs, so no block that uses the crate resolves it. The blocks that failed, and the content errors that turned up, are listed under Fixed. release.yml's gate also runs the crate's own doctests now (cargo test --doc --features duckdb-1-5-4).

  • New CI jobs and checks (ci.yml):

    • book builds the book with mdBook on every pull request, not only after a merge to main, then runs scripts/check-book-links.py, which checks every relative link and anchor in the book and README and every docs.rs link against this checkout's rustdoc. mdBook checks neither.
    • autoload-entries runs scripts/check-autoload-entries.py, which fails when a DuckDB release from v1.5.0 on autoloads an extension for a type or collation name missing from the lists behind the catalog-lookup refusal.
    • abi-guard-layout exercises the real LayoutMismatch path: a duckdb-1-5 build stamped C_STRUCT must be refused by DuckDB v1.5.0 and v1.4.4 and must load into the release its bindings come from. The existing abi-guard job only ever reached the declared-version check.
    • dependency-floor resolves duckdb / libduckdb-sys to the declared 1.4.4 floor and runs cargo test --lib, on stable: at that floor the dependency tree needs a newer rustc than the 1.86.0 MSRV.
    • extension-load runs scripts/check-hello-ext.py, which executes every statement documented in examples/hello-ext/README.md against the loaded extension and compares the output with examples/hello-ext/sql_checks.txt.
    • semver fails when cargo-semver-checks ran 0 checks against the published baseline for a bump that is not breaking; before, 0 checks passed. For a breaking bump it runs none by design, so a new informational step compares against origin/main with --release-type patch and lists every breaking change on the branch in the log and the job summary.
    • Integration tests fail when a documented "N pitfalls" count differs from LESSONS.md, when the source trees in CONTRIBUTING.md and the book miss a file under src/ or tests/ or list one that does not exist, and when a documentation paragraph mentions panic = "abort" without warning against it.

Changed

Fifth audit

  • Breaking: FfiState<T> stores a small T in DuckDB's state bytes. A T of at most 256 bytes, aligned no more strictly than usize, is kept inline; a larger one is still boxed. In 0.16.0 FfiState<T> was a one-word #[repr(C)] struct with a public inner: *mut T field that always boxed T; its layout is now private — a tag derived from T's TypeId (see Security), then T or the box. Use its callbacks and with_state / with_state_mut as before. See Fixed.
  • Breaking: AggregateState requires Sync. A window's segment tree shares its states between threads as combine sources. A state type holding a Cell or RefCell must switch to atomics or a Mutex, or keep that data outside the state.
  • Documentation: bind-time arguments are seen before the cast to the parameter type (ScalarBindInfo::argument); a Safety clause on data_chunk_from_arrow (validity bitmaps are read one byte past their rows); Known Limitations entries for aggregate states DuckDB never destroys, abandoned streams and out-of-memory aborts.
  • Value::as_str, display_string and Debug render only types whose every payload DuckDB can render; everything else (VARIANT, GEOMETRY, a type quack-rs does not know, and ARRAY / UNION values that could hold one) gets UNRENDERABLE. See Fixed.
  • Breaking: scalar function bind and init callbacks take their own argument types, RawScalarBindInfo and RawScalarInitInfo (#[repr(transparent)] over duckdb_bind_info / duckdb_init_info), in ScalarBindFn, ScalarInitFn, scalar_bind_callback!, scalar_init_callback!, ScalarBindInfo::new and ScalarInitInfo::new. A table function's bind and init callbacks receive the same C types, but DuckDB casts them to a different, larger struct, so table_bind_callback! output registered on a scalar function (which safe code could do) wrote past the scalar bind info when it reported a panic. The two kinds no longer type-check in each other's slots (compile_fail doctests). A hand-written raw scalar callback changes its parameter type only.
  • AggregateFunctionBuilder::ffi_state::<T>() and AggregateOverloadBuilder::ffi_state::<T>() install FfiState<T>'s state_size, init and destructor callbacks together. Set one by one, a size callback for one T with an init callback for a larger one wrote past DuckDB's allocation; the setters remain, and the aggregate register methods and Registrar now state that pairing as a Safety obligation (the Registrar docs said a call on the entry point's connection was "always sound"). The README and prelude examples use the new method; the prelude's had no destructor, so it leaked every state.
  • Breaking: ArrowConvertedSchema::from_raw takes the Arrow schema the handle was built from (&ArrowSchema) instead of a column count, and is no longer const: data_chunk_from_arrow checks each array against that schema's shape (see Fixed).
  • data_chunk_to_arrow refuses a chunk holding a value DuckDB would export as a different value (see Fixed).
  • Breaking, not named until now (found by cargo semver-checks against 0.16.0, run as a patch release so that it reports every break): Appender and Connection are no longer RefUnwindSafe (they gained a Cell and a RefCell in earlier passes: the appender's row bookkeeping and the collision check's catalog snapshot), and validate::description_yml::DescriptionYml is #[non_exhaustive], so it cannot be built with a struct literal. For catch_unwind, wrap a closure that captures an Appender or a Connection in AssertUnwindSafe.
  • Breaking: a callback macro's body must have type () (except cast_callback!'s, which returns the cast's bool). The body runs inside catch_unwind, and its value was discarded: a body that used ? compiled, and the error it returned was dropped without being reported. It is now a type error; report the error through the callback's set_error.
  • Breaking: a callback macro's body is no longer an unsafe context. The body was a closure inside the generated unsafe extern "C" fn, and a closure inherits its function's unsafe context, so a raw-pointer dereference or an unsafe fn call in a "safe" callback compiled with no unsafe keyword and, on edition 2021 (which the scaffold generates), no warning. The body is now expanded as a nested ordinary fn; wrap unsafe operations in unsafe blocks, as the documented examples already do. compile_fail doctests pin it.
  • Breaking: catalog lookups are refused in a catalog DuckDB does not implement itself. For a catalog a storage extension attaches, duckdb_catalog_get_entry starts that extension's transaction and runs its schema lookup with no try, so an exception there would abort the process. CatalogEntry::lookup and Catalog::get_entry now return an error for any catalog whose type is not "duckdb". DuckDB's own catalogs (the database's, temp and system) all have that type and are unaffected. Look up entries in another catalog with SQL instead, as the error says.
  • CombineFn documents that combine must leave its source states unchanged (see Fixed), and no longer advises moving out of them.
  • Catalog::type_name no longer gives "system" as an example type: the system and temp catalogs are of type "duckdb".
  • data_chunk_from_arrow's Safety section requires a fixed-width dictionary's values buffer to be readable one element past its length when the indices can be NULL: DuckDB points NULL indices at a sentinel entry there, and the flattening copy reads it (found with an AddressSanitizer-built libduckdb; item 29). Its Errors section now lists the layout refusals.
  • Documentation corrections from a mechanical check of the docs' universal and numeric claims: the Arrow conversion functions are already in DuckDB 1.4.4's C API (not added in 1.5.0); the unstable API region did not change "in every recent release" (1.4.4 and 1.4.5, 1.5.0 and 1.5.1, and 1.5.2 to 1.5.5 are byte-identical); it was not changed by middle insertions "in four of the last four" versions (two).
  • LESSONS.md and the book's pitfall catalogue gained L15 (combine must leave its source states unchanged), L16 (a valid Arrow array is not always one DuckDB imports correctly), L17 (a COPY … FROM reader must not declare result columns) and L18 (a LIST reserve moves every buffer below its child): 30 documented pitfalls. The README's and the book's summary tables, which stopped at L14, list all 30.
  • DestroyFn, FfiState and the book's known limitations document that DuckDB never destroys one aggregate state per row of a window frame with EXCLUDE (item 35: 5000 of a 5000-row window, every release from 1.4.4). DestroyFn's docs said it was called for every state DuckDB created.
  • query::prepare and the book's known limitations document that an allocation failure inside duckdb_prepare can hand back a statement DuckDB has already freed, which the error path then reads and frees (item 36: the last 3 of its allocations, every release from 1.4.4, found by failing each allocation in turn).
  • ClientContext::config_option documents that DuckDB's function has no try, and why no built-in setting's getter throws through it in practice.
  • FfiInitData::set and FfiLocalInitData::set name the callback each must be called from: both store through duckdb_init_set_init_data, so the wrong one sets the other kind of init data, which the other get then reads as the wrong type. FfiBindData::set says a typed table function's bind sets the bind data itself. ReplacementScanBuilder::register says DuckDB calls delete_callback even when extra_data is null, unlike its other destructor slots (a new test observes the call).
  • data_chunk_from_arrow says which column holds its claim on the Arrow array: DuckDB gives release to column 0 alone, so a vector made to reference another column (reference_vector) does not keep the producer's buffers alive (a new test observes both cases).
  • StructWriter's write_*, set_null and set_valid state in their Safety sections that no field writer was replaced through field_mut: swapping two writers is safe code, and the next write would go to the wrong child vector.
  • OwnedConnection's Send justification said DuckDB forbids concurrent use of one connection; it serialises it. The comment now gives the real reasons, and a test queries and drops a connection on another thread.
  • docs/upstream-duckdb-reports.md gained items 20 to 37, and item 16 gained a HUGEINT reproducer.
  • Test gaps the mutation sweeps exposed. The full sweep left 22 mutants alive, and the end-to-end run over the files mutants.toml excludes left more. Each is now killed by a test, excluded with the reason in mutants.toml, or recorded in AUDIT.md section 10 as equivalent with the argument. The new tests cover FfiState's tag structure, a release profile without panic = "unwind", an overload's own return type, every typed Value getter against a value of its own type, the TIMETZ and TIME_NS range guards, the Appender methods a no-op replacement survived (another schema, column_type, clear_columns, append_default_to_chunk), StructWriter's child vectors, InMemoryDb::execute's row count, QueryResult::result_kind for a statement that returns nothing, a created ErrorData, MockVectorWriter::len, and append_metadata's handling of = in a path, a lone -, signed version numbers, commit hashes of the wrong length, case or alphabet, and a footer whose magic field is wrong. tests/handle_leaks.rs kills the Drop mutants that survived every functional test. Two comparisons moved into small const fns so that a unit test can reach their boundary: the aggregate-state salt and the appender's u32::MAX-inclusive length limit.
  • The ABI layout table covers DuckDB v1.5.6, which shares the 546-slot layout of v1.5.2 – v1.5.5. Without it a duckdb-1-5 extension was refused by v1.5.6 as UnknownEngineVersion under the default AbiPolicy::Strict. v1.5.6 declares every slot stable for extensions targeting C API v1.5.6; quack-rs still targets v1.2.0, which v1.5.6 loads.
  • scripts/check-abi-table.py reads the v1.5.6 header correctly: it evaluates the new DUCKDB_API_VERSION_AT_LEAST(...) bands, no longer folds a preprocessor continuation line into the first declaration's fingerprint, and checks that the first STABLE_API_SLOT_COUNT declarations are identical in every release (the v1.4.0 varint → bignum rename aside) instead of requiring an unchanged stable count, which v1.5.6 raised from 357 to 404.
  • CI runs what the release workflow runs before a tag is pushed. The first v0.18.0 tag failed on Windows while main was green, because no PR job ran the duckdb-1-5* tests on macOS or Windows: test-bundled now also runs cargo test --all-targets --all-features on all three, and on Linux cargo doc --all-features, the only rustdoc pass over the test-only items. The test job runs the unit tests with Cargo's default release codegen (16 codegen units, no fat LTO), the build that exposed the FfiState salt defect (see Fixed); every other build is debug or has one codegen unit and fat LTO. The release workflow's test steps and test-bundled's pass --no-fail-fast, so one failing test binary no longer hides failures in the ones after it.
  • release.yml: gh release create --verify-tag, so a missing tag fails instead of being created at the default branch's HEAD; a re-run of publish recognises current cargo's "already exists on crates.io index" (it matched only the older "already uploaded"); every job has a timeout-minutes.
  • The unit tests run on wasm32, the one 32-bit target quack-rs supports: the wasm job installs emsdk 6.0.10 and runs cargo test --lib under node, with and without duckdb-1-5-4. Before, it only compiled for wasm32, so the 32-bit size limits were only ever tested with 64-bit values. Two tests assumed a 64-bit target and were corrected. criterion is now a non-wasm dev-dependency (its rayon dependency does not build for wasm32).
  • tests/file_handle_close.rs tests FileHandle's Drop when the close fails, by swapping stubs into the C API dispatch table; it fails on the old Drop, which destroyed without closing first.
  • tests/ffi_roundtrip/file_errors.rs runs on every platform: a failed write through a read-only handle, and a seek past i64::MAX, which is asserted to be refused with InvalidInput. Only the /dev/full sync failure, which has no portable trigger, stays Linux-only. The over-4 GiB string test runs on macOS as well as Linux.
  • .gitattributes: text files are LF in every checkout, so Windows CI tests the same bytes as Linux; fuzz seeds and images are binary.
  • The book's aggregate examples and hello-ext register state with ffi_state::<T>() instead of wiring FfiState's three callbacks by hand.
  • The book was reviewed page by page against the source. Corrected, among others: pages that said every extension binary is tied to one DuckDB release and should pin libduckdb-sys exactly (true only for builds that use the unstable API), an out-of-bounds ChunkWriter example, the instance cache's lifetime (it holds weak references), AbiPolicy::Warn (it prints to stderr), and links that rendered as literal brackets.
  • The site's custom <head> was never rendered (book.toml did not point mdBook at book/theme), so every page shipped an empty description and no canonical or Open Graph tags. scripts/seo-postbuild.py, run by docs.yml and by CI's book job, now writes a canonical URL and a specific description to each page and fails on a missing or duplicate one; the preview image is a PNG; headings are no longer rendered at half opacity; diagrams follow the theme; mermaid is pinned to 11.17.2.
  • The crate, README, book, logo and social-preview image no longer call quack-rs "production-grade" or the logo "FAST"; neither was backed by evidence, and the audits behind this release found serious defects. The crates.io description is now "Rust SDK for building DuckDB loadable extensions on the DuckDB C Extension API". A new README Status section and the FAQ state what the record shows instead: pre-1.0 API churn, the defects each audit found (AUDIT.md), DuckDB's own limits, and what CI checks.

Fourth audit

  • CI's "the refused extension registered nothing" checks could not fail (duckdb -c stops at the failed LOAD); they feed the statements on stdin and require a marker row.
  • AbiPolicy handling is a pure policy_verdict, tested against every check result. The docs of enforce_abi_policy no longer say Warn goes through set_error.
  • Arrow export documents the INTERVAL and UHUGEINT values DuckDB corrupts.
  • docs/upstream-duckdb-reports.md gained items 7 to 19.

Earlier passes (0.17.0, second and third audits)

  • Breaking: CopyFunctionBuilder is no longer Send or Sync. Supporting COPY … FROM gave it two new fields that each carry a raw pointer — copy_from: Option<TableFunctionHandle> (a duckdb_table_function) and extra_info: Option<ExtraInfo> (a *mut c_void) — and a raw pointer is neither Send nor Sync. cargo-semver-checks classifies this as auto_trait_impl_removed, a major break, one of the reasons this release bumps the minor version: for a pre-1.0 crate Cargo treats the leftmost non-zero component as the major, so the minor position is where a break goes (RELEASING.md, "Semantic versioning policy").

    The change aligns the type with its three siblings rather than making it an outlier — ScalarFunctionBuilder, TableFunctionBuilder and AggregateFunctionBuilder were already !Send + !Sync in 0.16.0, each because it holds a raw DuckDB handle. A builder is a short-lived object constructed and registered inside duckdb_init_c_api, and no DuckDB C API accepts one from another thread, so the traits were never usable for anything real. Code that moved a CopyFunctionBuilder between threads must now construct it on the thread that registers it.

  • Flaky tests: a shared counter raced across parallel tests. The extra_info tests and the new arrow ones each reset one static AtomicUsize and then asserted on it, but cargo test runs tests in parallel — so one test's reset could land between another's drop and its assertion. dropping_an_untransferred_extra_info_frees_it failed in CI while the identical job on the identical commit passed. Each test now owns its counter: extra_info carries it in the allocation, and arrow carries it in the record's own private_data, which is what that field is for. Verified with 100 repeat runs, zero failures.

  • 128-bit splitting and reassembly in Value and PreparedStatement was open-coded. Value::as_i128 / as_u128 / as_uuid / as_decimal / uuid and PreparedStatement::bind_i128 / bind_u128 / bind_decimal each did their own << 64 / >> 64 word arithmetic, where being silently wrong is easy — a shift in the wrong direction still compiles, still round-trips zero, and still round-trips anything that fits in 64 bits. Those call sites now route through four pub(crate) helpers (hugeint_from_i128 / hugeint_to_i128 / uhugeint_from_u128 / uhugeint_to_u128) that live next to each other and are unit-tested in both directions, at the extremes and against hand-built records. Mutation testing confirms all seven shift mutants across the three modules are now killed. The appender, datetime and the vector reader and writer still split and reassemble 128-bit values themselves.

  • Test gaps the mutation sweep exposed, once it could see the files. Chief among them the shift direction in hugeint_from_i128 / uhugeint_from_u128: swapping >> for << still compiles, still round-trips zero and still round-trips anything that fits in 64 bits, so nothing in the suite noticed. Also TypeId::composite_constructor_hint's per-variant arms, LogicalType::check_slot's rejection path, composite_message, LogicalTypeError::api_func, the arrow accessors against a populated record rather than only an empty one, and map2_str's NULL propagation when just one argument is NULL.

    With the three gate defects fixed and the two FFI-wrapper modules excluded, the incremental sweep over this branch's 40 changed source files reports 404 mutants — 250 caught, 154 unviable, none missed. What survives after that is annotated in the source with #[mutants::skip] and a reason, rather than filtered out of sight: the bare FFI reads (DataChunk::size / column_count, ScalarBindData::set, ScalarLocalState::set), two Drop impls whose effect is only visible in freed memory or to a leak checker (OwnedVector, SecretEntry), the deprecated FfiBindData::get_from_bind whose mutant is the function (it returns None unconditionally, because DuckDB has no duckdb_bind_get_bind_data), map2 / map2_str whose per-row NULL check only runs inside DuckDB's expression executor, and Value::as_str_or_default, whose null-handle answer is exactly the mutant's String::new().

    Two of those turned into real work rather than an annotation. SecretEntry's zeroize-on-drop is a security property the crate advertises and nothing asserted — its body is now SecretEntry::zeroize_in_place, tested directly, with Drop left as a one-line delegation. And hugeint_to_i128 combined its halves with |; because the halves occupy disjoint bits, | → ^ cannot change the result for any input, so that mutant was unkillable by construction. The halves are added instead: identical here, incapable of overflowing (i64::MIN << 64 is exactly i128::MIN, and the round trips at the extremes would panic in a debug build if that were wrong), and - or * in its place dies at once.

  • The mutation-testing gate always reported 100% and always passed. cargo mutants --output DIR writes its results to DIR/mutants.out/, so --output mutants.out put them in mutants.out/mutants.out/. The report step counted mutants.out/caught.txt and friends, found nothing, computed SCORE=100% from TOTAL=0, and skipped its exit 1 because MISSED was also 0. The run on this PR printed MUTATION SCORE: 100% and passed while cargo-mutants' own summary line in the same log read 229 missed, 238 caught, 218 unviable. Both jobs now pass --output ..

  • Incremental mutation testing skipped every top-level src/*.rs. The job selected changed files with the pathspec src/**/*.rs, but git's default wildmatch lets * cross /, so that pattern requires at least one directory component after src/ and matches no top-level file at all. src/value.rs, src/query.rs, src/appender.rs and every other module directly under src/ were silently excluded while the job reported success. It now filters to .rs in the shell.

  • The mutation gate's Display/Debug filter matched nothing. mutants.toml carried exclude_re = ["^fmt::"], commented "Display/Debug impls". cargo-mutants matches that regex against the whole line it prints for a mutant — src/x.rs:1: replace <impl core::fmt::Debug for T>::fmt -> … with … — which begins with the file path, so a pattern anchored at fmt:: can never match and seventeen unkillable Debug::fmt -> Ok(Default::default()) mutants survived every sweep. cargo-mutants' own documented form, impl Debug, does not match this crate either: it writes impl core::fmt::Debug for T, and the qualified path lands in the mutant name. The pattern is now impl [a-z:]*Debug for. Display is deliberately not excluded — those impls render error text that tests assert on, so their mutants die and belong in the gate.

    The incremental job now also re-applies exclude_re from mutants.toml as CLI flags, the way it already did for exclude_globs: cargo-mutants only combines a CLI --exclude-re with the config file from 27.0.0 onwards, and that job passes one.

  • src/value.rs and src/query.rs are excluded from mutation testing, and their pure logic moved out so that it is not. Every function left in those two files wraps a DuckDB C call, which the gate's --cargo-arg=--lib run — no live engine — structurally cannot reach; that is the same rationale mutants.toml already carried for nine sibling modules, and between them the two files accounted for 206 of the 220 survivors. Excluding them wholesale would have swallowed the pure code too, so it moved to siblings that stay in the gate: src/value/hugeint.rs (the four 128-bit word helpers), src/value/defaults.rs (the fourteen as_*_or accessors, as a second inherent impl Value) and src/query/cstr.rs (to_c_sql, c_str_to_owned). Seven of the as_*_or accessors had no unit test at all, and nothing exercised c_str_to_owned's non-null path — all of them were among the 220 — so nine new tests turn those survivors into kills rather than hiding them. No public item moved: the re-export keeps every existing path.

  • Four CI jobs never actually ran. rust-toolchain.toml pins channel = "stable", and a rustup toolchain file overrides the default dtolnay/rust-toolchain sets — so the miri, leak-check, fuzz and nightly jobs all resolved a bare cargo to stable. Miri and LeakSanitizer failed loudly (the 'miri' component ... is not available for the 'stable-...' toolchain; the -Z flag is only accepted on the nightly channel); the informational nightly job failed silently, re-testing stable. All four now invoke cargo +nightly explicitly, with a comment saying why.

    The test-bundled-prebuilt job's clippy step was also missing the env: block its sibling test step has, so build.rs panicked looking for duckdb.hpp before clippy ran.

  • cargo doc with -D warnings failed. Six intra-doc links were broken or redundant — QueryResult::column_logical_type, scalar::typed, vector::ops (twice), TableFunctionBuilder::build_handle — and vector::ops carried a doc comment on both the pub mod declaration and the file's own //! header, so its module docs were resolved in the parent's scope and every link into the module failed with no source location. Three more links in appender and table_description pointed at duckdb-1-5-gated methods and so broke the default-feature doc build. cargo doc is now clean under -D warnings on all four feature sets.

  • Pitfall L9 — duckdb_data_chunk_from_arrow takes the array even when it fails. duckdb.h reads like a success-path statement; arrow-c.cpp nulls arrow_array->release inside the per-column loop, before the work that can throw. Guessing either way gives you a bug — double release, or a leaked Arrow buffer tree when a zero-column schema means the loop never runs. Documented in LESSONS.md and the book, and made impossible by the by-value signature.

    The pitfall count in the README, the crate docs and the FAQ was stale at 17, and the book's catalogue was missing L8 and L9. All four now agree with LESSONS.md, and an integration test fails when a documented count differs from it (see Added).

  • A copy function may now implement COPY … FROM alone. CopyFunctionBuilder::register used to require bind, sink and finalize, which made a read-only format impossible even though duckdb_register_copy_function accepts one: it decides what a copy function supports from info.sink != nullptr and copy_from_bind != nullptr independently. Leaving all three unset is now valid when copy_from is set; setting only some of them is an error that says so. Existing COPY … TO functions are unaffected.

  • CI gains four jobs: miri (546 unit tests under the interpreter), leak-check (LeakSanitizer over the end-to-end suite against a real libduckdb, now leak-clean), fuzz (cargo-fuzz over the description.yml parser, the duckdb_string_t decoder and the validators) and semver (cargo-semver-checks). tests/ffi_roundtrip.rs is now linted — it is feature-gated, so the plain clippy job had been compiling it away to nothing.

  • extension-load is now a matrix over DuckDB v1.4.4, v1.5.0, v1.5.5 and latest, rather than one releases/latest run that silently retargeted whenever DuckDB shipped. That is the README's compatibility claim, proven rather than sampled.

  • New end-to-end coverage for nested types as scalar-function input — LIST offsets that are cumulative rather than uniform, NULL elements inside a list, a MAP key miss, and an ARRAY's fixed stride across rows. Writing them was already covered; reading them does raw offset arithmetic against a layout only DuckDB defines, which a mock cannot check.

  • AUDIT.md records the full review: what was read, what was probed, what was verified correct, and what is still open.

  • CI: the AddressSanitizer job is now blocking, as planned when it was added. It passed on main and over the merged 125-test end-to-end suite with no reports and no suppressions.

  • CI: the informational beta clippy job now also lints the end-to-end tests, with bundled-test-prebuilt,duckdb-1-5-4 against a pre-built libduckdb. Those tests compile only with bundled-test-prebuilt, so beta's new assert_is_empty lint fired on four of them while the job stayed green.

  • Mutation testing: the configuration moved to .cargo/mutants.toml. At the repository root cargo-mutants never read it — so the full sweep ran without its exclusions or features (2,164 mutants listed instead of 1,348) — and it carried two keys cargo-mutants 27.1.0 rejects (cap_timeout, jobs). Its examine_globs is gone too: once the file was read, that key overrode the incremental job's --file flags instead of being narrowed by them.

  • Breaking: TableFunctionBuilder::with_state requires S: Clone + Send: the state bind returns is a template and every execution scans a fresh clone. Use the new with_bind_init for state that cannot be cloned.

  • Breaking: TypedTableFunctionBuilder::projection_pushdown is removed — the typed scan closure cannot learn the projection, so enabling it returned the wrong columns. Use the raw TableFunctionBuilder for pushdown.

  • Breaking: FfiBindData::set / FfiInitData::set require T: Send + Sync, FfiLocalInitData::set T: Send; ReplacementScanBuilder::register_with_data and Connection::register_replacement_scan_with_data require T: Send + Sync.

  • Breaking: every scalar Value getter returns Option<T> (as_i8 … as_u128, as_f32, as_f64, as_bool, the date/time/timestamp family, as_interval, as_uuid, as_decimal; the new as_enum_index does too): None for a null handle, SQL NULL, a non-scalar value or a failed cast. A failed cast used to return a sentinel (T::MIN, NaN) indistinguishable from a real value. The as_*_or(default) forms keep their signatures and now cover all four cases.

  • Breaking: Value::as_blob accepts only a BLOB (DuckDB's cast of anything else to BLOB could throw) and errors on SQL NULL; Value::as_str errors on SQL NULL.

  • Breaking: FileSystem<'ctx> borrows its ClientContext. Code that stores a FileSystem needs a lifetime parameter, and the ClientContext must outlive it.

  • Breaking: DuckDbErrorType gains Autoload, Sequence and InvalidConfiguration (40–42), which used to map to Invalid.

  • Breaking: SelectionVector::new returns Result<Self, ExtensionError>.

  • Breaking: datetime::date_to_days, time_from_micros, time_tz_from_bits, timestamp_from_micros, timestamp_to_micros, time_tz_bits and decimal_to_f64 return Option.

  • Breaking: VectorWriter::set_null / set_null_range (and so DataChunk::propagate_nulls) on a STRUCT or ARRAY vector also null the row's fields / elements, recursively, as DuckDB's FlatVector::SetNull does. To reuse such a row, call set_valid on it first, which restores what set_null nulled below it, then write any field or element NULLs.

  • Breaking: SqlMacro::to_sql emits double-quoted identifiers (CREATE OR REPLACE MACRO "add"("a", "b") AS (a + b)). Calling the macro is unchanged: DuckDB resolves quoted identifiers case-insensitively.

  • ScalarFunctionBuilder::varargs, varargs_logical and volatile no longer require duckdb-1-5: both C functions are in the stable v1.2.0 API.

  • Appender is re-exported from the prelude without a feature, matching the module, and use quack_rs::prelude::* now brings entry_point! / entry_point_v2! into scope, as the prelude's own documentation said.

  • AggregateFunctionSetBuilder::overloads is unchanged, but the builder its closure receives is now named AggregateOverloadBuilder, for symmetry with ScalarOverloadBuilder. The old name remains as a deprecated type alias (quack_rs::aggregate::builder::OverloadBuilder) and still compiles.

  • AggregateOverloadBuilder is exported from quack_rs::aggregate and from the prelude. The old OverloadBuilder was reachable only at quack_rs::aggregate::builder::, and had no public constructor, so a caller could not build one outside an overloads closure.

  • AggregateOverloadBuilder moved out of set.rs to src/aggregate/builder/overload.rs.

  • Breaking: QueryResult::next_chunk returns Result<Option<OwnedDataChunk>, ExtensionError>. duckdb_fetch_chunk returns null both at the end of the rows and when the fetch fails, and next_chunk returned None for both, so a streaming result that failed part way — a runtime error in the query, or another statement run on the same connection — read as a complete, shorter result. The error DuckDB recorded is now returned as Err, and keeps being returned on later calls; a clean end stays Ok(None).

  • Breaking: Value::time_ns and Value::timestamp return Result<Value, ExtensionError> and refuse a payload outside the range DuckDB's own SQL produces; so do the new Value::time, time_tz, timestamp_tz, timestamp_s, timestamp_ms and timestamp_ns (see Added). DuckDB stores any 64-bit payload unchecked, and rendering an out-of-range one crashed or aborted the process (Value::time_ns(i64::MAX), Value::timestamp(i64::MIN)) or printed garbage. Value::date and the new Value::interval are infallible: DuckDB renders every value of those types.

  • Breaking: CatalogEntry::lookup and Catalog::get_entry return Result<Option<_>, ExtensionError>: Ok(None) is "not found", Err is "refused". A Type or Collation lookup of a name that makes DuckDB autoload an extension (inet, json, an ICU collation name) is refused without calling DuckDB while the autoload_known_extensions setting is on: a failed autoload inside duckdb_catalog_get_entry aborted the process, and a successful one loaded an extension as a side effect of a lookup. The entry types that were already refused (Schema, Database, PreparedStatement, Invalid) are now an Err too, instead of a None that looked like "not found".

  • Breaking: FileSystem::open returns FileHandle<'_>, which borrows the FileSystem. A FileHandle kept after its database was closed read freed memory (valgrind: an invalid read in duckdb::FileHandle::Read). FileHandle::from_raw returns a handle whose lifetime the caller chooses.

  • Breaking: InitInfo::projected_column_index returns Option<usize>, None past the end of the projection, where DuckDB returns 0 — a real column index. CopyBindInfo::column_type returns Option<LogicalType>, bounds-checked like BindInfo::result_column_type; it used to wrap the null handle DuckDB returns for an out-of-range index.

  • Breaking: TypedTableFunctionBuilder::build returns an error when projection_pushdown(true) was set on the TableFunctionBuilder before with_state / with_bind_init. That sequence bypassed the typed builder's no-pushdown rule, and the scan returned the wrong columns (SELECT b got column a's values).

  • Breaking: ConfigOptionBuilder::register refuses an option with no default value (DuckDB then reports it as an unrecognized configuration parameter), an option of a pseudo-type (ANY, SQLNULL, INTEGER_LITERAL, STRING_LITERAL) or of a type the typed Value constructors cannot build (BIGNUM, GEOMETRY, VARIANT, …), and a name that duckdb_settings() already lists, compared case-insensitively and including aliases: an option named threads used to register and shadow the built-in setting. The default is now handed to DuckDB already typed (see Security).

  • Breaking: TableFunctionBuilder::register refuses a name that already belongs to a table function or table macro, and CopyFunctionBuilder::register a format name that already exists (csv, parquet, …). DuckDB dropped such a registration while reporting success, so the function was never called. The C API cannot overload table functions.

  • Breaking: ScalarFunctionBuilder::register and ScalarFunctionSetBuilder::register (and so map1, map2 and the other typed constructors) refuse a parameter signature that duckdb_functions() already lists under that name, built-ins included. DuckDB merges a scalar registration into the existing entry with override on, so the new overload silently replaced the old one for every query in the database — after registering abs(BIGINT), abs(-5::BIGINT) returned 995 — or, with a different return type, made every call ambiguous. Signatures containing STRUCT, UNION, ENUM or a literal pseudo-type are registered unchecked, as documented on both builders.

  • Breaking: SqlMacro::register refuses a body that holds more than one statement, counted with DuckDB's own parser, and executes nothing. The body 1); DROP TABLE t; SELECT (1 used to register and drop the table. A scalar body that ends in a -- comment, which used to comment out the closing parenthesis, now works.

  • Breaking: Appender no longer loses rows silently. DuckDB's appender cannot take back a value, so a row() closure that failed after appending part of a row left the row half-written, and close() then returned Ok and wrote none of the buffered rows. Such a row() now poisons the appender: every later append_*, row, end_row, flush and close returns an error saying how many buffered rows were not written. close() with a row started but not ended is an error; every mutating method is an error after a successful close() (DuckDB accepted appends after close and wrote them at the next flush); append_chunk in the middle of a row is refused. With duckdb-1-5, clear() resets DuckDB's state and clears the poison. A value rejected first in its row loses nothing and does not poison.

  • Breaking: the LogicalType STRUCT and UNION constructors (struct_type, struct_type_from_logical, union_type, union_type_from_logical and their try_ forms) apply the rules DuckDB's binder applies to the same types in SQL: names unique ignoring ASCII case (empty names exempt) and at most MAX_UNION_MEMBERS union members. The C API checks nothing, so a scalar returning such a type registered and then failed every call with "duplicate name in struct". The try_ forms return an error; the others panic.

  • Breaking: VectorWriter::write_varchar / write_blob and StructWriter::write_varchar / write_blob panic for a value longer than MAX_STRING_LEN instead of storing a truncated one (see Fixed). Inside scalar_callback! and the typed scalar constructors the panic becomes a SQL error; try_write_varchar / try_write_blob return it instead.

  • Breaking: MockVectorWriter behaves like a real output vector, so a test that passed for a callback that is wrong in DuckDB now fails. set_null followed by a write_* leaves the row NULL (set_valid undoes a NULL); a row never written is valid, not NULL (is_written tells a test whether the loop wrote it); a write past capacity panics instead of growing the mock; a string over MAX_STRING_LEN panics, as VectorWriter now does. The docs no longer claim one function can be called with both the mock and the real writer: the types differ.

  • Breaking: generate_scaffold validates the free text it writes into the generated files: description and maintainer must be non-empty, must not start or end with whitespace and must not contain control or bidirectional characters (maintainer must also be one line), github_repo must be owner/repo, and git_ref a commit hash or tag. Those fields went into description.yml unquoted, so Fast: analytics, Analytics #1 or a newline produced a file that parsed to something else or not at all; they are now YAML double-quoted scalars. Each line of a multi-line description gets its own //! prefix; the second and later lines used to fall outside the doc comment and fail to compile.

  • Breaking: validate_spdx_license rejects nesting deeper than 64 parentheses. It recursed once per ( with no limit, so a license field of a million ( overflowed the stack and aborted the process, reachable from parse_description_yml on untrusted input.

  • TypeId::TimeNs, Any, Varint, SqlNull, IntegerLiteral and StringLiteral no longer require duckdb-1-5. All six exist in DuckDB 1.4.4, this crate's floor, and 1.4.4 produces TIME_NS and BIGNUM columns, so with default features LogicalType::get_type_id panicked on those columns.

  • Breaking: TypeId::Varint.sql_name() returns "BIGNUM", DuckDB's name for the type since 1.4, instead of "VARINT". Both names parse as SQL, but code that compares the returned string sees a different value.

  • SecretEntry's Debug output shows [REDACTED] for a non-empty scope: the scope (a bucket or URL prefix) is zeroized on drop as sensitive, but Debug printed it.

  • InMemoryDb checks, before first use, that the bundled-test-prebuilt C++ shim was compiled against headers whose duckdb_ext_api_v1 has the size the libduckdb-sys bindings expect, and panics naming both slot counts, DUCKDB_LIB_DIR and cargo tree -i libduckdb-sys if not. With a DuckDB 1.5.0 library under 1.10505 bindings, a slot the older headers lack was left as stack garbage and the tests ran on; larger headers would overrun the buffer.

  • scripts/check-abi-table.py lists every upstream vX.Y.Z tag from v1.2.0 on instead of a hard-coded list, which is how v1.4.5 went missing from the layout table (see Security).

  • RELEASING.md: a release is done only once the tag is on origin, crates.io's newest version is the Cargo.toml version and docs.rs has built it, each with a command to check; version snippets are bumped in the release pull request itself.

  • Internal layout, with no change to any public path: src/value.rs is split into value/composite.rs, value/nested.rs, value/scalars.rs, value/temporal.rs and value/temporal_checks.rs; the LogicalType constructors moved to types/logical_type/construct.rs; and ScalarOverloadBuilder moved to scalar/builder/overload.rs. src/query.rs, src/appender.rs, src/arrow.rs and src/testing/mock_vector.rs are split the same way into private submodules.

  • Every unsafe block in library code now states, in a // SAFETY: comment, the invariant it relies on and why it holds; 149 did not. clippy::undocumented_unsafe_blocks is enabled so CI keeps it that way (test code is exempt).

  • docs/architecture.md matches the crate again: its module table listed neither abi, arrow, callback, chunk_writer, datetime, query, secrets, tls nor warning, and said appender, table_description, ScalarFunctionBuilder::varargs and volatile need duckdb-1-5 (they do not). An integration test now fails when the table and src/lib.rs disagree.

Fixed

Fifth audit

  • On a 32-bit target the Arrow layout walk's refusal of an oversized row count named vector::ops::MAX_CAPACITY (2^28 - 1) as the limit while refusing exactly that many rows: a list child's reserve rounds up to a power of two, so the limit is 2^27. The message now states the effective limit. Found by the first run of the unit tests on wasm32.
  • Rendering a value aborted the process for values ordinary SQL builds: a VARIANT holding an out-of-range timestamp, a DECIMAL(38, 0) holding i128::MIN (from sum over two in-range values), and a GEOMETRY built from malformed WKB each made DuckDB's cast to text throw through the C API; a DECIMAL(38, 38) holding 1.2 rendered with an unwritten first byte, and aborted too when that byte was not valid UTF-8. The render guard is now an allow-list, and DECIMAL payloads are checked against their width.
  • ListBuilder aborted the process on a large list of a wide type. DuckDB's ceiling is 2^37 bytes per child buffer, not elements, checked after it rounds the reservation up to a power of two, so a BIGINT list of 2^34 + 1 elements, or an INTEGER[1000] list of 2^25 + 1, reached a reserve that throws through the C API. The builder respects the rounded byte ceiling; a row past it is NULL.
  • Aggregate states DuckDB moved were never dropped (unreleased fourth-audit code; 0.16.0 had no tag). That pass's FfiState tag was derived from the slot's address, and radix repartitioning copies states to new rows, so every moved state was skipped and its T leaked (8 of them after one ungrouped query on eight threads in the regression test). The tag follows the slot's contents.
  • FfiState's type tag was salted with the address of T's type-name string (unreleased fifth-audit code), which is not unique: with more than one codegen unit and no fat LTO (Cargo's default release profile) the init, access and destroy callbacks could see different addresses, so with_state found no state and destroy skipped every one: in such a build every FfiState aggregate returned NULL or garbage (10 of the 279 end-to-end tests fail that way). The salt is now a hash of TypeId::of::<T>().
  • Aggregate states DuckDB never destroys leaked a box each. A grouped aggregate's states that a stopped scan never reached (a LIMIT above it, an error, an interrupt) are never destroyed by DuckDB 1.4.4 to 1.5.5; under LIMIT 10 over 300,000 groups, 297,952 boxed Ts leaked. A small T is now stored in DuckDB's own state bytes, so it leaks nothing unless it owns heap memory itself.
  • ListBuilder::with_element_limit called after the first row could raise the limit past DuckDB's ceiling for the child type, which the builder applies once, at the first row; a later row past the ceiling then reached duckdb_list_vector_reserve and aborted the process (a BIGINT row of 2^34 + 1 in tests/ffi_roundtrip/list_limits.rs). The ceiling is kept apart and always applies. On a 32-bit target a row of more than 2^31 elements under a larger limit made the reservation 0 (next_power_of_two overflowing) while the closure still wrote the row, past the child; the reservation is now the limit there.
  • TableDescription::column_name, column_type and column_has_default aborted the process for index u64::MAX on DuckDB 1.5.0 to 1.5.5: the C API converts the index to an optional_idx, whose constructor throws for that value outside any try (upstream item 30). They return None for it without calling DuckDB, as for any other index past the last column.
  • data_chunk_to_arrow exported different values without an error. An INTERVAL of more than i64::MAX / 1000 microseconds wrapped when DuckDB converted it to nanoseconds; a UHUGEINT of 2^127 or more came out negative; a HUGEINT or UHUGEINT of 39 digits was exported as a decimal128(38, 0) it does not fit. Each is now refused, at any nesting depth; a HUGEINT is accepted when arrow_lossless_conversion exports it as a 16-byte binary.
  • data_chunk_to_arrow could return an array that contradicts its schema. Before 1.5.5, BIGNUM (and from 1.5.0 GEOMETRY) exported under arrow_output_version = '1.4' are written as binary views while the schema declares plain binary; a consumer reads the views as offsets, and appending DuckDB's own re-import of the batch crashed on 1.4.4 to 1.5.4 (item 34). The export is now checked against the declared schema and refused on a mismatch.
  • data_chunk_from_arrow imported valid Arrow arrays from the wrong rows, or read and wrote out of bounds. DuckDB mishandles offsets below the top level: a struct inside an offset struct or list, a union's members, a run-end-encoded array's value validity, and a dictionary's validity under a list. It also mishandles overlapping or gapped list views, dictionaries whose values are dictionary-encoded, sparse unions whose type codes are not 0, 1, …, and a run-end-encoded array where it reads a plain one. A dictionary-encoded array of more than 2048 rows under a struct with NULL rows had its validity copied past a 2048-row heap mask; the regression test aborted with glibc's corrupted size vs. prev_size. Each layout is refused, naming the node, and the neighbouring layouts DuckDB does import correctly are still accepted (tests/ffi_roundtrip/arrow_layout.rs).
  • data_chunk_from_arrow accepted three more layouts DuckDB imports wrongly, found by the fifth audit's review of the new layout walker: a fixed-size list with NULLs whose child is a STRUCT with a dictionary field of more than 2048 rows (the list's NULLs are broadcast into the struct and reach the 2048-row mask of item 25; valgrind reports 156 invalid accesses in SetInvalid past the 256-byte mask on 1.5.5); a dictionary with null_count = -1, whose NULL rows came back as values (item 31); and a sparse union with a nonzero null_count, whose type ids were read as validity, so every row came back NULL (item 32). Each is refused, as is such a layout below a node DuckDB converts as zero rows (the walk used to stop there, though DuckDB still expands a run-end-encoded descendant; valgrind showed its values' validity read 11 bytes past a 1-byte bitmap on 1.5.5, item 24). So is a geoarrow.wkb column read as more than 2048 rows: DuckDB 1.5 copies its storage into a 2048-row vector before the cast to GEOMETRY (SIGSEGV on 4096 rows, item 33).
  • A null-typed field below the top level of an imported Arrow column read as valid after its first row: DuckDB imports it as a constant vector, which the readers index as flat. The column is now flattened whenever its type holds NULL at any depth, not only at the top.
  • The CombineFn documentation advised a combine that gave wrong window results. It said to move out of the source states. A window's segment tree combines one state into every frame that covers it, so a combine that consumed its source gave 4985 of 5000 rows wrong over ROWS BETWEEN 100 PRECEDING AND CURRENT ROW (pinned by tests/ffi_roundtrip/agg_window.rs). It now says to leave the source unchanged.
  • A typed table function answered with the wrong column under projection pushdown switched on after build(). build refused pushdown switched on before with_state, but not on the raw builder it returns, and SELECT b then returned column a's value. Registering such a builder (and MockRegistrar::register_table) now fails.
  • A typed table function's callbacks could be replaced after build() with the raw builder's safe bind, init, local_init and scan setters (or extra_info), after which the typed trampolines read one type's data as another's: an init callback setting a u64 as init data made the scan lock it as a Mutex<S>. Registering such a builder (and MockRegistrar::register_table) now fails; both extra_info Safety sections now say the pointee must be what the installed callbacks read.
  • A row closure that panicked after its first value left the appender unpoisoned, so finishing the row by hand committed a row half written by the closure that panicked. It is poisoned, as an error there poisons it.
  • Dropping a FileHandle could abort the process when the close threw: duckdb_destroy_file_handle calls Close() with no try. Only an extension's file system can throw there; DuckDB's local close cannot. The drop now closes through duckdb_file_handle_close, which catches, and destroys the handle only if that succeeded; after a failed close it leaks the handle rather than retry the close unguarded.
  • FileHandle::seek clamped a position past i64::MAX to i64::MAX, which a file system that accepts that offset (tmpfs) took, so the call returned Ok at a position the caller never asked for. It is now an InvalidInput error.
  • append_metadata named the wrong default platform on OpenHarmony (*-linux-ohos): DuckDB appends _musl there, as for any musl-based Linux, and the tool did not.
  • MockRegistrar still accepted builders whose types the real registration refuses: a composite or literal TypeId in any parameter, varargs, return, named-parameter, cast or config-option slot, and an ANY return type. Its module doc said these checks need DuckDB; they do not. Each builder now runs one sequence of type checks, with DuckDB when registering and without it in the mock, so the messages are the same.
  • Four writer contracts named only a LIST/MAP vector's direct child as moved by a reserve on that LIST/MAP. VectorWriter::from_vector, StructWriter::new, StructVector::field_writer and ValidityBitmap::ensure_writable now say the same about every STRUCT field and ARRAY element vector below that child, down to the next LIST or MAP: DuckDB 1.5.5 reallocates all of their data and validity buffers (tests/ffi_roundtrip/nested_reserve.rs measures which move).
  • Arrow import: a fixed-size list format DuckDB accepts could bypass the layout checks. DuckDB reads the size in +w:N with std::stoi, so +w:2x, +w: 2 and +w:+2 are fixed-size lists to it; quack-rs parsed the size strictly, took them for leaves and skipped every check below them. A fixed-size list → struct → dictionary layout the checks refuse under +w:2 crashed the process (SIGSEGV) under +w:2x. The size is now parsed as stoi parses it.
  • Arrow import: offsets below the top level were not checked for being negative, and their sums could overflow (a panic in a debug build). Every node's length and offset must be non-negative, and a sum past i64::MAX is refused.
  • ArrowArray::release and ArrowSchema::release called a producer's callback again at drop when the callback did not null itself, as the Arrow specification requires it to; they now null it themselves.
  • entry_point! and entry_point_v2! aborted the process when an argument expression panicked. $policy and $register were evaluated in the generated extern "C" function before the panic guard ("panic in a function that cannot unwind"); they are now evaluated under it, and the load fails with the panic's message.
  • A typed table function used as a COPY … FROM reader could invalidate the database. Its bind declares columns, and under COPY … FROM DuckDB appends each to the INSERT's own expected types, so every chunk reaching the table is too wide: a DuckDB built with assertions fails chunk.ColumnCount() == types.size() and invalidates the database; a release build drops the column. The typed bind now fails with a message when a column is declared there (upstream item 37).
  • VARCHAR and BLOB values were decoded as little-endian. duckdb_string_t holds its length and pointer in the target's own byte order; DuckStringView, read_duck_string and read_duck_blob read both as little-endian, so on a big-endian target an inline string's length read as length << 24 and its inlined bytes were followed as a pointer (reproduced under Miri with --target s390x-unknown-linux-gnu). Both are now read natively, and the pointer is read as a pointer at the target's width, which also keeps its provenance. No DuckDB extension platform is big-endian, and nothing changes on a little-endian one; bytes passed to DuckStringView::inline_from_bytes are now read in native order too.
  • The Arrow export check read HUGEINT / UHUGEINT rows with their halves swapped on a big-endian target. DuckDB stores them as {lower, upper}; the check read the 16 bytes as a native i128, so on s390x it passed 10^38 (read as about 1.27 * 10^37) into a lossy export and refused 2^63. It now reads DuckDB's struct (reproduced under Miri on s390x).
  • On a 32-bit target (wasm32), ListBuilder and OwnedVector::new could ask DuckDB for a buffer whose size wraps. DuckDB computes a buffer's size as a 64-bit idx_t and passes it to malloc, which narrows it to a 32-bit size_t unchecked. The limits bounded the element count by usize::MAX, not the bytes, so a BIGINT or VARCHAR list child could reserve 2^32 elements and malloc receive 0 bytes, and OwnedVector::new(HUGEINT, 2^28) succeeded with the same wrap. The list limit now also fits one allocation, and vector::ops::MAX_CAPACITY is 2^28 - 1 on a 32-bit target. 64-bit targets are unchanged.
  • data_chunk_from_arrow passed on lengths no vector can hold. DuckDB sizes the chunk from the array's length before its error handling starts, and sizes each list child it reserves from the list's element count; a run-end-encoded column declares any length with a few bytes of buffers. On a 32-bit target a length of 2^28 (VARCHAR) or 2^29 (BIGINT) had its byte size narrowed by malloc, and the import then wrote past the buffer; on any target a length past 2^37 threw through the C API. The layout walk now refuses, at any node, a row count whose next power of two exceeds vector::ops::MAX_CAPACITY (the rounding a list child's reserve adds).
  • A run-end-encoded child of a zero-row list crashed the Arrow import when the list's offset was not 0. DuckDB treats a list it converts as zero rows as empty whatever its offsets say, and reads an empty list's child as a plain array (upstream item 24): a run-end-encoded child there is read from buffers it does not have (SIGSEGV on every release from 1.4.4). The layout walk refused this only when the list's offset was 0, and a valid array (an inner list under an empty outer row, whose offsets need not start at 0) got through. It now decides as DuckDB does, by the row count.
  • A dictionary-encoded child of a fixed-size list crashed the Arrow import when the fixed-size list was a MAP value. DuckDB verifies a map by flattening its entries, and flattening an ARRAY flattens its child over the child vector's capacity, which a map allocates at the chunk's row count rather than the entries it holds; the dictionary child's selection vector was built only for the entries converted, so the flatten read it out of bounds (SIGSEGV on 1.5.0-1.5.2, an AddressSanitizer heap-buffer-overflow on 1.5.5; upstream item 24). The layout walk refuses a dictionary-encoded fixed-size-list child under a map; the same shape at the top level or under a plain list, which is not over-flattened, still imports.

Fourth audit

  • A C API aggregate in a running window returned the wrong answer. Without a destructor DuckDB streams it and re-reads the first row (sum-like: 1 2 3 4 5 for 1 3 6 10 15). Every aggregate now registers one (a no-op when none is given).
  • Arrow import with a nonzero parent offset imported the wrong rows (DuckDB ignores the offset for values); it is refused.
  • ListBuilder overwrote rows already in the list vector when it started on a non-empty one; it appends after them.
  • VectorWriter::set_valid left a nested row's children NULL; it restores them when the row was NULL.
  • Scalar collision check. It missed signatures containing a type alias, could be shadowed by a user macro named like the catalog functions it queries (it now qualifies them with system.main.), and accepted overloads that DuckDB's binder finds ambiguous with varargs. Breaking: such overlapping overloads are refused at registration. Connection keeps a snapshot of existing scalars, so 300 registrations through Registrar take 18.6–35.6 ms instead of 5.2–6.3 s (release build, three runs each).
  • Keyword names. Breaking: validate_function_name refuses the 53 keywords (DUCKDB_UNCALLABLE_KEYWORDS) that DuckDB cannot call unquoted — coalesce(x) silently ran the built-in — and SqlMacro parameters use the new validate_parameter_name (79 keywords).
  • Catalog entries outlived their handle. Breaking: CatalogEntry copies the name and type at lookup and owns nothing afterwards.
  • Secret scopes were parsed from a string; they are read as a list, NULL-safely, from system.main.duckdb_secrets().
  • CopyGlobalInitInfo::get_file_path truncated at a NUL. Breaking: it returns Result<String>; get_file_path_bytes returns the raw bytes.
  • interval_to_micros reported overflow for totals that fit (an intermediate sum overflowed); it computes the exact total in i128.
  • Appender: a failed automatic flush (every 204,800 rows) poisoned the appender with a false "half-written row" message; the row counts as ended and the constraint error is reported as it is.
  • On DuckDB 1.4.x, registering a scalar under any existing name (a new overload of abs, say) failed with no reason, because the C API registers with CREATE before 1.5.0; quack-rs refuses it first and says why. And LogicalType::try_new(TypeId::TimeNs) returned an INVALID type there (the 1.4.x C API does not know TIME_NS); it is an error now, as is any type id the running engine hands back changed.
  • LogicalType::register blamed a taken name when the type contained ANY; it names the cause. try_decimal validates width and scale itself (DuckDB does only from 1.5.4). Breaking: try_array(_, 0), an empty union_type and an empty Value::array_value are refused, as in SQL.
  • MockRegistrar accepted builders the real registration refuses. Breaking: it runs the same checks (missing callback or return type, empty function set, incomplete copy function, config option without type or default) and records nothing on failure.
  • description.yml: a quoted, flow or block value on the line after its key kept its quotes; values that PyYAML (YAML 1.1) reads as a boolean, number, date or null draw a warning; the scaffold quotes name, github and ref. The "text after a comment" error named the line the value started on rather than the line of the comment that ended it.
  • Scaffold: the generated lib.rs failed cargo fmt --check for names of 4 characters or fewer or 40 or more; the generated CI's Linux SQLLogicTest step was skipped by extension-ci-tools.
  • append_metadata refuses a platform group (linux) and a wasm_* platform without --wasm, and recognises a footer with an empty ABI field.
  • validate_semver refuses numeric pre-release identifiers with leading zeros; validate_spdx_license names the canonical spelling for a case-only difference.
  • ABI refusal for a development engine no longer suggests a declaration the build ignores; build.rs warns about a malformed QUACK_RS_TARGET_DUCKDB_VERSION.

Earlier passes (0.17.0, second and third audits)

  • Pitfall L8 — DEFAULT_NULL_HANDLING does not propagate NULLs for scalar functions. quack-rs documented that DuckDB "automatically returns NULL if any argument is NULL, without your function callback being called". For a scalar function registered through the C API that is false at run time: CAPIScalarFunction calls the callback for every row including NULL ones and never inspects the result's validity, and the only NULL check in ExpressionExecutor::Execute is VerifyNullHandling, whose entire body is inside #ifdef DEBUG. A callback that ignores validity therefore returns a non-NULL answer for a NULL input, silently, in every release build.

    SELECT f(NULL) still returns NULL — a literal NULL is constant-folded before the function is reached — which is why the bug survives review. From a column it does not. New DataChunk::propagate_nulls / any_null restore SQL semantics in one line; the new typed constructors (see Added) get it right by construction; the docs, the book chapter and LESSONS.md now state what DuckDB does, with the source quoted. A regression test pins the behaviour. Aggregates are no different: their update receives NULL rows under either setting too (Pitfall L12; see Added and the aggregate entry below).

  • Composite TypeIds silently produced an invalid type. duckdb_create_logical_type "returns an invalid logical type" for DECIMAL, ENUM, LIST, STRUCT, MAP, ARRAY and UNION — a non-null handle wrapping LogicalTypeId::INVALID, so the existing null check never fired. .param(TypeId::Struct) failed much later with a message that named neither the parameter nor the fix, and get_type_id() on one panicked. New TypeId::is_composite / composite_constructor_hint; LogicalType::new asserts, try_new errors, and every builder validates before allocating any DuckDB handle.

  • extra_info leaked when a builder was not registered. DuckDB only takes ownership at duckdb_*_set_extra_info; a dropped builder dropped the pointer. This reached users through APIs that never mention a pointer — TableFunctionBuilder::with_state boxes two closures. Found by Miri.

  • Two stale-borrow bugs in src/secrets.rs's own tests, which took a pointer into a String, called a &mut method, then read through the stale pointer. The library's zeroize_string was correct throughout.

  • The reference example disabled every panic guard in the crate. examples/hello-ext/Cargo.toml shipped panic = "abort" — the setting validate_release_profile rejects outright and the scaffold refuses to generate, because it makes every catch_unwind in quack-rs inert. CI built that example, loaded it into a real DuckDB, and held it up as the way to do this. Two new tests hold the example and the scaffold's generated profile to quack-rs's own validator, so the generator and the validator cannot drift.

  • The ScaffoldConfig example in the README and two book pages did not compile. The struct gained three fields and the exhaustive literals were never updated; rustdoc examples are compiled by cargo test but Markdown code fences are not, so the copy a new user reaches for was the broken one. All now use ..ScaffoldConfig::default().

  • Registration failures now name what to check. duckdb_register_*_function reports failure as a bare DuckDBError with no message. There are exactly three causes, and a name collision with a DuckDB built-in (list_sum, array_sum, …) looks identical to a type error. The message names all three and points at SELECT * FROM duckdb_functions() WHERE function_name = '<name>'.

  • Wrong answers, no error:

    • A NULL row of a STRUCT result kept its fields valid, so (f(x)).a returned the stale field value instead of NULL.
    • A typed table function failed the second time a plan ran (PREPARE … ; EXECUTE p; EXECUTE p;, or a recursive CTE): init moved the state out of bind data DuckDB reuses for every execution.
    • Value getters cast the value in place: Value::double(1.5).as_i32() turned the value into DOUBLE 2.0. Getters now read a private copy.
    • BindInfo::add_result_column with a type containing ANY/INVALID was dropped by DuckDB, shifting every later column; it is now a bind error.
    • datetime::time_tz_bits silently corrupted an out-of-range offset.
  • Duplicate overload signatures in a scalar or aggregate function set are rejected at register, naming both overloads, instead of registering and then failing every call with "Could not choose a best candidate function".

  • Expression::fold returns Err for a non-foldable expression instead of Ok with a null-handle Value.

  • CastFunctionBuilder::register leaked extra_info when DuckDB rejected an ANY/INVALID type; those are now rejected before anything is handed over.

  • ReplacementScanInfo::set_error("") was ignored by DuckDB, so the query fell through to "table does not exist".

  • SqlMacro with a SQL keyword as a name or parameter produced a parser error.

  • Documentation that was false against the code or DuckDB: set_max_threads (it does not need local_init); DbConfig::set accepts unknown option names; Pitfall L4 (a skipped ensure_validity_writable silently drops the NULL rather than segfaulting); the book's panic = "abort" advice (it must be "unwind", which the crate itself enforces); several README and book examples that did not compile; and the stable ABI prefix, which is ABI-identical since v1.2.0 but not byte-identical (two slots were renamed varint → bignum in v1.4.0).

  • ClientContext's constructors now state that the context must not outlive its connection: DuckDB's wrapper holds a reference, not an owner.

  • Pitfall L10 (LESSONS.md, book/src/reference/pitfalls.md) — scalar bind data is dropped when DuckDB copies a bound expression. The book's pitfall summary table was also missing L8 and L9; all three rows are now there.

  • ScalarBindInfo::set_bind_data_copy — scalar bind data was silently lost whenever DuckDB copied a bound expression. CScalarFunctionBindData::Copy() populates the copy's bind data only if a copy callback is registered, and quack-rs never exposed the setter, so get_bind_data could return null on a copied expression: a wrong answer, not a crash.

  • An invalid mutants.yml that silently disabled the mutation gate. A shell comment added earlier in this branch wrote out an empty workflow expression while explaining not to pass values that way. GitHub evaluates expressions anywhere in the file, including inside shell comments, so the empty one invalidated the whole workflow:

    Invalid workflow file: .github/workflows/mutants.yml
    (Line: 203, Col: 14): An expression was expected
    

    This fails silently by design: GitHub records a run with zero jobs and the workflow stops running. mutants-incremental therefore stopped executing on pull requests while every other check stayed green — the same "a gate that is not actually running" failure this release is otherwise about, introduced by this branch rather than found in it.

    scripts/check-workflow-expressions.py now runs in the doc job and rejects this class. It fails on the exact commit that broke and passes on the fix. Neither yaml.safe_load nor a JSON-Schema check catches it, because both treat the run: block as an opaque string.

  • 22 assert!(x.is_empty()) / assert!(!x.is_empty()) assertions that beta clippy's new assert_is_empty / assert_is_not_empty lints reject, across 10 files. These are not style noise: clippy-beta was running stable clippy before this release, so it had never reported them, and when beta promotes to stable the blocking clippy job inherits every one. Each now uses assert_eq! / assert_ne! against an empty value, which is what the lint asks for and prints the actual value on failure.

  • Four CI quality gates were testing nothing, each verified against the files rather than inferred:

    • The MSRV job (ci.yml) and the release gate's MSRV entry ran a bare cargo check after selecting 1.86.0. rust-toolchain.toml pins channel = "stable" and a toolchain file overrides the rustup default, so both ran stable. Now cargo +1.86.0 check.
    • clippy-beta ran stable clippy for the same reason.
    • Miri ran with default features, cfg-ing out every duckdb-1-5* module — including src/arrow.rs, the largest block of pure-Rust unsafe here.
    • Doctests were never compiled: every invocation used --all-targets, which excludes them. All 183 passed when the gate was added.
  • release.yml still used fail-fast: true, the setting LESSONS.md blames for a release that shipped with two platforms broken.

  • MUTANTS_EXIT captured tee's status rather than cargo-mutants'.

  • SECURITY.md recommended panic = "abort"; that makes catch_unwind inert and disables the crate's entire panic-containment mechanism. Cargo.toml has always set unwind, and validate_release_profile rejects abort.

  • RELEASING.md Step 1 listed 9 check names against 30 CI jobs, and .github/workflows/README.md listed 14. A maintainer following either could tag with Miri, LeakSanitizer, osv-scan or semver red. The workflow README now carries a table generated from ci.yml, with the generator inline.

  • An aggregate's update receives NULL rows; the documentation said it did not. NullHandling, the null_handling setter on both aggregate builders, the null-handling book page and the first-extension tutorial said that DuckDB skips NULL rows before an aggregate's update unless SpecialNullHandling is set. DuckDB's CAPIAggregateUpdate passes every row of the chunk to update, NULL rows included, under DefaultNullHandling and SpecialNullHandling alike; for an aggregate, DuckDB reads the setting only when decorrelating a correlated subquery. An update that reads a row without checking its validity reads a meaningless value for every NULL row. The docs now say so, the new Pitfall L12 gives the symptom and the fix (skip rows whose is_valid is false), and the end-to-end test aggregate_update_receives_null_rows_under_either_null_handling pins the behaviour. L12 also records that, under either setting, a count-like aggregate in a correlated subquery returns NULL, not 0, for an outer row with no match; DuckDB rewrites that NULL to 0 only for its own count.

  • Wrong answers, no error, third pass:

    • Once a row exceeded ListBuilder's child-capacity ceiling, push_row / push_map_row wrote no list entry for it or for any later row of the chunk. DuckDB reuses output vectors, so those rows returned the previous chunk's lists as valid values. A refused row is now NULL.
    • VectorWriter::write_varchar / write_blob stored a value longer than u32::MAX bytes as a truncated string, for VARCHAR possibly cut inside a UTF-8 sequence: a 4 GiB + 1 byte value came back one byte long.
    • A panic in a cast_callback! body under TRY_CAST left the output vector as it was, so TRY_CAST returned zeros or a previous query's values rather than NULL: DuckDB ignores a cast's return value in TRY mode and nulls only the rows passed to set_row_error. The macro now marks every row of the chunk as an error.
    • Two overloads of a scalar set that differ only in their varargs type, such as f(BIGINT) and f(BIGINT, BIGINT...), were rejected as duplicates. The varargs type is now part of the signature, as it is in DuckDB.
  • A cast_callback! body that returned false without calling set_error failed a regular CAST with Conversion Error: and no text. The macro now sets callback::CAST_FAILED_WITHOUT_MESSAGE before running the body; the body's own set_error replaces it. Hand-written cast callbacks are unchanged.

  • set_error("") on a table function, cast or copy function reached the user as an error with no text after the prefix (Binder Error: , Conversion Error: ); the new EMPTY_ERROR_PLACEHOLDER is reported instead.

  • An error message containing a NUL byte lost everything after it on some paths and kept it on others. Every path now replaces the NUL with ?: set_error on scalar, aggregate, table, cast, copy and replacement-scan functions, ExtensionError::to_c_string, ErrorData::new, and the entry point's report of a failed registration.

  • Expression::fold errors carried DuckDB's exception serialized as JSON ({"exception_type":"Conversion","exception_message":…}) with the type InvalidInput. They now carry the plain message and the matching DuckDbErrorType (Conversion, OutOfRange, …).

  • With duckdb-1-5, a typed table function whose bind declares no column fails with an ordinary bind error; DuckDB raised an INTERNAL Error with a C++ stack trace.

  • FileFlag::CreateNew did not create a file: it set only DuckDB's exclusive flag, which is ignored without the create flag, so an existing file opened and a missing one failed. set_flag(FileFlag::CreateNew, true) now also sets Create. On Windows an existing file is still opened rather than refused: DuckDB's Windows file system ignores the exclusive flag, which is now documented on FileFlag::CreateNew.

  • WarningCollector dropped every warning, silently, once a panic had poisoned its lock. It now recovers the lock and keeps working.

  • ScalarFunctionBuilder::varargs / ScalarOverloadBuilder::varargs with a composite TypeId (TypeId::List, …) panicked inside the setter. register now refuses it with an error naming the varargs slot.

  • ScalarFunctionSetBuilder and AggregateFunctionSetBuilder check every overload for a return type and its required callbacks before creating any DuckDB handle, and the error names the overload index; the scalar set used to report "overload missing function callback" with no index.

  • AbiPolicy::Strict's refusal of an unknown engine recommended AbiPolicy::AllowUnknownEngine; it now recommends a build that uses only the stable C API. The layout-mismatch message no longer advises rebuilding a duckdb-1-5 extension against a pre-1.5 engine.

  • append_metadata rejected a 32-byte footer value, which DuckDB reads in full without a terminating NUL, and it now also accepts --option=value.

  • examples/parse_descriptions.rs, pointed at a community-extensions checkout (extensions/<name>/description.yml), found no files and reported "0 parsed, 0 rejected" with exit status 0. It now reads that layout, fails when it finds no files, and exits non-zero when any file is rejected.

  • Documentation that was false against the code or DuckDB, third pass:

    • Aggregates: the state_size callback runs whenever an operator sizes a state buffer, not once at registration; finalize runs once per result batch; destroy also runs on the source states of combine; combine targets are initialised by the init callback (StateInitFn; T::default() for FfiState<T>), not zeroed; update receives one state pointer per input row, not per group.
    • Scalars: the null_handling setters repeated the claim Pitfall L8 disproved (that DuckDB skips the callback for NULL arguments). Identical calls share bind data: for SELECT f(i), f(i) the bind callback runs twice, but both columns use the data from its first run, so a bind that reads a counter, clock or RNG needs volatile. The register methods now state DuckDB's collision rules: an aggregate cannot take a name already used by a scalar function, aggregate or macro, while a scalar merges into an existing function of the same name (an identical signature is now refused, see Changed). The entry-point docs say that registration is not transactional: functions registered before the registration closure fails stay registered.
    • Table and copy functions: set_cardinality's is_exact works the other way round from duckdb.h; local_init does not enable parallelism, only set_max_threads does; table extra_info must point to Send + Sync data; CopyBindInfo::options' shape is now described.
    • Casts: set_row_error requires row < count, because DuckDB checks that bound only in debug builds; the cast docs said false becomes NULL under TRY_CAST, which holds only for rows passed to set_row_error.
    • Queries and values: a multi-statement string passed to query() / execute() runs every statement and returns the first row-producing statement's result; every prepared parameter has a name (a positional ? is named by its position); DbConfig::get_flag returns the extension's name, not a description, for an extension setting; check_valid_utf8 agrees with std::str::from_utf8 rather than being stricter.
    • Dates and intervals: date_from_days of infinity gives 5881580-07-11; DuckDB's 30-day month applies to interval comparison and epoch_us, not to interval arithmetic or epoch; DuckInterval's Eq compares fields, so 1 month differs from 30 days here although SQL calls them equal.
    • InstanceCache::get_or_create returns an error for a different config on an already-open database (the docs said the config was ignored), and never caches in-memory paths. FileSystem's set_flag(flag, false) does not clear a flag. SqlMacro documents where the macro is created, that it persists in a database file and can replace a user's macro or shadow a built-in; its parameter errors say "parameter name".
    • Table macros read a table parameter through query_table(tbl): the SqlMacro examples wrote FROM tbl, which DuckDB binds at creation time to a table literally named tbl, so the module example failed to register.
    • Loading: the getting-started pages and Pitfall P3 said to LOAD a bare .so, which every supported DuckDB refuses; they now show the append_metadata step and duckdb -unsigned, and no recipe sets allow_extensions_metadata_mismatch any more, since a correctly stamped C_STRUCT build loads without it. Pitfall P2's symptom was wrong: a bad -dv is stamped without complaint and LOAD refuses the file.
    • The book's Known Limitations page suggested approximating window semantics with aggregate functions, the exact shape that crashes every C API aggregate (Pitfall L11); that page, the README and the hello-ext README now warn about it. The hello-ext README also stopped telling readers to verify panic = "abort" and to add duckdb with bundled as a dev-dependency (Pitfall P9).
    • The installation page's MSRV rationale, SECURITY.md's claim that TlsConfigProvider enforces TLS 1.2+ by default (min_tls_version is a required method), the entry-point page's macro expansion, the crate's and FAQ's pitfall counts, and the README's validated-fields table (an unlisted SPDX id is a warning, not a rejection).
    • Book examples that did not compile or did not do what they said: among them a DuckStringView example using the deprecated from_bytes (it returns None for any string over 12 bytes), a my_bind that leaked the duckdb_value from duckdb_bind_get_parameter, a README description.yml example that did not compile, and scaffold examples that, run as written, overwrote the current package's Cargo.toml and src/lib.rs. Broken links in the book are fixed, and hand-kept test counts, several of them wrong, are removed from the repository trees.
  • # Safety contracts that allowed undefined behaviour (found while documenting every unsafe block; contract text only, no signature or behaviour change):

    • FfiLocalInitData::get / get_mut, FfiInitData::get_mut and FfiBindData::get_from_init / get_from_function did not require T to be the type passed to set; set::<u8> then get_mut::<[u64; 64]> met every stated clause and wrote 512 bytes through a 1-byte allocation. The init-data getters also did not bound the returned lifetime by the scan call. Both requirements are now stated.
    • MapVector::set_size did not require the size to equal the entries written, so a larger size made DuckDB read past the child vectors.
    • read_duck_blob did not require the vector to outlive the returned slice, which borrows the vector's own buffer for a blob of 12 bytes or fewer.
  • Documented: when an aggregate's finalize reports an error, DuckDB 1.5.5 does not destroy every state the query created (ungrouped: 2 initialised, 1 destroyed; grouped: 4 and 2), so whatever those states own leaks, and for an FfiState<T> whose T is boxed (see Changed) the box too. Nothing in an extension can detect it; Known Limitations and AggregateFunctionInfo::set_error now say so, and an end-to-end test pins it. Two of this release's own tests leaked on these paths and made the LeakSanitizer job fail; both now use states that own nothing.

  • cargo doc failed with default features on a link to PreparedStatement::execute_streaming, which needs duckdb-1-5. CI now also builds the docs with default features, the set a dependent crate documents.

  • Cargo.toml's duckdb-1-5 description gave the requirement as libduckdb-sys >= 1.5.0 (the crate is versioned 1.10500.0) and claimed the feature has no effect against 1.4.x, which was never checked; the hello-ext README said varargs and volatile need duckdb-1-5.

Security

Fourth audit

  • An Arrow import with a large dictionary wrote past a heap buffer. data_chunk_from_arrow on a dictionary-encoded array with NULLs and more than 2048 entries — a LIST of 1025+ two-element dictionary-encoded lists is enough — made DuckDB overflow a validity mask (valgrind: invalid write in GetValidityMask; SIGSEGV or SIGABRT on 1.4.4, 1.5.0 and 1.5.5). Such arrays are refused before DuckDB is called.
  • An Arrow import read out of bounds. A dictionary-encoded or null-type column came back as a dictionary or constant vector, which every quack-rs reader reads as flat: wrong values, then reads past the buffer. Those columns are copied into flat vectors before data_chunk_from_arrow returns.
  • Rendering a timestamp from SQL aborted the process. make_timestamp(-9223372036854775808) passed to a table function, then Value::as_str, display_string or {:?}: DuckDB's rendering threw through the C API (exit 134, 1.4.4 to 1.5.5), also inside a LIST, STRUCT or MAP. Every temporal payload is checked first; as_str returns Err(UNRENDERABLE). as_time and the other converting getters no longer hand such a payload to DuckDB's cast, which overflowed a signed multiply.
  • A scalar bind callback inspecting a subquery argument aborted the process on DuckDB 1.5.0 to 1.5.4. Those releases copy the argument outside any try; SELECT f((SELECT 1)) threw a C++ exception through the callback. Breaking: ScalarBindInfo::argument now asks for nothing on those releases and fails the bind with an explanation; get_argument's Safety section states the requirement.
  • An out-of-range TIME or timestamp crashed DuckDB later. PreparedStatement::bind_time / bind_timestamp / bind_timestamp_tz and Appender::append_time / append_timestamp accepted any payload; i64::MIN as a TIME segfaulted when rendered, and at the append itself into a VARCHAR column. They are refused (Breaking). bind_str, bind_blob and append_bytes refuse more than 4 GiB, which DuckDB stored modulo 2^32 (a 4 GiB + 3 byte blob became 3 bytes).
  • A panic in a scalar function's bind or init callback aborted the process. New scalar_bind_callback! / scalar_init_callback! macros (duckdb-1-5) catch it and fail the query with its message; a panic with an empty message reports EMPTY_PANIC_PLACEHOLDER.
  • A null get_api or access from DuckDB panicked or crashed the entry point. init_extension checks both before libduckdb-sys unwraps them.
  • A literal type in a registration invalidated the database. TypeId::StringLiteral / IntegerLiteral as a parameter, return or registered type: the first query that used it raised an internal error and every later query failed. LogicalType::try_new refuses both.
  • FfiState destroyed states init never ran on. After a state_init error, DuckDB passes never-initialised states to destroy; each state now carries a tag derived from T's TypeId (and, for a boxed T, the box's address) that destroy checks (Breaking: see the FfiState<T> entry under Changed).
  • SecretEntry left secret bytes in spare capacity after truncation; zeroisation now covers the whole allocation.

Earlier passes (0.17.0, second and third audits)

  • A panicking Drop in extension state aborted the process. Every FFI destructor quack-rs generates — FfiState<T>::destroy_callback, FfiBindData / FfiInitData / FfiLocalInitData::destroy, replacement_scan::drop_box, TypedCallbacks::destroy_extra — dropped a Box<T> of arbitrary user data directly inside an extern "C" fn. Since Rust 1.81 an unwind across that boundary is a guaranteed process abort. Reproduced against DuckDB 1.5.4: an aggregate whose state type has a panicking Drop killed the process with SIGABRT from inside duckdb::RowOperations::DestroyStates, on a task-scheduler thread. All of them now run under the new callback::catch_ffi_panic, which is public so extensions writing their own extern "C" destructors get the same containment.

    Where DuckDB offers an error channel the panic is now reported instead of swallowed: CAPIAggregateStateInit checks the error flag and throws, so a panicking Default::default() becomes an ordinary SQL error rather than a silent NULL. The state destructor has none (CAPIAggregateDestructor takes no info and returns nothing), so there the message is discarded.

    FfiState::init_callback also no longer forms a &mut Self over the possibly-uninitialised allocation DuckDB hands it.

  • Soundness: safe code could corrupt memory, race, or read freed memory.

    • FileSystem held a raw ClientContext pointer with no lifetime, so it could be used after its connection closed and read freed memory (confirmed under valgrind). Breaking: FileSystem<'ctx>.
    • SelectionVector::new exposed uninitialised memory through the safe as_slice() (stale 0xDEADBEEF observed), and a large length made DuckDB compute a wrapped allocation size, so safe indexing segfaulted. Breaking: it returns Result, rejects lengths above MAX_LEN before DuckDB is called, and zeroes the buffer.
    • datetime::decimal_to_f64 read past DuckDB's powers-of-ten tables for scale > 38.
    • FfiBindData, FfiInitData and replacement-scan data are shared across threads by DuckDB. Breaking: FfiBindData::set / FfiInitData::set / register_with_data require T: Send + Sync, FfiLocalInitData::set requires T: Send.
  • Process aborts from ordinary input. Each of these let a DuckDB C++ exception unwind into Rust ("Rust cannot catch foreign exceptions"), killing the host process; each is now validated in Rust first and reported as an error or None:

    • every Value::as_* getter (and the _or forms) on a SQL NULL — e.g. a table function called with f(n := NULL); a null handle was dereferenced;
    • datetime::date_to_days on an invalid date, timestamp_from_micros / timestamp_to_micros on infinities and the far-negative range;
    • datetime::time_from_micros / time_tz_from_bits on a time outside 00:00:00–24:00:00, in a DuckDB built with assertions (a debug build, as the bundled-test feature compiles), which fails Time::Convert's D_ASSERT; a release build returned out-of-range fields;
    • SelectionVector::new above DuckDB's allocation limit;
    • a config option whose default does not cast to its type;
    • catalog lookups for Schema, Database, PreparedStatement and Invalid entry types;
    • a panic whose payload's own Drop panics (panic_any(value)), in every callback macro, catch_ffi_panic, the typed scalar and typed table trampolines and the entry point's registration guard;
    • the entry points dereferenced a NULL duckdb_database* when DuckDB's get_database failed.
  • Documented, not fixable here: C API aggregates crash under agg(x) OVER () and agg(x ORDER BY y). DuckDB's CAPIAggregateUpdate does not flatten the state vector, and the window-constant and sorted-aggregate executors pass a one-element state array with count > 1, so every aggregate registered through the C API — not only quack-rs's — reads out of bounds. Reproduced in plain C against DuckDB 1.4.4, 1.5.0 and 1.5.5; reported upstream as duckdb/duckdb#26109. Documented on AggregateFunctionBuilder, AggregateFunctionSetBuilder, FfiState, the aggregate book pages and as Pitfall L11.

  • The one active advisory suppression is gone, because the crate behind it is. RUSTSEC-2026-0235 (rkyv 0.7.46) was suppressed in osv-scanner.toml, reachable only as quack-rs → duckdb → rust_decimal → rkyv. In duckdb 1.10505.0 rust_decimal became an optional dependency (it was required in 1.10504.0), and quack-rs does not enable it — so neither crate is in Cargo.lock any more. Both the suppression and the CI step that re-proved it have been removed.

  • cargo deny was scanning the wrong dependency graph. deny.toml had graph.all-features = false; since duckdb is optional and the crate declares no default feature, neither the advisory scan nor the license scan ever evaluated duckdb, arrow, chrono or rust_decimal. Now all-features = true.

  • ci.yml and mutants.yml — the two workflows that build and execute pull-request code — had no permissions: block, so the token inherited the repository default. Both now take contents: read.

  • mutants.yml interpolated a PR-derived file list straight into a run: block, where $(...) expands before bash parses the script. Moved to env:.

  • persist-credentials: false on all 48 actions/checkout steps.

  • Soundness, third pass: more ways safe code, or a malformed input, could corrupt or read freed memory.

    • A FileHandle could outlive its database and read freed memory. Breaking: FileHandle<'fs> (see Changed).
    • DuckDB v1.4.5 was missing from the ABI layout table, so a duckdb-1-5 build loaded into it was treated as an unknown engine instead of a layout mismatch, and under AbiPolicy::AllowUnknownEngine it loaded and segfaulted. v1.4.5 is now in the table, so such a build is reported as a layout mismatch there, as it is on v1.4.4.
  • Process aborts, third pass. Each of these killed the host process:

    • ClientContext::config_option on an option whose value is NULL: enable_profiling before it is set, or any option after SET <option> = NULL. It now returns None.
    • A catalog Type or Collation lookup of a name that makes DuckDB autoload an extension, when the autoload fails. Breaking: now refused (see Changed).
    • A config option default that SQL's TRY_CAST accepts but DuckDB's built-in cast does not, such as a TIMESTAMPTZ default with a time-zone name while ICU is loaded. ConfigOptionBuilder::register now converts the default through the connection's own TRY_CAST and hands DuckDB a Value of the option's type, so no cast runs inside the C API.
    • Value temporal getters whose cast DuckDB implements by throwing: for example as_time() of 'infinity'::TIMESTAMP (reachable from a table function's named parameter) or as_timestamp_ns() of any TIMESTAMP after 2262. They now return None where DuckDB would throw, and also for a result outside the target type's range, which DuckDB could produce but not render.
    • Rendering a temporal Value built from an out-of-range payload. Breaking: the constructors now validate (see Changed).
    • The AbiPolicy::Warn diagnostic used eprintln!, which panics when writing to stderr fails, outside the entry point's panic guard. A failed write now loses the warning instead.
    • validate_spdx_license on deeply nested parentheses (stack overflow). Breaking: nesting is capped at 64 (see Changed).
  • SqlMacro::register ran every statement in a body. A body of 1); DROP TABLE t; SELECT (1 registered and dropped the table. Breaking: a body with more than one statement is now refused (see Changed).

  • SecretEntry::with_field on an existing key, with_provider and with_scope freed the value they replaced without zeroizing it (and with_field also the duplicate key), although Drop zeroizes those same values. Each now zeroizes before replacing. A test that inspects every buffer as it is freed covers the replaced value, provider and scope; the duplicate-key case is covered only by reading the code.

Dependencies

  • libduckdb-sys / duckdb 1.10504.0 → 1.10505.0 (DuckDB 1.5.4 → 1.5.5), cc 1.2.64 → 1.4.7, arrow 58.1.0 → 58.4.0. The relock removed 100 packages net: libduckdb-sys 1.10505.0 swapped its reqwest build-dependency for ureq, taking the hyper/tokio/quinn/rustls trees with it. MSRV is unchanged at 1.86.0.
  • No duckdb-1-5-5 feature was added, deliberately. DuckDB 1.5.5 adds no C Extension API surface: extension_api.hpp is byte-identical between v1.5.4 and v1.5.5 (sha256 0232a22a…3017031, 89456 bytes, 546 function pointers in each). There would be nothing to gate. See the note in Cargo.toml.
  • GitHub Actions, each SHA resolved against the upstream tag: actions/checkout v7.0.0 → v7.0.1, Swatinem/rust-cache v2.9.1 → v2.9.2, codecov/codecov-action v7.0.0 → v7.1.1, actions/attest-build-provenance v4.1.1 → v4.2.2, actions/deploy-pages v5.0.0 → v5.0.1. The actions/configure-pages pin was already v6.0.0; only its comment said v5.0.0.
  • Pinned the four CI tools installed unpinned (cargo-mutants 27.1.0, cargo-semver-checks 0.50.0, cargo-llvm-cov 0.9.1, cargo-fuzz 0.13.2). mutants.yml reasons about behaviour introduced in cargo-mutants 27.0.0, which held only by luck of whatever cargo install fetched. Every cargo install step now also clears the workflow-wide RUSTFLAGS: "-D warnings", which was compiling third-party trees with warnings-as-errors.
  • dtolnay/rust-toolchain is pinned to a commit on the action's master branch in every workflow and in the workflow generate_scaffold writes. The old pin was a commit of its regenerated stable branch that no ref reaches any more, so GitHub may garbage-collect it. Every use now passes toolchain: explicitly, as master's action.yml requires; the pin does not pin the Rust version.

0.16.0 — 2026-08-19

Security

  • New abi module: duckdb_ext_api_v1 layout verification. DuckDB hands a loadable extension a struct of function pointers. Its first 357 slots — the "stable prefix" — have been byte-for-byte identical in every release from v1.2.0 through v1.5.5, but everything past that is the unstable region, and DuckDB inserts new entries in the middle of it between releases (duckdb_appender_clear at slot 410 in v1.5.0, duckdb_geometry_type_get_crs at slot 493 in v1.5.2). Every quack-rs wrapper behind the duckdb-1-5 / duckdb-1-5-3 features — 105 C API functions covering scalar bind/init, copy functions, catalog access, ErrorData, FileSystem, Expression, SelectionVector, config options, table descriptions and the client context — lives in that region.

    DuckDB does not catch this: an extension stamped C_STRUCT + v1.2.0 (the default) is accepted by any DuckDB whose C API version is at least v1.2.0 and then handed the whole struct, unstable region included. Loading such a build into a DuckDB with a different layout silently dispatches to the wrong function pointers. Verified end-to-end: an extension built against DuckDB 1.5.0's headers, stamped C_STRUCT/v1.2.0, loaded into DuckDB 1.5.5 aborts the process with double free or corruption.

    abi::check compares the slot count of the compiled-in layout against the layout the running engine uses (resolved from duckdb_library_version(), which sits at stable slot 7 and is therefore always dispatched correctly). init_extension / init_extension_v2 and the entry_point! / entry_point_v2! macros now run that check under the new AbiPolicy::Strict default whenever duckdb-1-5 is enabled, turning the memory corruption above into a LOAD error that names the mismatch and the remedy. AbiPolicy::Warn and AbiPolicy::Trust opt out; Trust is the right choice for binaries stamped C_STRUCT_UNSTABLE, where DuckDB already pins the release. Extensions that stay on the stable prefix are unaffected and keep their forward compatibility.

    scripts/check-abi-table.py re-derives the layout table from every upstream release header and runs in CI, so the table cannot drift as DuckDB releases.

  • DuckStringView::from_bytes was unsound. It was safe to call yet dereferenced the heap pointer embedded in bytes 8–15 of a pointer-format duckdb_string_t, so safe code holding attacker-influenced bytes could read arbitrary memory. Replaced by two honest constructors: from_raw (unsafe, honours pointer format — what callbacks want) and inline_from_bytes (safe, returns None for pointer-format values). from_bytes is deprecated and no longer dereferences.

  • CopyGlobalInitInfo::get_file_path corrupted the heap. It called duckdb_free on the pointer from duckdb_copy_function_global_init_get_file_path, which returns info_ref.file_path.c_str() — the interior pointer of a C++ std::string DuckDB still owns and destroys itself. Every COPY ... TO through a quack-rs copy function handed the allocator a pointer it never issued; the first live test of the path aborted with corrupted size vs. prev_size in fastbins.

    Every other duckdb_free call site in the crate was then audited against DuckDB's implementation, and all twelve are correct. The signature is not sufficient to decide: char * returns are owned and const char * returns are usually borrowed, but duckdb_parameter_name is declared const char * and returns strdup(...), so it is owned. Recorded as LESSONS.md P11 with the full table.

Portability and feature-combination breakage

  • MAX_LIST_CHILD_CAPACITY was typed usize, making the crate fail to compile for wasm32. DuckDB's DConstants::MAX_VECTOR_SIZE is 1ULL << 37ULL — an idx_t, not a pointer-sized value. As a usize const, 1 << 37 is a const-eval overflow wherever pointers are 32 bits, which is every wasm32 target — and DuckDB's own extension CI builds three of them. Now typed u64, with a separate usize-clamped constant for the capacity arithmetic; on a 32-bit target the ceiling is larger than any allocation usize can describe, so usize::MAX is the real limit.

  • Three unit tests called duckdb-1-5-gated methods without a cfg gate, breaking --features bundled-test on its own — a combination that builds the live-DuckDB tests but not the 1.5 wrappers. Value::display_string and TableDescription::column_count / column_type are the gated methods; the assertions around them are now gated too.

    Both defects compiled cleanly under every other feature combination. scripts/check-matrix.sh now runs the combinations CI runs — including bundled-test alone and the wasm32 legs — in one command.

Security scanning

  • Security (OSV / GHSA) failed on RUSTSEC-2026-0235 (rkyv 0.7.46), an advisory that no build of this crate can reach. osv-scanner reads Cargo.lock, and Cargo pins optional dependencies there whether or not their feature is enabled, so the job flagged a crate that is never compiled. The chain is quack-rs → duckdb (optional dependency, enabled by bundled-test) → rust_decimal 1.40.0 → rkyv (optional feature of rust_decimal, not enabled). rust_decimal itself is built; rkyv is not. It is also not fixable here — rust_decimal 1.40 constrains rkyv to ^0.7, and the fixed version is 0.8.17. The advisory is present on main as well.

    Suppressed via a new osv-scanner.toml carrying the full reachability argument. The suppression does not rest on that comment staying true: the osv-scan job now re-derives it on every run, failing the build if an rkyv node ever appears in cargo tree --all-features --target all. The check tests the forward tree, because cargo tree -i exits 0 with "nothing to print" for a lockfile-only package and so cannot tell "not built" from "built" by exit code; it also anchors on rust_decimal being present, so a truncated or failed tree reports as unverified rather than as clean.

CI guards

  • scripts/check-abi-table.py treated an unreachable release header as proof the release did not exist. Its fetch swallowed every exception — 404, timeout, DNS, 5xx alike — and returned None, which the caller printed as "not published (skipped)" and dropped from the derivation. One transient failure on v1.4.4 therefore narrowed the derived range from v1.4.0–v1.4.4 to v1.4.0–v1.4.3 and failed CI reporting src/abi.rs as stale. Following that advice would have shrunk the layout table and made the runtime guard refuse DuckDB versions it should accept — the exact failure the table exists to prevent.

    A definite 404 is now distinguished from every other failure, transient errors are retried, and a tag that could not be downloaded suspends the staleness comparison (exit 2, "could not check") instead of failing it. The other three guards each fetch a single file, so an empty fetch already meant no data rather than partial data.

  • All four guard jobs treated exit 2 as a failure. The scripts document it as "upstream unreachable, could not check", but the workflow ran them bare, so any non-zero failed the job — making every guard a network-flake away from a red build. They now surface exit 2 as a warning and fail only on exit 1.

Generated CI

  • The generated CI workflow left one action unpinned. Three of its four actions were SHA-pinned; dtolnay/rust-toolchain@stable was not, justified by a comment claiming its SHA "changes with each Rust release". That is not how the action works — it reads the toolchain from rust-toolchain.toml or its toolchain: input at run time, so pinning the action's SHA does not pin the Rust version. quack-rs's own CI SHA-pins the same action and gets current stable. A branch is a moving target its owner can repoint, and a workflow step runs arbitrary code in the user's CI. All four are now pinned to the same SHAs quack-rs itself uses, and a test asserts every uses: in the generated workflow carries a 40-character hex ref.

Fixed

The release-profile validator required the setting that breaks panic safety

  • validate_release_profile required panic = "abort", which makes every one of quack-rs's panic guards inert. quack-rs wraps every extern "C" entry point — the extension entry point and every scalar/table/aggregate/cast/copy callback macro — in catch_unwind, so a panic in an extension's code becomes a DuckDB error instead of a crash. catch_unwind catches nothing under panic = "abort": the runtime aborts before unwinding starts. Demonstrated directly rather than assumed —

    rustc -O            panic_probe.rs  →  caught, process survived,  exit 0
    rustc -O -C panic=abort  …          →  Aborted,                   exit 134
    

    — so the validator was telling extension authors to configure the one setting that turns a recoverable SQL error into a SIGABRT that kills the user's whole DuckDB session.

    The crate already disagreed with itself: the scaffold has generated panic = "unwind" since the panic-safety work in this release, with a comment explaining why. validate_release_profile now requires "unwind" and rejects "abort" with that explanation; ReleaseProfileCheck::panic_abort is renamed panic_unwind. A new test asserts the scaffold and the validator agree, so they cannot drift apart again.

    The original justification — "panics across FFI boundaries are undefined behavior" — is also out of date: Rust defines an unwind escaping extern "C" as an abort, and quack-rs catches panics before the boundary regardless.

    quack-rs's own [profile.release] also said panic = "abort". Cargo ignores a dependency's profile so it changed nothing downstream, but it contradicted the crate's own advice; it now says "unwind".

  • validate_function_name rejected mixed-case names, and it gates try_new — so ScalarFunctionBuilder::try_new("myFunc") returned Err and the function could not be registered through quack-rs at all. DuckDB itself ships formatReadableSize and formatReadableDecimalSize, and registering a camelCase name through the C API succeeds: verified against DuckDB 1.5.5, where the function is then callable as formatReadableThing, formatreadablething and FORMATREADABLETHING, because DuckDB identifiers are case-insensitive.

    The rule was justified as avoiding "catalog issues"; that test disproves it. Letters of either case are now accepted. Everything that would genuinely break is still rejected — a name needing quotes in SQL (my-func, my func, my.func), one starting with a digit, one over 256 characters, one with an interior NUL. snake_case remains the right convention and is documented as one, rather than enforced as a rule that blocks a legal name.

    The same relaxation applies to AggregateFunctionBuilder, TableFunctionBuilder and SqlMacro parameter names, which share the validator.

    A regression test now runs validate_function_name over every function in duckdb_functions() (746 of them) and validate_extension_name over every entry in duckdb_extensions(), asserting that everything identifier-shaped is accepted and every operator is not. That is how the defect was found.

A documented convention that was not being followed

  • "Every unsafe block inside this crate has a // SAFETY: comment" was not true. clippy::undocumented_unsafe_blocks reports 180 blocks in the library. Most are inside an unsafe fn and merely forward that function's own documented contract — unsafe_op_in_unsafe_fn is denied crate-wide, so those blocks are required syntax rather than new assertions — but around forty were in safe functions, where the crate rather than the caller is asserting the invariant, and those had nothing.

    The claim is replaced with the convention actually worth following, and that convention is now met: every unsafe block in a safe function carries a // SAFETY: comment. Auditing them also turned up three comments that described the wrong thing — two duckdb_free calls and a duckdb_destroy_value annotated as if they were uses of the enclosing handle; those now say which allocation they own and why, cross-referencing LESSONS.md P11.

The scaffold generated a description.yml that would be rejected

  • repo.ref was generated as main. DuckDB's community-extension documentation is explicit: "Provide the hash of the latest commit on the branch targeting stable as ref". The repository builds exactly that revision and signs the result, so a branch makes the build unreproducible. Of the 43 published extensions sampled, 41 pin a full 40-character hash and two pin a tag; none uses a branch.

    ScaffoldConfig gains git_ref, defaulting to REF_PLACEHOLDER ("REPLACE_WITH_COMMIT_HASH") — deliberately not a valid revision, so it cannot be submitted by accident the way main silently could. The generated file carries a comment saying why, and a commented-out ref_next.

  • DescriptionYml silently dropped repo.ref_next. It is a documented field: while a new DuckDB release is being prepared, the community repository tests an extension against both the latest stable release and main, and ref_next names the revision compatible with main. Now parsed into git_ref_next, empty when absent.

  • The generated description.yml had no docs: section. All 43 published extensions have one — it is what renders on the community-extensions documentation site. The scaffold now emits hello_world and extended_description stubs.

Two more documented behaviours that were not the real ones

  • ClientContext::catalog documented an empty name as "the default catalog"; DuckDB rejects it outright. duckdb_client_context_get_catalog starts with if (!context || !name || strlen(name) == 0) return nullptr; — an empty string is the one value guaranteed to fail. The catalog of an in-memory database is named memory; a file database's is the file's stem. The doc now says so, along with the other None case the C API imposes and quack-rs never mentioned: DuckDB checks transaction.HasActiveTransaction(), so this works inside a callback but not on an idle auto-commit connection. Both verified against 1.5.5 by a live test.

  • ClientContext::config_option aborts the process when asked for a setting that does not exist — on a DuckDB built with debug assertions. duckdb_client_context_get_config_option calls TryGetCurrentSetting(...).GetScope() without first checking the lookup succeeded, and GetScope() asserts scope != SettingScope::INVALID. A release DuckDB compiles the assertion out and the function's own default: arm returns NULL as documented, so this never reproduces for end users and always reproduces in a test suite linking a debug DuckDB.

    This is a DuckDB defect, not a quack-rs one, but it makes the obvious "does the user have this setting?" probe unsafe. Documented on the method with the source lines, recorded as LESSONS.md P12, and the abort-free alternative given: SELECT count(*) FROM duckdb_settings() WHERE name = ?.

Documentation claimed a bridge that cannot exist

  • The secrets module described itself as bridging into DuckDB's secrets system. There is no such bridge, and there cannot be. The extension C API has zero secret functions — not one duckdb_secret_* among the 546 slots of duckdb_ext_api_v1 in DuckDB 1.5.5. An extension cannot ask DuckDB for a credential through the C API at all.

    The only route is the duckdb_secrets() table function, and DuckDB redacts sensitive fields there. Verified against 1.5.5:

    CREATE SECRET s (TYPE s3, KEY_ID 'AKIAEXAMPLE', SECRET 'super-secret-value');
    SELECT secret_string FROM duckdb_secrets();
    -- ...;key_id=AKIAEXAMPLE;secret=redacted
    

    The module docs now say this plainly, and say what SecretsManager actually is: a trait over the extension's own credential source, carrying the redacting Debug, zeroize-on-drop and absent PartialEq that credential handling needs, rather than a route to DuckDB's store.

    The zeroize claim is also narrowed to what is true: it covers the buffers a SecretEntry owns, not a String the caller still holds or one a String abandoned when it grew.

The description.yml validator rejected 84% of real extensions

  • parse_description_yml rejected 36 of the 43 published community extensions it was tested against. Its entire purpose is to tell an author their submission is valid before they open a PR, and it told almost everyone they were invalid. Four independent causes:

    1. requires_toolchains was treated as required. It is not — only 14 of the 43 set it, and the community-extensions documentation does not list it as required. This alone rejected half the corpus. It is now optional; validate_rust_extension still requires rust in it when present.

    2. YAML quotes were not stripped. parse_kv deliberately returned quoted values with their quotes and left stripping to each caller, and only excluded_platforms did. 12 of 43 files write version: '2025120401', so the parser saw '2025120401' — quotes included — and every version check failed on it. parse_kv now unquotes, with a real balanced-quote check rather than trim_matches, which would also eat ""doubled"" and a trailing a".

    3. validate_extension_version imposed a format DuckDB does not. It accepted only semver or a git hash; 11 of 43 published extensions use a date-based build id (2025120401). DuckDB's community-extension documentation specifies no version format at all — it says the descriptor carries "the version of the extension" and points at existing extensions as examples. The check is now what would actually break something: empty, over 64 characters, or containing anything outside [A-Za-z0-9._+-] (whitespace, path separators, control characters). classify_extension_version is unchanged — DuckDB's three-tier stability scheme is documented and is strict, and that function is where it belongs.

    4. windows_amd64_rtools was rejected. It is the R-tools Windows build (DuckDBPlatform() emits it under DUCKDB_PLATFORM_RTOOLS), it is not in the distribution matrix, and 14 of 43 published extensions exclude it. DUCKDB_PLATFORMS now also accepts it and the four group names (linux, osx, wasm, windows — the top-level keys of distribution_matrix.json), while the new DUCKDB_CI_PLATFORMS keeps the matrix-derived list the guard script checks. Empty segments from a trailing ; — which five real files have — are skipped rather than reported as a platform named "".

    All 43 now parse, with every name matching its directory.

  • Prose in the docs: section was parsed as metadata. The scan was flat, so a version: or license: line inside docs.extended_description — free-form prose in 42 of the 43 files — silently overwrote the extension's real values. Demonstrated: a license: FAKE-LICENSE line inside a documentation block made a valid file fail validation, and the same mechanism could have made an invalid one pass. The parser is now section-aware (only extension: and repo: are read) and understands block scalars: key: | and key: > bodies are captured as the field's value — literal blocks keeping line breaks, folded blocks joined — instead of being scanned for mappings.

  • Three doc examples showed indented YAML that was not indented. A \ line-continuation in a Rust string literal eats the following line's leading whitespace, so description.yml examples in parse_description_yml, validate_description_yml_str and validate_rust_extension were parsing fully-unindented text. They only passed because the parser ignored indentation; making it section-aware exposed them. Rewritten as real multi-line literals.

Validators were giving wrong answers

  • The DuckDB platform list was stale in both directions. validate::platform rejected linux_amd64_musl and linux_arm64_musl — real, currently-built targets — so an extension that legitimately cannot support musl could not declare it. And it accepted linux_amd64_gcc4, which DuckDB retired: DuckDBPlatform() in duckdb/common/platform.hpp now raises a compile error for the legacy CXX ABI rather than emitting a _gcc4 suffix, and it is absent from the distribution matrix. Excluding it was a silent no-op.

    The list is now derived from config/distribution_matrix.json in duckdb/extension-ci-tools — the file the community-extensions build actually reads — and scripts/check-platform-table.py plus a CI job fail when the two diverge. Adds DUCKDB_OPT_IN_PLATFORMS and is_opt_in_platform, because three of the twelve (linux_amd64_musl, linux_arm64_musl, windows_arm64) are only built on request, so excluding one of those is also a no-op. linux_amd64_gcc4 gets a targeted error saying what happened to it, rather than "not a recognized DuckDB build target".

  • validate_spdx_license claimed valid licenses did not exist. COMMON_SPDX_LICENSES is a 42-entry shortlist of a 733-entry registry, but the rejection message read "is not a recognized SPDX identifier" — false for CC0-1.0, Python-2.0, BSD-4-Clause and roughly 690 others. It now says the identifier is not on quack-rs's shortlist and points at the registry.

    Every entry was checked against spdx/license-list-data: all 42 are real and none are deprecated. scripts/check-spdx-list.py and a CI job keep it that way, and flag any newly-added identifier that is not OSI-approved (SSPL-1.0 is listed and deliberately is not). The list is now sorted, with a test keeping it so. Also fixes the module doc, which called the field extension.licence; real description.yml files — and quack-rs's own parser — use license.

Silent data corruption

  • The UUID accessors disagreed about which 128 bits they meant, and the documentation said they agreed. A UUID column is physically a HUGEINT, but DuckDB stores it with the top bit flipped so that signed integer ordering matches UUID string ordering (BaseUUID::FromUHugeint in src/common/types/uuid.cpp subtracts 2^63 from the upper half). So:

    AccessorReturnedFor '11111111-…'::UUID
    VectorReader::read_uuid (old)raw storage0x9111…
    Value::as_uuidtextual bits0x1111…

    Both were documented as "matching" the other. Handing one to the other — the obvious thing to do when a table function reads a UUID and builds a Value from it — silently changed the UUID's first hex digit.

    read_uuid / write_uuid (on VectorReader, VectorWriter, StructReader, StructWriter and both mocks) now apply the flip and take/return u128 textual bits, the same convention as Value::uuid / Value::as_uuid and every Rust Uuid type. Value::uuid / as_uuid move from i128 to u128 for the same reason. The type change is deliberate: it turns a silent behaviour change into a compile error at every affected call site.

    read_i128 / write_i128 still read and write the raw storage, and the new vector::uuid_from_storage / vector::uuid_to_storage convert between the two. Pinned by a live test that asserts the raw storage and the textual bits really do differ, so the conversion cannot quietly become a no-op.

Wrong results and unloadable builds

  • ChunkWriter no longer hardcodes a 2048-row capacity. DuckDB can be built with a different STANDARD_VECTOR_SIZE, which is exactly why the C API exposes duckdb_vector_size(); assuming 2048 against a smaller build overruns the output vectors. ChunkWriter::new now reads the running engine's value. ChunkWriter::new and DataChunk::into_chunk_writer are consequently no longer const fn.

  • The scaffold produced an extension DuckDB refuses to load. The generated Makefile set DUCKDB_PLATFORM_VERSION, which extension-ci-tools does not read, alongside USE_UNSTABLE_C_API=1. TARGET_DUCKDB_VERSION therefore fell back to its v0.0.1 default and the binary was stamped C_STRUCT_UNSTABLE/v0.0.1, which DuckDB rejects with "The file was built specifically for DuckDB version 'v0.0.1'". The generated Makefile now sets EXTENSION_NAME (not EXT_NAME, which base.Makefile ignores), TARGET_DUCKDB_VERSION and USE_UNSTABLE_C_API from the new ScaffoldConfig fields, and defines the all/configure/debug/release/ test/clean targets its own README and CI invoke.

  • The scaffold generated panic = "abort", which makes the catch_unwind in scalar_callback!, table_scan_callback! and the extension entry point inert — so any panic in extension code killed the whole DuckDB process instead of surfacing as a SQL error. Now generates panic = "unwind".

  • The scaffold pinned quack-rs = "0.13" regardless of the generating crate's version. It now tracks the current major.minor.

  • A freshly scaffolded project failed its own generated CI. cargo clippy --all-targets -- -D warnings (which the generated workflow runs) rejected the generated src/lib.rs for clippy::redundant_closure and src/wasm_lib.rs for special_module_name. Both are fixed; a new scaffold-e2e CI job builds the generated project, stamps its metadata footer, loads it into a real DuckDB, asserts the query result, and runs the generated lint gate.

  • The generated CI referenced a nonexistent action (duckdb/duckdb-build@v1) and ran make test without make configure / make release, so it could not have passed. Replaced with a workflow that configures, builds and tests through extension-ci-tools.

  • The extension entry point ran user registration code without catch_unwind. A panic in a registration closure unwound to the extern "C" entry point, aborting the process; it now becomes a LOAD error. An api_version containing an interior NUL is also rejected up front instead of panicking inside libduckdb-sys.

Behaviour documented after verification

  • Value::display_string renders a SQL literal, not display text: Value::varchar("hello") gives 'hello' and Value::date(0) gives '1970-01-01'::DATE. Now documented with a table, since silently getting quotes and a cast suffix in a diagnostic is surprising.
  • Value::as_str truncates at an interior NUL, because duckdb_get_varchar returns a NUL-terminated char *. DuckDB stores the full bytes; only this read path is limited. Documented on both as_str and Value::varchar, and pinned by a test.

Documentation

  • The crate documented an "architectural limitation" that does not exist. Cargo.toml, testing::in_memory_db and the book all stated that VectorReader, VectorWriter and Connection::register_* "cannot be called in cargo test" because they route through the dispatch table. Opening an InMemoryDb populates that table for the whole process, after which the entire C API — registration included — works. The new tests/ffi_roundtrip.rs registers real scalar functions and round-trips every vector type through SQL: every integer width at its extremes, HUGEINT/UHUGEINT at theirs, floats and NaN, strings across the 12-byte inline/pointer boundary and multi-byte UTF-8, blobs containing NUL and non-UTF-8 bytes, all temporal types cross-checked against DuckDB's own rendering, UUID, INTERVAL's three fields, DECIMAL at all four physical widths, NULL in and out, multi-chunk scans, and a panicking callback surfacing as a SQL error.

  • Documentation examples pinned quack-rs = "0.13".

Added

Live tests for every previously untested C API path

  • Copy functions and replacement scans had no live tests at all. Between them they had 19 unit tests, none of which registered anything against a running DuckDB — which is how a heap-corrupting free survived in a shipped API. Both now have end-to-end coverage:

    • A COPY ... TO 'f' (FORMAT my_format) over 5000 rows, threading bind data and global state through all four lifecycle phases, asserting the sink saw every row and that both destructors ran exactly once (a leak or a double free is invisible without counting).
    • A replacement scan rewriting SELECT * FROM '10.myfmt' into a table function call, plus the decline path — an identifier the callback ignores must still reach DuckDB's own error handling — and a panicking scan surfacing as a SQL error.
  • Six more modules had unit tests but no live registration: scalar bind/init/local state, Expression::fold, catalog lookup, config options, selection vectors and the instance cache. All now run against a real DuckDB, which turned up two more documentation defects (below) and confirmed the rest.

  • copy_bind_callback!, copy_global_init_callback!, copy_sink_callback! and copy_finalize_callback!. Every other callback kind had a panic-safe macro; the four copy-function phases did not, so a panic in one of them had nothing to catch it. Each routes the message through that phase's own duckdb_copy_function_*_set_error.

  • TypeId::try_from_duckdb_type — returns Option<TypeId> instead of panicking on a type value this build does not know. Extensions routinely meet these: a column of a type added in a newer DuckDB, or a 1.5.x type reaching a build without duckdb-1-5. from_duckdb_type still panics and now documents that callbacks should not use it.

  • Fallible LogicalType constructors that previously panicked on an interior NUL in a caller-supplied name: try_struct_type_from_logical, try_union_type, try_union_type_from_logical, try_enum_type, try_set_alias.

  • entry_point! / entry_point_v2! accept an optional AbiPolicy as their second argument; init_extension_with_policy / init_extension_v2_with_policy are the function-level equivalents.

  • examples/scaffold_to_dir.rs — writes a scaffolded project to disk, used by the new scaffold-e2e CI job.

Panic safety

  • A panic-safe wrapper macro for every callback kind. Only scalar_callback! and table_scan_callback! existed, so the other six kinds — table bind, table init, aggregate update/combine/finalize/destroy, cast, and replacement scan — were unguarded, and a panic in any of them aborted the DuckDB process. The aggregate ones are the worst case: they run on worker threads, so the abort comes from a thread the user never sees. New macros: table_bind_callback!, table_init_callback!, aggregate_update_callback!, aggregate_combine_callback!, aggregate_finalize_callback!, aggregate_destroy_callback!, cast_callback!, replacement_scan_callback!. Each routes the panic message to that callback kind's own set_error; cast_callback! also returns false so TRY_CAST yields NULL. The aggregate destructor has no error channel in the C API, so its panic is caught and dropped — leaking beats aborting during query teardown. Verified end-to-end: a panicking aggregate update and a panicking cast both surface as SQL errors and leave the connection usable.

  • The two existing macros now share callback::panic_message and callback::message_to_c_string with the new ones. The latter replaces an interior NUL rather than dropping the diagnostic, which the old if let Ok(c_msg) = CString::new(msg) silently did.

  • TypedTableFunctionBuilder reported every panic as the same fixed string. It now includes the payload, so the user learns which assertion failed.

  • Deprecated FfiBindData::get_from_bind, which always returned None and always will: DuckDB exposes no duckdb_bind_get_bind_data. Being safe and returning Option, it silently sent if let Some(..) down the wrong branch.

Capabilities

  • ListBuilder for LIST and MAP output vectors. duckdb_list_vector_reserve takes a total capacity and reallocates the child vector when it grows, so a VectorWriter obtained beforehand is left dangling. That makes the natural "reserve as you go, keep one writer" loop a use-after-free. ListBuilder re-fetches the child writer after every reserve, tracks the running offset, writes each parent {offset, length} entry, and grows geometrically so building a list is not quadratic. push_map_row does the same for MAP. It also refuses capacities above MAX_LIST_CHILD_CAPACITY (duckdb::DConstants::MAX_VECTOR_SIZE), above which DuckDB throws a C++ exception that its own C API does not catch — an exception unwinding into Rust would be undefined behaviour. Covered by tests building 2000 lists and 1500 maps of varying length through real SQL.

  • Value gained the extractors and constructors it was missing. A table function declared with a TIMESTAMP or LIST parameter handed the bind callback a duckdb_value that could only be read via as_str() and reparsed. Adds as_date, as_time, as_time_tz, as_timestamp, as_timestamp_tz, as_timestamp_s/ms/ns, as_interval, as_uuid, as_decimal, as_u128, list_len / list_child / list_items, struct_child, map_len / map_key / map_value, and the constructors boolean, bigint, double, date, timestamp, varchar, uuid, null_value.

  • query module — running SQL from inside an extension. The C API has everything needed (duckdb_query, duckdb_prepare, duckdb_bind_*, duckdb_fetch_chunk) and it is all in the stable prefix, but each handle has a destroy that must run exactly once, including on error paths. QueryResult, OwnedDataChunk, PreparedStatement and OwnedConnection are RAII wrappers for those; Connection gains query, execute, prepare and open_connection.

    OwnedConnection covers the case the borrowed registration connection cannot: a duckdb_connection holds its own reference to the database instance, so one opened during load stays valid afterwards — for a callback or a background thread. Verified by a test that closes the duckdb_database handle and keeps querying.

  • datetime module — calendar conversions. DATE, TIME and TIMESTAMP move through vectors as raw integers; turning those into year/month/day meant reimplementing the proleptic Gregorian calendar and DuckDB's infinity sentinels. DuckDB already exposes the conversions in the stable API, so this wraps them: date_from_days/date_to_days, time_from_micros/time_to_micros, timestamp_from_micros/timestamp_to_micros, time_tz_bits/time_tz_from_bits, the four is_finite_* predicates, and HUGEINT/UHUGEINT/DECIMAL ↔ f64.

    Also exports the exact sentinel values as constants. -infinity is -i32::MAX / -i64::MAX, not i32::MIN / i64::MIN — i32::MIN is an ordinary finite date, and treating it as infinity would silently drop real rows.

  • VectorWriter caches its validity bitmap. set_null called duckdb_vector_ensure_validity_writable + duckdb_vector_get_validity on every row; both are now resolved once per vector (2 FFI calls instead of 4096 for an all-NULL 2048-row vector). Adds set_null_range for the batched case.

  • Vector accessors for the remaining physical layouts: write_u128/read_u128 (UHUGEINT), write_decimal/read_decimal (which select i16/i32/i64/i128 from the declared width the way DuckDB does), write_time_tz/read_time_tz, and TIMESTAMPTZ / TIMESTAMP_S / TIMESTAMP_MS / TIMESTAMP_NS accessors. VectorReader::contains bounds-checks an index against the row count.

  • Callback signature aliases are re-exported at their module roots: scalar::ScalarFn (plus ScalarBindFn / ScalarInitFn under duckdb-1-5) and aggregate::{StateSizeFn, StateInitFn, UpdateFn, CombineFn, FinalizeFn, DestroyFn}, matching what table already did.

  • The prelude re-exports AbiPolicy, the datetime types and the query types.

  • Registrar::register_config_option — the trait already covered scalar, scalar set, aggregate, aggregate set, table, SQL macro, cast and copy functions, but not config options, so an extension registering one could not have its whole registration closure exercised through MockRegistrar. Added, with config_option_names / has_config_option on the mock.

  • secrets::list_duckdb_secrets — reads the secret metadata DuckDB does expose, via duckdb_secrets(): name, type, provider, persistence, storage, scope prefixes and the redacted secret_string. Enough to pick a scope, warn that a required secret is missing, or choose a provider. It returns a DuckDbSecretInfo, deliberately not a SecretEntry, so nothing suggests it carries credentials. A live test asserts both halves: the metadata comes through, and the credential provably does not.

  • The appender is no longer behind duckdb-1-5, and gained the row-at-a-time API it never had. duckdb_appender_* occupies slots 281–291 and 330–356 — the frozen stable prefix, unchanged since v1.2.0 — yet the whole module was gated on duckdb-1-5, whose wrappers live in the unstable region. Using the appender therefore forced an extension onto the version-pinned unstable ABI, for functionality that has been portable for four minor releases. Only three methods actually need 1.5 and stay gated: error_data, clear and append_default_to_chunk.

    The 24 row-at-a-time functions were wrapped for the first time: append_bool / _i8 / _i16 / _i32 / _i64 / _i128 / _u8 / _u16 / _u32 / _u64 / _u128 / _f32 / _f64 / _str / _bytes / _date / _time / _timestamp / _interval / _value / _null / _default, end_row, column_count, column_type, add_column, clear_columns, and a row(|row| …) helper that calls end_row for you. Previously the only way to insert a row was to build a whole DataChunk.

    Three details that are easy to get wrong and are handled here: append_str uses duckdb_append_varchar_length, so interior NUL bytes survive; that function narrows its length to uint32_t with an unchecked cast in DuckDB's release builds, so longer strings are refused rather than truncated; and duckdb_append_value dereferences its argument with no null check, so a null Value handle is refused. Covered by live tests that append every scalar type at its extremes, 5000 rows across several vectors, a short row, a constraint violation surfacing at close, and a DEFAULT-filled column subset.

    New appender::AppendError is ErrorData with duckdb-1-5 and ExtensionError without, so enabling the feature upgrades the error type in place without changing any method's shape — existing duckdb-1-5 code is unaffected.

  • table_description is no longer behind duckdb-1-5 either. Slots 292–297 are stable; only column_count and column_type are 1.5 additions and stay gated. Adds TableDescription::with_catalog (duckdb_table_description_create_ext, for tables in another catalog) and column_has_default (duckdb_column_has_default) — the latter being the only way to know whether Appender::append_default will succeed.

  • FileHandle gained the looping I/O helpers, and size/tell became fallible. duckdb_file_handle_read and duckdb_file_handle_write return "the number of bytes actually read/written" — a single call can come up short, which over httpfs is routine rather than theoretical. Adds read_exact, read_to_end and write_all, which loop. size() and tell() changed from i64 to Result<u64, ErrorData>: the C API signals failure with a negative return, and the previous signature made handle.size().max(0) as usize — silently treating an error as an empty file — the obvious thing to write. It was in this crate's own documentation.

  • Value::type_id(). Value had forty as_* accessors and no way to ask what the value actually is, so reading a VARCHAR with as_i64() returned garbage rather than an error. Wraps duckdb_get_value_type (stable prefix, slot 137, unchanged since v1.2.0), returning None for a null handle or a type id newer than this build knows.

  • Every public type implements Debug. 58 of them did not, which is Rust API guideline C-DEBUG and not cosmetic: Result::unwrap, Result::expect_err, assert_eq!, and #[derive(Debug)] on any downstream struct storing a quack-rs type all fail to compile without it. LogicalType and Value print decoded state (type id, alias, DECIMAL width/scale, DuckDB's own rendering) rather than a pointer; builders print set/unset per callback, which is the question you have when register reports a missing function; WarningCollector uses try_lock so printing can neither block nor deadlock. missing_debug_implementations is now enabled crate-wide, and CI's -D warnings makes it an error. testing::InMemoryDb was a 59th, only visible once the lint ran with bundled-test on.

Changed

MSRV

  • MSRV lowered 1.87.0 → 1.86.0. DuckDB's reusable _extension_distribution.yml — the workflow the community-extensions repository builds every extension with — pins dtolnay/rust-toolchain@… # 1.86.0 for the WebAssembly job. quack-rs required 1.87.0, so Cargo refused, and no quack-rs extension could be built for wasm_mvp / wasm_eh / wasm_threads by the official pipeline — despite the crate advertising wasm32-unknown-emscripten support since 0.14.0.

    The entire 1.87 requirement was five const fn accessors calling Vec::len (stabilised as const in 1.87). None can be reached in a const context — MockVectorWriter, StructReader and StructWriter are all built at runtime — so dropping const costs nothing. 1.86.0 is now the floor for the library, its dev-dependencies (criterion needs 1.86) and the hello-ext example, all verified.

    New scripts/check-msrv-vs-duckdb-ci.py and a CI job re-derive DuckDB's pinned toolchains from that workflow and fail if the MSRV creeps back above them.

  • Breaking: ScaffoldConfig gains target_duckdb_version and use_unstable_c_api. ScaffoldConfig now implements Default, so existing struct literals can add ..ScaffoldConfig::default(). generate_scaffold rejects combinations that produce an unloadable binary — a C_STRUCT build claiming a DuckDB release as its -dv, or a C_STRUCT_UNSTABLE build claiming the C API version.

CI / tooling

  • New abi-table job: scripts/check-abi-table.py verifies src/abi.rs's layout table against every upstream DuckDB release header.
  • New abi-guard job: builds an extension against DuckDB 1.5.0's header layout, stamps it C_STRUCT, and asserts the load is refused with a layout diagnostic — a regression test for the corruption described above.
  • New scaffold-e2e job (see above).
  • extension-load now stamps a real metadata footer and asserts query results rather than grepping the log for the word "error"; loading a bare .so bypassed DuckDB's metadata validation entirely.

0.15.0 — 2026-07-16

Added

  • Value::as_blob() for copying arbitrary binary data from a duckdb_value. (Thanks @adonm.)

Fixed

  • VectorReader::read_blob() now preserves non-UTF-8 bytes instead of returning an empty slice. (Thanks @adonm.)

Changed

  • Dev/CI DuckDB bumped to 1.5.4 — libduckdb-sys / duckdb 1.10503.1 → 1.10504.0 in the root lockfile, the hello-ext example lockfile, and the bundled-test-prebuilt CI download (v1.5.3 → v1.5.4). 1.5.4 is a bugfix release in the 1.5.x line; its C extension API version is unchanged (v1.2.0, verified from duckdb_extension.h), so DUCKDB_API_VERSION is unchanged and the public libduckdb-sys dependency range (>=1.4.4, <2) is untouched — downstream consumers are unaffected.

Security

  • crossbeam-epoch 0.9.18 → 0.9.20 (root lockfile), resolving RUSTSEC-2026-0204 (invalid pointer dereference in the fmt::Pointer impl). Reaches the tree only as a dev-dependency via criterion → rayon → crossbeam-deque.
  • quinn-proto 0.11.14 → 0.11.15 (root and example lockfiles), resolving RUSTSEC-2026-0185 (CVSS 7.5). Reaches the lockfiles via libduckdb-sys → reqwest → quinn (feature-union only; the loadable-extension build never links it).

CI / tooling

  • Refreshed SHA-pinned GitHub Actions via Dependabot: actions/checkout v6.0.2 → v7.0.0, codecov/codecov-action v6.0.1 → v7.0.0, actions/cache v5.0.5 → v6.1.0, and actions/attest-build-provenance v4.1.0 → v4.1.1. Also bumped the cc build-dependency 1.2.63 → 1.2.64.

0.14.0 — 2026-06-07

Added

  • wasm32-unknown-emscripten support (the DuckDB-WASM target). The crate no longer hard-rejects non-64-bit targets with a top-level compile_error!, and the duckdb_string_t pointer slot is read as a u64 then narrowed to usize — lossless on 64-bit, and on wasm32 it yields the low 4 bytes of the 8-byte slot (the upper 4 are zero padding in DuckDB's 16-byte layout). The full public API, including the duckdb-1-5-3 surface, cargo checks for wasm32-unknown-emscripten; CI now guards this. (Thanks @killzoner.)
  • bundled-test-prebuilt feature — links a pre-built libduckdb instead of compiling DuckDB from C++ source, for a much faster test build. Supply the library via DUCKDB_DOWNLOAD_LIB=1 (libduckdb-sys downloads the upstream release zip) or DUCKDB_LIB_DIR=... (a libduckdb tree you already have). bundled-test continues to compile DuckDB from source. (Thanks @killzoner.)
  • InMemoryDb::open_unsigned() opens an in-memory database with allow_unsigned_extensions=true, allowing downstream extension crates to LOAD their own locally-built (unsigned) .duckdb_extension artifact for integration testing. (Thanks @killzoner.)

Changed

  • duckdb is now a purely optional dependency, activated only by bundled-test / bundled-test-prebuilt. It is no longer a dev-dependency, and there is no default bundled feature. As a result, a plain cargo test — and every downstream consumer's Cargo.lock — no longer pulls the DuckDB + arrow tree, and the default test build no longer compiles DuckDB.

Security

  • tar 0.4.45 → 0.4.46 in both the root and example lockfiles, resolving GHSA-3pv8-6f4r-ffg2 ("PAX header desynchronization", Moderate). tar is a libduckdb-sys build-dependency, so it appears in both Cargo.lock files and raised one Dependabot alert each — the two moderate alerts reported on main. This advisory is published in the GitHub Advisory Database (GHSA) but not the RustSec database, so cargo deny did not flag it; the new OSV scan below closes that gap.
  • Bumped cc 1.2.62 → 1.2.63 (which moves shlex 1.3.0 → 2.0.1) and refreshed the codecov/codecov-action pin to v6.0.1.

CI / tooling

  • Added an OSV / GHSA advisory scan to CI (osv-scanner, pinned to v2.3.8 via a checksum-verified binary) covering both Cargo.lock files. cargo deny consults only the RustSec database; OSV.dev aggregates GHSA and RustSec, so GHSA-only advisories (such as the tar one above) now fail CI alongside the existing cargo-deny gate.

0.13.0 — 2026-05-24

Added

New safe wrappers for the DuckDB 1.5.0+ C extension API, all gated behind the duckdb-1-5 feature, plus a new duckdb-1-5-3 feature that surfaces the two DuckDB 1.5.3 type-enum values. DuckDB 1.5.3's C extension function-pointer API (version v1.2.0) is unchanged from 1.5.2; the one new C addition — the DUCKDB_TYPE_VARIANT (41) type-enum value — is now exposed as TypeId::Variant behind the duckdb-1-5-3 feature (see below). So the additions below mostly expose 1.5.x capabilities the SDK had not previously wrapped rather than anything new to 1.5.3 specifically.

  • error_data module — ErrorData, an RAII wrapper over duckdb_error_data (the structured error type returned by several 1.5 APIs). Carries a DuckDbErrorType category and a message, and converts into ExtensionError. Adds the free function check_valid_utf8, exposing DuckDB's own UTF-8 validator.
  • expression module — Expression, an RAII wrapper over duckdb_expression, with return_type, is_foldable, and fold. This closes a real gap: ScalarBindInfo already returned a raw, unusable duckdb_expression from get_argument; the new ScalarBindInfo::argument returns a safe Expression, so bind callbacks can inspect argument types and pre-fold constant arguments once at bind time.
  • file_system module — FileSystem, FileHandle, FileOpenOptions, and FileFlag: read and write files through DuckDB's virtual file system (honouring httpfs, in-memory files, and other registered file systems) instead of reaching for std::fs.
  • appender module — Appender: bulk row insertion (create, append a DataChunk, flush, close) plus the 1.5 additions clear (revert buffered rows), error_data (structured errors), and append_default_to_chunk.
  • selection_vector module — SelectionVector: allocate and fill zero-copy row-index selection vectors.
  • instance_cache module — InstanceCache: share one underlying database instance across repeated opens of the same path.
  • Value gains display_string (canonical string rendering of any value, via duckdb_value_to_string) and TIME_NS accessors Value::time_ns / Value::as_time_ns (pairing with the existing TypeId::TimeNs).
  • Catalog gains type_name (the catalog's storage type, e.g. "duckdb" or a storage extension's name).
  • All new public types are re-exported from the prelude behind the duckdb-1-5 feature.
  • duckdb-1-5-3 feature + TypeId::Variant / TypeId::Geometry — a new feature flag (duckdb-1-5-3, which implies duckdb-1-5) exposes the DUCKDB_TYPE_VARIANT (41, added in DuckDB 1.5.3) and DUCKDB_TYPE_GEOMETRY (40) type-enum values as TypeId::Variant and TypeId::Geometry, with the matching to_duckdb_type / from_duckdb_type / sql_name / Display coverage. It is a separate gate because these constants postdate the duckdb-1-5 feature's 1.5.0 floor and require libduckdb-sys >= 1.10503.1; keeping them out of duckdb-1-5 preserves compatibility for consumers pinned to libduckdb-sys 1.5.0–1.5.2.
  • ErrorData is now a first-class error type — implements std::fmt::Display and std::error::Error, gains a structured Debug impl, and converts into ExtensionError via From (alongside the existing into_extension_error) so it propagates through ?. DuckDbErrorType now implements Display (backed by a new pub const fn as_str).
  • TableDescription::as_raw() — exposes the raw handle, matching the accessor convention of the other 1.5 wrappers.

Changed

  • duckdb / libduckdb-sys 1.10502.0 → 1.10503.1 (DuckDB 1.5.2 → 1.5.3) in both the workspace and examples/hello-ext Cargo.lock. DuckDB 1.5.3 is a bugfix release (announcement); since the >=1.4.4, <2 constraint already permitted it, the bundled fixes are picked up purely by the lock-file update with no source changes required for the bump itself.
  • cc → 1.2.62 in both Cargo.lock files — workspace (1.2.61 → 1.2.62, folding in Dependabot PR #89, the patch-updates group) and examples/hello-ext (1.2.57 → 1.2.62, re-syncing the example lock's older cc). Build-dependency; no API impact.
  • MSRV corrected to 1.87.0. The crate declared rust-version = "1.84.1", but libduckdb-sys (1.5.x line, a non-optional dependency) is edition = "2024" / rust-version = "1.85.1" — so quack-rs has in fact required Rust ≥ 1.85.1 since before this release (cargo +1.84.1 check cannot even parse the manifest). The declared MSRV, the CI MSRV job (now explicitly pinned with toolchain: "1.87.0" so it genuinely gates instead of silently falling back to the rust-toolchain.toml stable channel), the release matrix, and all docs/badges are updated to 1.87.0 — a small headroom margin above the 1.85.1 floor.

Fixed

  • TypeId::from_duckdb_type no longer panics on the duckdb-1-5 type-enum values. It previously recognised only the base (1.4) values and panic!ed on everything else — including the duckdb-1-5 values (TIME_NS, ANY, BIGNUM/VARINT, SQLNULL, INTEGER_LITERAL, STRING_LITERAL). Because the public LogicalType::get_type_id() calls it, inspecting such a type inside a bind callback could panic across the FFI boundary (Pitfall L3). It now maps every variant available in the active feature set (plus the duckdb-1-5-3 GEOMETRY / VARIANT values when that feature is enabled).
  • TableDescription's Drop now null-checks the handle before destroying it, matching every other RAII wrapper in the crate.

Documentation

  • New book section "DuckDB 1.5+ APIs" — dedicated guide pages for the error_data, expression, appender, file_system, selection_vector, and instance_cache modules, wired into SUMMARY.md.
  • Refreshed the reference docs (docs/architecture.md, docs/ffi-reference.md, the TypeId reference, CONTRIBUTING.md/book source trees) to cover the new modules, and updated the VARIANT/GEOMETRY entries in Known Limitations, concepts/types.md, and the TypeId reference to document the new duckdb-1-5-3 gate (previously tracked as a follow-up).
  • Added // SAFETY: comments to previously-undocumented unsafe blocks in the get_client_context accessors (scalar, copy_function) and TableDescription::create, and SPDX headers to benches/interval_bench.rs and the test submodule files — closing the last gaps against the crate's own "every file / every unsafe block" conventions.
  • Corrected the README install note (it claimed v0.11.0 was the latest published crate; v0.12.1 was in fact already on crates.io) and bumped install-example version references throughout the README, book, and scaffold template to 0.13.

CI

  • docs.rs now builds with duckdb-1-5-3 ([package.metadata.docs.rs]), so the feature-gated modules (appender, error_data, file_system, …) and the new TypeId variants render on docs.rs and the README's docs.rs links resolve. Previously docs.rs built the empty default feature set and omitted them.
  • CI exercises the duckdb-1-5-3 feature — the feature job now runs check / test / clippy for duckdb-1-5-3 alongside duckdb-1-5, and the Clippy (beta) and doc jobs use duckdb-1-5-3.
  • Fixed the Nightly CI job silently running stable — the SHA-pinned dtolnay/rust-toolchain step lacked with: toolchain: nightly, so it fell back to the rust-toolchain.toml stable channel (the same class of bug previously fixed for the MSRV job).
  • Mutation testing scoped to testable code — DuckDB FFI-wrapper modules whose methods require a live runtime (and whose tests are bundled-test-gated or absent) are excluded from cargo mutants, since their mutants can't be killed by unit tests. This extends the existing exclusion pattern to the 1.5.x wrappers — expression, file_system, appender, selection_vector, instance_cache, table_description, and the scalar/copy *Info accessors. Pure-logic code (e.g. DuckDbErrorType, the TypeId conversions) stays in scope. The mutants feature set is bumped to duckdb-1-5-3.

0.12.1 — 2026-05-01

Security

Closes nine GitHub Dependabot alerts (two High, seven Low) split across the workspace Cargo.lock and examples/hello-ext/Cargo.lock.

  • rustls-webpki 0.103.10 → 0.103.13 — picks up the fix for three RustSec advisories reachable via the bundled DuckDB build's transitive reqwest → rustls chain:

    • RUSTSEC-2026-0098 (GHSA-965h-392x-2mh5) — nameConstraints with URI name restrictions were silently ignored instead of enforced. Patched in 0.103.12+; the URI-name path is not on the public Web PKI, so impact is limited to private-PKI consumers.
    • RUSTSEC-2026-0103 (GHSA-xgp8-3hg3-c2mh) — name-constraint enforcement accepted certificates asserting a wildcard subject name. Reachable only after signature verification and requires misissuance. Patched in 0.103.12+.
    • RUSTSEC-2026-0104 — reachable panic when parsing certificate revocation lists with a syntactically valid empty BIT STRING in the onlySomeReasons element of an IssuingDistributionPoint CRL extension. Affects only applications that use CRLs. Patched in 0.103.13+.

    Neither path is exercised by quack-rs itself, but the advisories trip cargo deny for any downstream consumer that has not yet bumped, so shipping a release that resolves them is the path of least friction.

  • rand 0.9.2 → 0.9.4 / 0.8.5 → 0.8.6 — picks up the fix for RUSTSEC-2026-0097 (GHSA-cq8v-f236-94qc) — ThreadRng could produce an aliased &mut BlockRng<ReseedingCore> (Stacked-Borrows UB) when a custom logger reentered rand::rng() from inside a reseed at trace-level logging. Triggering the unsoundness requires a custom global logger that pulls from rand::rng() while reseeding, which is not a pattern quack-rs uses, but the advisory matches by version range so resolving it removes the alert noise. Patched on every affected line: 0.8.6+, 0.9.3+, 0.10.1+.

Changed

  • Workspace Cargo.lock bumps —
    • cc 1.2.59 → 1.2.61 (build-dep; no API impact)
    • duckdb / libduckdb-sys 1.10501.0 → 1.10502.0 (latest patch release; no API impact for quack-rs)
    • rand 0.8.5 → 0.8.6 (transitive via rust_decimal; security)
    • rand 0.9.2 → 0.9.4 (transitive via proptest dev-dep; security)
  • examples/hello-ext Cargo.lock bumps —
    • libduckdb-sys 1.10501.0 → 1.10502.0
    • rand 0.9.2 → 0.9.4 (security)
    • rustls-webpki 0.103.10 → 0.103.13 (security; matches workspace)

CI

  • GitHub Actions pin updates —

    • actions/cache v5.0.4 → v5.0.5
    • actions/upload-artifact v7.0.0 → v7.0.1
    • actions/upload-pages-artifact v4.0.0 → v5.0.0

    All updates retain SHA-pinned references for supply-chain integrity.

  • New informational Clippy (beta) job — runs the same cargo clippy --all-targets --features duckdb-1-5 -- -D warnings invocation on the beta Rust toolchain. Marked continue-on-error so a beta-only lint regression does not block the merge queue, but surfaces six weeks before the lint reaches stable. Originally added in response to clippy::map_unwrap_or graduating to stable in Rust 1.95.0 and biting src/warning.rs after the toolchain rolled forward.

Fixed

  • clippy::map_unwrap_or on WarningCollector::len — self.warnings.lock().map(|w| w.len()).unwrap_or(0) rewritten as self.warnings.lock().map_or(0, |w| w.len()). Behaviour-preserving; fixes Clippy and Test duckdb-1-5 feature jobs under Rust 1.95.0.
  • clippy::map_unwrap_or_default on WarningCollector::snapshot (defensive) — same rewrite for the sibling map(|w| w.clone()).unwrap_or_default() call. Caught proactively alongside the above; otherwise would have surfaced the next time the lint promotion round-trips through pedantic or nursery.

0.12.0 — 2026-04-09

Added

  • TypedTableFunctionBuilder<S> with closure-based bind/scan — new high-level layer on top of TableFunctionBuilder that lets extensions register table functions via two safe Rust closures instead of hand-rolled unsafe extern "C" fn trampolines. Entry point is TableFunctionBuilder::with_state::<S, _>(|bind| Ok(S { ... })), followed by .scan(|state, chunk| { ... Ok(()) }) and .build()? to recover a fully configured TableFunctionBuilder usable with any Registrar. Highlights:

    • The bind closure receives &BindInfo, declares the output schema, reads parameters, and returns the typed scan state S: Send + 'static.
    • The scan closure receives &mut S and a DataChunk for the output chunk. Returning with chunk size zero signals end-of-stream.
    • Panics in user closures are caught via std::panic::catch_unwind and surfaced through duckdb_bind/init/function_set_error; the scan forces chunk size to zero on panic so the query terminates safely.
    • Scan state is carried from bind through init into init_data so the scan callback can hold &mut S without extra ceremony.
    • Because S is only required to be Send, scans are serialised by calling set_max_threads(1). Extensions that need true multi-worker parallelism should continue to use the raw TableFunctionBuilder with local_init.
    • Re-exported from the prelude as TypedTableFunctionBuilder.

    This is proposal A from the duck_net "quack-rs enhancements" list and eliminates the raw bind/init/scan trampolines that every FFI-heavy extension would otherwise write by hand.

  • ExtensionError: additional From impls — From<std::io::Error>, From<std::ffi::NulError>, and From<std::fmt::Error> allow the ? operator to propagate common error types directly in register_all() without .map_err(). This eliminates the need for panic!() when operations like tokio runtime allocation fail during extension initialization.

  • tls module — TlsConfigProvider trait for type-erased TLS client configuration injection. HTTP-capable extensions (e.g., duck_net) implement this trait to supply custom CA bundles, client certificates for mTLS, or restricted cipher suites through a uniform interface. Uses std::any::Any so quack-rs has no dependency on any specific TLS library. Security hardened: client_config() returns Result for fallible config creation, accepts_invalid_certs() and min_tls_version() enable security auditing, config_type_name() allows safe pre-downcast verification, and audit_tls_provider() integrates with the warning module to automatically flag CWE-295 (cert validation bypass) and CWE-327 (deprecated TLS versions). Includes TlsVersion enum with is_deprecated() and Ord ordering.

  • warning module — structured security warning API with ExtensionWarning, WarningSeverity (Info/Low/Medium/High/Critical), and WarningCollector. Extensions that touch external resources emit warnings with machine-readable codes and optional CWE identifiers. WarningCollector is thread-safe (Mutex-backed) and supports emit(), snapshot(), drain(), and clear().

  • secrets module — SecretsManager trait and SecretEntry type for bridging into DuckDB's native CREATE SECRET storage. Extensions implement SecretsManager to provide get_secret(), list_secrets(), and remove_secret() through a safe Rust interface. SecretEntry uses a builder pattern with with_provider(), with_scope(), and with_field(). Security hardened: Debug redacts field values, Drop zeroizes sensitive data via write_volatile, PartialEq intentionally omitted to prevent timing side-channels, and fields are private with accessor methods.

  • StructWriter::child_list_vector(field_idx) — semantic alias for child_vector() that makes the intent clear when a struct field has LIST type. Returns the raw duckdb_vector handle for use with ListVector methods (reserve, set_entry, set_size, child_writer, etc.).

  • Prelude additions — TlsConfigProvider, ExtensionWarning, WarningSeverity, WarningCollector, SecretEntry, SecretsManager re-exported from quack_rs::prelude.

0.11.0 — 2026-03-30

Added

  • StructWriter::child_vector(field_idx) — returns the raw duckdb_vector handle for a struct field, enabling ListVector/MapVector/ArrayVector operations on nested complex types without raw FFI calls.

  • StructReader::child_vector(field_idx) — read-side counterpart for accessing nested complex type fields within STRUCT input vectors.

  • ChunkWriter::vector(col_idx) — raw duckdb_vector access for complex column types (LIST, MAP, ARRAY) from within a ChunkWriter.

  • ChunkWriter::column_count() — returns the number of columns in the chunk without needing a separate DataChunk.

  • VectorWriter::set_valid(row) — marks a row as non-NULL, undoing a previous set_null() call. Calls ensure_validity_writable automatically.

  • StructWriter::set_valid(row, field_idx) — batched version of VectorWriter::set_valid() for STRUCT fields.

  • ReplacementScanInfo::add_parameter_raw(value) — adds any duckdb_value as a parameter to a replacement scan redirect, enabling non-VARCHAR parameter types (INTEGER, BIGINT, BOOLEAN, etc.).

  • ReplacementScanInfo::add_i64_parameter(value) — convenience method for adding BIGINT parameters to replacement scan redirects.

  • ReplacementScanInfo::add_bool_parameter(value) — convenience method for adding BOOLEAN parameters to replacement scan redirects.

Changed

  • table_scan_callback! error reporting — the macro now extracts the panic message and reports it to DuckDB via duckdb_function_set_error before setting chunk size to 0. Previously, panics silently ended the stream with no error message visible to the user.

0.10.0 — 2026-03-29

Added

  • StructWriter (vector::struct_writer module) — batched, typed writer for STRUCT output vectors. Pre-creates VectorWriters for all fields at construction, then exposes write_bool, write_varchar, write_i64, write_date, write_timestamp, write_time, write_blob, write_uuid, set_null, etc. Eliminates ~120 raw duckdb_struct_vector_get_child calls across typical extensions.

  • StructReader (vector::struct_reader module) — batched, typed reader for STRUCT input vectors. Read-side counterpart to StructWriter with read_bool, read_str, read_i64, read_date, read_timestamp, read_blob, read_uuid, is_valid, etc.

  • ChunkWriter (chunk_writer module) — auto-sizing chunk writer for table function scan callbacks. Tracks rows via next_row() and automatically calls duckdb_data_chunk_set_size on Drop, preventing forgotten-set-size bugs.

  • scalar_callback! / table_scan_callback! macros (callback module) — wrap unsafe extern "C" callbacks with std::panic::catch_unwind, preventing undefined behaviour from panics unwinding across the FFI boundary. Scalar errors are reported via duckdb_scalar_function_set_error; table scan panics set chunk size to 0 (end of stream).

  • Value extraction methods — as_i8(), as_i16(), as_u8(), as_u16(), as_u32(), as_u64(), as_i128() covering every DuckDB integer type via duckdb_get_int8/int16/uint8/uint16/uint32/uint64/hugeint. Plus as_str_or(), as_str_or_default(), and _or(default) null-safe variants for all types.

  • VectorReader — read_date(), read_timestamp(), read_time(), read_blob(), read_uuid() semantic methods for DATE, TIMESTAMP, TIME, BLOB, and UUID column types.

  • VectorWriter — write_date(), write_timestamp(), write_time(), write_blob(), write_uuid() semantic methods matching reader additions.

  • DataChunk convenience methods — struct_writer(col, fields), struct_reader(col, fields), struct_field_reader(col, field), into_chunk_writer() bridging to the new StructWriter, StructReader, and ChunkWriter types.

  • ChunkWriter::struct_writer(col, fields) — convenience bridge to StructWriter from within a ChunkWriter.

  • MockVectorWriter — write_blob(), write_date(), write_timestamp(), write_time(), write_uuid() matching real VectorWriter additions.

  • MockVectorWriter / MockVectorReader — try_get_i8(), try_get_i16(), try_get_u8(), try_get_u16(), try_get_u32(), try_get_u64(), try_get_f32(), try_get_i128(), try_get_blob(), try_get_uuid() closing the type coverage asymmetry between mock and real vector types.

  • MockVectorReader constructors — from_i8s(), from_i16s(), from_u8s(), from_u16s(), from_u32s(), from_u64s(), from_f32s(), from_i128s(), from_intervals(), from_blobs() for every MockDuckValue variant.

  • MockDuckValue::Blob(Vec<u8>) — new variant for BLOB testing.

  • Prelude additions — StructReader, StructWriter, ChunkWriter re-exported.

Changed

  • TableDescription::column_type() now returns Option<LogicalType> (RAII) instead of raw duckdb_logical_type, eliminating manual destroy calls by callers.

  • Version references updated — all documentation, examples, scaffold templates, and book pages now reference quack-rs = "0.10" (was "0.9").

Fixed

  • FFI callback panic safety — replaced 13 CString::new(...).expect(...) calls in FFI callback contexts (table/info, scalar/info, cast/builder, aggregate/info, copy_function/info, replacement_scan) with non-panicking str_to_cstring() that truncates at interior null bytes. Fully honours the "no panics across FFI" design principle (Pitfall L3).

  • Non-idiomatic &mut { expr } syntax — replaced 8 instances in builder register() methods and 1 in replacement scan with idiomatic &raw mut.

0.9.0 — 2026-03-29

Added

  • Value RAII wrapper (value module) — owned wrapper around duckdb_value with automatic cleanup via Drop. Typed extraction methods: as_str(), as_i64(), as_i32(), as_f64(), as_f32(), as_bool(). Eliminates manual duckdb_destroy_value calls and prevents memory leaks in bind parameter extraction.

  • DataChunk wrapper (data_chunk module) — ergonomic non-owning wrapper around duckdb_data_chunk with reader(col), writer(col), size(), set_size(n), column_count(), and vector(col) methods. Eliminates raw duckdb_data_chunk_get_vector / duckdb_data_chunk_set_size calls in scan callbacks.

  • VectorWriter::write_str(idx, value) — alias for write_varchar for discoverability. Extension authors searching for write_str now find it immediately.

  • BindInfo::get_parameter_value(index) — returns an owned Value instead of a raw duckdb_value, preventing memory leaks.

  • BindInfo::get_named_parameter_value(name) — same for named parameters.

  • MapVector::key_writer(vector) / value_writer(vector) — create VectorWriter instances for MAP key and value child vectors directly.

  • MapVector::key_reader(vector, count) / value_reader(vector, count) — create VectorReader instances for MAP key and value child vectors.

  • MockVectorWriter::write_str(idx, value) — alias for write_varchar matching the VectorWriter API addition.

  • Prelude additions — Value, DataChunk, and ValidityBitmap are now re-exported from quack_rs::prelude.

Changed

  • Version references updated — all documentation, examples, scaffold templates, and book pages now reference quack-rs = "0.9" (was "0.7").

0.8.0 — 2026-03-28

Added

  • LogicalType::from_raw(ptr) — construct a LogicalType from an existing raw duckdb_logical_type handle, taking ownership.

  • LogicalType complex type constructors — decimal(width, scale), array(element, size), array_from_logical(element, size), union_type(members), union_type_from_logical(members), enum_type(members).

  • LogicalType _from_logical variants — struct_type_from_logical, list_from_logical, map_from_logical accept LogicalType values for nested complex types that cannot be expressed as simple TypeId.

  • LogicalType introspection methods (20 methods) — get_type_id, get_alias, set_alias, decimal_width, decimal_scale, decimal_internal_type, enum_internal_type, enum_dictionary_size, enum_dictionary_value, list_child_type, map_key_type, map_value_type, struct_child_count, struct_child_name, struct_child_type, union_member_count, union_member_name, union_member_type, array_size, array_child_type.

  • TypeId::from_duckdb_type(raw) — reverse conversion from raw DUCKDB_TYPE C enum to TypeId.

  • ScalarFunctionBuilder::extra_info(data, destroy) — attach arbitrary data to a scalar function, accessible via duckdb_function_get_extra_info in callbacks.

  • ScalarOverloadBuilder::extra_info(data, destroy) — same for scalar function set overloads.

  • AggregateFunctionBuilder::extra_info(data, destroy) — attach arbitrary data to an aggregate function.

  • TableFunctionBuilder::param_logical(logical_type) — add a positional parameter with a complex LogicalType.

  • TableFunctionBuilder::named_param_logical(name, logical_type) — add a named parameter with a complex LogicalType.

  • CastFunctionBuilder::new_logical(source, target) — construct a cast builder using LogicalType values for complex source/target types.

  • ScalarFunctionInfo — callback wrapper with get_extra_info(), set_error(), and (duckdb-1-5) get_bind_data(), get_state().

  • ScalarBindInfo (duckdb-1-5) — scalar bind callback wrapper with argument_count(), get_argument(), get_extra_info(), set_bind_data(), set_error(), get_client_context().

  • ScalarInitInfo (duckdb-1-5) — scalar init callback wrapper with get_extra_info(), get_bind_data(), set_state(), set_error(), get_client_context().

  • AggregateFunctionInfo — aggregate callback wrapper with get_extra_info() and set_error().

  • CopyBindInfo (duckdb-1-5) — copy bind callback wrapper with column_count(), column_type(), get_extra_info(), set_bind_data(), set_error(), get_client_context().

  • CopyGlobalInitInfo (duckdb-1-5) — copy global init callback wrapper with get_bind_data(), get_extra_info(), get_file_path(), set_global_state(), set_error(), get_client_context().

  • CopySinkInfo (duckdb-1-5) — copy sink callback wrapper with get_bind_data(), get_extra_info(), get_global_state(), set_error(), get_client_context().

  • CopyFinalizeInfo (duckdb-1-5) — copy finalize callback wrapper with get_bind_data(), get_extra_info(), get_global_state(), set_error(), get_client_context().

  • BindInfo::get_parameter(index) — retrieve positional parameter value in table function bind callbacks.

  • BindInfo::get_named_parameter(name) — retrieve named parameter value in table function bind callbacks.

  • BindInfo::get_extra_info(), InitInfo::get_extra_info(), FunctionInfo::get_extra_info() — access extra info from table function callbacks.

  • get_client_context() — available on BindInfo (table), ScalarBindInfo, ScalarInitInfo, CopyBindInfo, CopyGlobalInitInfo, CopySinkInfo, CopyFinalizeInfo. Returns a ClientContext RAII wrapper.

  • ArrayVector — helper for fixed-size array vectors with get_child().

  • vector_size() — returns the default DuckDB vector size (typically 2048).

  • vector_get_column_type(vector) — returns the LogicalType of a vector.

  • Prelude additions — StructVector, ListVector, MapVector, ArrayVector, ScalarFunctionInfo, AggregateFunctionInfo now re-exported from quack_rs::prelude.

Changed

  • CastFunctionBuilder::source() / target() now return Option<TypeId> instead of TypeId, returning None when the builder was created via new_logical(). This is a breaking change.

  • CastRecord::source / target fields changed from TypeId to Option<TypeId> to match the builder change.

0.7.1 — 2026-03-27

Added

  • TypeId::Any — wildcard type for function overload resolution. Maps to DUCKDB_TYPE_ANY in the C API. Requires duckdb-1-5 feature.

  • TypeId::Varint — variable-length arbitrary-precision integer. Maps to DUCKDB_TYPE_BIGNUM in the C API, exposed as VARINT in SQL. Requires duckdb-1-5 feature.

  • TypeId::SqlNull — explicit SQL NULL type representing the type of a bare NULL literal before type resolution. Maps to DUCKDB_TYPE_SQLNULL in the C API. Requires duckdb-1-5 feature.

  • TypeId::IntegerLiteral — internal type for unresolved integer literals during overload resolution. Maps to DUCKDB_TYPE_INTEGER_LITERAL. Requires duckdb-1-5 feature.

  • TypeId::StringLiteral — internal type for unresolved string literals during overload resolution. Maps to DUCKDB_TYPE_STRING_LITERAL. Requires duckdb-1-5 feature.

  • MockVectorReader/MockVectorWriter tests — 12 new tests covering from_i32s, from_f64s, from_bools constructors, typed getters (i32, f64, bool), u16/i128/interval round-trips, wrong-type returns None, and is_empty.

  • DuckDB v1.5.1 compatibility evaluation — comprehensive analysis of all 80+ changes in DuckDB v1.5.1 against quack-rs. See docs/duckdb-v1.5.1-evaluation.md.

Fixed

  • ARM64 / aarch64 build — replaced all .cast::<i8>() and *const i8 pointer casts with std::os::raw::c_char, which resolves to i8 on x86-64 and u8 on ARM64 (where C char is unsigned). Eliminates E0308/E0277 mismatched-types errors when cross-compiling or building natively on aarch64. Affected files: replacement_scan/mod.rs, types/logical_type.rs, vector/writer.rs.

Changed

  • DuckDB v1.5.1 compatibility — updated DUCKDB_API_VERSION doc comment and version range documentation to explicitly cover v1.5.1. The C API version remains "v1.2.0" (unchanged from v1.5.0). Users are strongly recommended to upgrade their DuckDB runtime to v1.5.1 for critical WAL corruption and ART index correctness fixes.

Internal

  • CI action update — dtolnay/rust-toolchain pinned to 631a55b12751854ce901bb631d5902ceb48146f7 (PR #59).

  • Mutation testing — mutants.toml now sets features = ["duckdb-1-5"] so that cargo mutants compiles and tests feature-gated code paths. Previously, four mutants in MockRegistrar::copy_function_names, has_copy_function, and total_registrations were unreachable because their tests were also feature-gated. Added mock_registrar_total_registrations_scalar_plus_copy_function to robustly kill the + with - mutation in total_registrations by using a non-zero base count.

0.7.0 — 2026-03-22

Added

  • duckdb-1-5 feature modules — the duckdb-1-5 feature flag is no longer a placeholder. When enabled, it gates five new modules wrapping DuckDB 1.5.0 C Extension API additions:

    • catalog — catalog entry lookup (CatalogEntry, Catalog, CatalogEntryType)
    • client_context — client context access (ClientContext) for retrieving catalogs, config options, and connection IDs from within registered function callbacks
    • config_option — extension-defined configuration options (ConfigOptionBuilder, ConfigOptionScope) registered via SET/RESET/current_setting()
    • copy_function — custom COPY TO handlers (CopyFunctionBuilder) with bind → global init → sink → finalize lifecycle
    • table_description — table metadata queries (TableDescription) for column count, names, and logical types
  • TypeId::TimeNs — new TIME_NS column type variant for nanosecond- precision time of day (DuckDB 1.5.0+, requires duckdb-1-5 feature)

  • ScalarFunctionBuilder::varargs() / varargs_logical() — mark a scalar function as accepting variadic arguments (requires duckdb-1-5)

  • ScalarFunctionBuilder::volatile() — mark a scalar function as volatile (re-evaluated for every row even with constant arguments, requires duckdb-1-5)

  • ScalarFunctionBuilder::bind() — set a bind callback invoked once during query planning for per-query state allocation (requires duckdb-1-5)

  • ScalarFunctionBuilder::init() — set an init callback invoked once per thread for per-thread local state allocation (requires duckdb-1-5)

Changed

  • DuckDB 1.5.0 support — upgraded default libduckdb-sys from 1.4.4 to 1.10500.0 (DuckDB 1.5.0) and duckdb from 1.4.4 to 1.10500.0. The version range ">=1.4.4, <2" in Cargo.toml is unchanged, preserving backward compatibility with DuckDB 1.4.x.

  • Transitive dependency updates — cc 1.2.56→1.2.57, tar 0.4.44→0.4.45, rustls-webpki 0.103.9→0.103.10, arrow 56.2.0→57.3.0, clap 4.5.60→4.6.0, tempfile 3.14.0→3.27.0, plus ~30 other minor/patch updates.

  • CI action updates — Swatinem/rust-cache v2.8.2→v2.9.1, actions/download-artifact v8.0.0→v8.0.1, actions/cache 5.0.3→5.0.4, codecov/codecov-action 5.4.3→5.5.3.

Fixed

  • COPY format handlers — previously listed as a known limitation (no C API counterpart). DuckDB 1.5.0 adds duckdb_create_copy_function and related symbols; the new copy_function module wraps them behind duckdb-1-5.

0.6.0 — 2026-03-12

Added

  • InMemoryDb dispatch table initialisation — InMemoryDb::open() now correctly initialises the loadable-extension dispatch table from bundled DuckDB symbols before opening a connection, allowing all three InMemoryDb unit tests to pass under cargo test --features bundled-test. Previously every call to InMemoryDb::open() panicked with "DuckDB API not initialized or DuckDB feature omitted" because the loadable-extension dispatch table was never populated in cargo test.

  • src/testing/bundled_api_init.cpp — thin C++ shim that wraps DuckDB's internal CreateAPIv1() function (from duckdb/main/capi/extension_api.hpp) as a C-linkage symbol (quack_rs_create_api_v1). Called once at test startup to populate all 459 AtomicPtr slots in the dispatch table with real bundled DuckDB function pointers.

  • build.rs — Cargo build script that, when the bundled-test feature is active, locates the libduckdb-sys build output directory, finds the bundled DuckDB include path, and compiles bundled_api_init.cpp via the cc crate.

  • CI: test-bundled job — new CI job runs cargo test --all-targets --features bundled-test on all three platforms (Linux, macOS, Windows) on every push and pull request, closing the gap that allowed this failure to reach the release workflow undetected.

  • Pitfall P9 documented — LESSONS.md now contains a full analysis of the loadable-extension dispatch table failure mode: root cause, the CreateAPIv1() solution, ABI compatibility details, risks of relying on DuckDB's internal C++ header, and a mitigation table.

Fixed

  • InMemoryDb::open() no longer panics when called in cargo test with the bundled-test feature enabled. This was a regression introduced when InMemoryDb was first shipped in 0.5.1 without the dispatch table initialisation step.

Changed

  • bundled-test feature documentation updated to accurately describe the dispatch table initialisation behaviour (previously claimed to "bypass" the dispatch mechanism; it now correctly initialises it).

0.5.1 — 2026-03-12

Added

  • Testing primitives (quack_rs::testing) — new mock types for unit-testing extension logic without a live DuckDB process:

    • MockVectorWriter — in-memory output buffer matching the VectorWriter API; use to test scalar/aggregate finalize/scan callbacks
    • MockVectorReader — in-memory input buffer with convenience constructors (from_i64s, from_strs, from_bools, from_f64s, from_i32s)
    • MockDuckValue — typed enum covering all DuckDB scalar types
    • MockRegistrar — implements the Registrar trait using interior mutability; records registered functions without any C API call
    • CastRecord — records source/target types for cast registrations
  • bundled-test Cargo feature — links the bundled DuckDB static library via the duckdb crate and enables InMemoryDb::open() for SQL-level assertions in cargo test. Does not initialize the loadable-extension dispatch table.

  • InMemoryDb — wraps duckdb::Connection for SQL-level integration tests; available behind the bundled-test feature.

  • Builder introspection accessors — pub fn name(&self) -> &str added to ScalarFunctionBuilder, ScalarFunctionSetBuilder, AggregateFunctionBuilder, AggregateFunctionSetBuilder, and TableFunctionBuilder. pub fn source(&self) -> Option<TypeId> and pub fn target(&self) -> Option<TypeId> added to CastFunctionBuilder.

Security

  • Bump quinn-proto 0.11.13 → 0.11.14 in root and examples/hello-ext Cargo.lock files (addresses RUSTSEC advisory).

0.5.0 — 2026-03-10

Added

  • param_logical(LogicalType) on all builders — register parameters with complex parameterized types (LIST(BIGINT), MAP(VARCHAR, INTEGER), STRUCT(...)) that TypeId alone cannot express. Available on AggregateFunctionBuilder, AggregateFunctionSetBuilder::OverloadBuilder, ScalarFunctionBuilder, and ScalarOverloadBuilder. Parameters added via param() and param_logical() are interleaved by position, so the order you call them is the order DuckDB sees them.

  • returns_logical(LogicalType) on all builders — set a complex parameterized return type. When both returns(TypeId) and returns_logical(LogicalType) are called, the logical type takes precedence. Available on AggregateFunctionBuilder, AggregateFunctionSetBuilder, ScalarFunctionBuilder, and ScalarOverloadBuilder. This eliminates the need for raw FFI when returning LIST(BOOLEAN), LIST(TIMESTAMP), MAP(K, V), or any other parameterized type.

  • null_handling(NullHandling) on set overload builders — per-overload NULL handling configuration for AggregateFunctionSetBuilder::OverloadBuilder and ScalarOverloadBuilder. Previously only available on single-function builders.

Notes

  • Upstream fix: duckdb-loadable-macros panic-at-FFI-boundary — the safe entry-point pattern developed in quack-rs (using ? / ok_or_else throughout instead of .unwrap()) was contributed upstream as duckdb/duckdb-rs#696 and merged 2026-03-09. All users of the duckdb_entrypoint_c_api! macro from duckdb-loadable-macros will receive this fix in the next duckdb-rs release. quack-rs users have always been protected via the safe entry_point! / entry_point_v2! macros provided by this crate.

0.4.0 — 2026-03-09

Added

  • Connection and Registrar trait — version-agnostic extension registration facade (src/connection.rs). Connection wraps the duckdb_connection and duckdb_database handles provided at initialization time. The Registrar trait provides uniform methods for registering all extension components (scalar, scalar set, aggregate, aggregate set, table, SQL macro, cast), making registration code interchangeable across DuckDB 1.4.x and 1.5.x. Replacement scans are exposed as direct methods on Connection since they require duckdb_database, not the connection handle.

  • init_extension_v2 — new entry point helper that passes &Connection to the registration callback instead of a raw duckdb_connection. Prefer this over init_extension for new extensions.

  • entry_point_v2! macro — companion macro to entry_point! that generates the #[no_mangle] unsafe extern "C" entry point using init_extension_v2.

  • duckdb-1-5 cargo feature — placeholder feature flag for DuckDB 1.5.0-specific C API wrappers. Currently empty; will be populated when libduckdb-sys 1.5.0 is published on crates.io.

Changed

  • DuckDB version support broadened to 1.4.x and 1.5.x — the libduckdb-sys dependency requirement was relaxed from an exact pin (=1.4.4) to a range (>=1.4.4, <2). DuckDB v1.5.0 (released 2026-03-09) does not change the C API version string (v1.2.0) used in duckdb_rs_extension_api_init; the existing DUCKDB_API_VERSION constant remains correct for both releases. Extension authors can now pin their own libduckdb-sys to either =1.4.4 or =1.5.0 and resolve cleanly against quack-rs. The scaffold template and CI workflow template were updated to default to DuckDB v1.5.0.

0.3.0 — 2026-03-08

Added

  • TableFunctionBuilder — type-safe builder for registering DuckDB table functions (the SELECT * FROM my_function(args) pattern). Covers the full bind/init/scan lifecycle with ergonomic callbacks, eliminating ~100 lines of raw FFI boilerplate. Helper types BindInfo, FfiBindData<T>, and FfiInitData<T> manage parameter extraction and per-scan state with zero raw pointer manipulation. See table and examples/hello-ext (generate_series_ext) for a fully-tested end-to-end example verified against DuckDB 1.4.4.

  • ReplacementScanBuilder — builder for registering DuckDB replacement scans (the SELECT * FROM 'file.xyz' pattern where a file path triggers a table-valued scan). The builder handles callback registration, path extraction, and bind-info population through a 4-method chain. See replacement_scan.

  • StructVector — safe wrapper for reading and writing STRUCT child vectors. get_child(vec, idx), field_reader(vec, idx, row_count), and field_writer(vec, idx) replace manual offset arithmetic over child vector handles.

  • ListVector — safe wrapper for reading and writing LIST child vectors. get_child, get_entry, set_entry, reserve, set_size, child_reader, and child_writer cover the complete LIST read/write workflow without raw pointer casts.

  • MapVector — safe wrapper for DuckDB MAP vectors (stored as LIST<STRUCT{key, value}>). keys(vec), values(vec), struct_child(vec), reserve, set_size, set_entry, and get_entry expose the full MAP interface.

  • vector::complex module — re-exports StructVector, ListVector, MapVector at quack_rs::vector::complex and documents the read-vs-write workflow for nested types with working code examples in the module doc.

  • prelude additions — TableFunctionBuilder, BindInfo, FfiBindData, FfiInitData, ReplacementScanBuilder, StructVector, ListVector, MapVector, CastFunctionBuilder, CastFunctionInfo, CastMode are now all re-exported from quack_rs::prelude.

  • CastFunctionBuilder — type-safe builder for registering custom type cast functions via duckdb_cast_function_*. Covers both explicit CAST(x AS T) and implicit coercions (with optional implicit_cost). The companion CastFunctionInfo wrapper exposes cast_mode(), set_error(), and set_row_error() inside callbacks, giving correct TRY_CAST / CAST error handling with zero raw pointer boilerplate. See cast for the full API.

  • DbConfig — RAII wrapper for duckdb_config (extension configuration parameters). Builder-style .set(name, value)? chain, automatic duckdb_destroy_config on drop, and flag_count() / get_flag(index) for enumerating all available options. Useful when an extension needs to open a secondary DuckDB database from within its callbacks. See config.

  • ScalarFunctionSetBuilder — builder for registering scalar function sets (multiple overloads under one name), mirroring AggregateFunctionSetBuilder.

  • TypeId variants — Decimal, Struct, Map, UHugeInt, TimeTz, TimestampS, TimestampMs, TimestampNs, Array, Enum, Union, Bit.

  • From<TypeId> for LogicalType — idiomatic conversion from TypeId.

  • #[must_use] on builder structs — ScalarFunctionBuilder, AggregateFunctionBuilder, AggregateFunctionSetBuilder, and OverloadBuilder now warn at compile time if constructed but never consumed.

  • NullHandling enum and .null_handling() builder method — configurable NULL propagation for scalar and aggregate functions via duckdb_scalar_function_set_special_handling / duckdb_aggregate_function_set_special_handling.

  • VectorWriter::write_interval — writes INTERVAL values to output vectors using the correct 16-byte { months: i32, days: i32, micros: i64 } layout.

  • append_metadata binary — native Rust replacement for the Python append_extension_metadata.py script, now shipping with the crate. Install with cargo install quack-rs --bin append_metadata.

  • hello-ext cast function demo — examples/hello-ext now registers a CAST(VARCHAR AS INTEGER) cast function using CastFunctionBuilder, demonstrating both CAST (abort-on-error) and TRY_CAST (NULL-on-error) code paths. Five unit tests cover parse_varchar_to_int, including boundary values and overflow.

Not implemented (upstream C API gap)

  • Window functions — duckdb_create_window_function and related symbols do not exist in DuckDB's public C extension API. They are implemented only in the C++ layer and are therefore not wrappable by quack-rs or any other C-API binding. Verified against the DuckDB stable C API reference and libduckdb-sys 1.4.4 bindings.

  • COPY format handlers — duckdb_create_copy_function and related symbols are similarly absent from the C extension API for the same reason.

Fixed

  • hello-ext gs_bind callback — replaced incorrect duckdb_value_int64(param) (wrong arity: takes 3 arguments) with duckdb_get_int64(param) (correct 1-argument form). The extension now builds cleanly and all 11 live SQL tests pass against DuckDB 1.4.4.

Changed

  • Bump criterion dev-dependency from 0.5 to 0.8.
  • Bump Swatinem/rust-cache GitHub Action from v2.7.5 to v2.8.2.
  • Bump dtolnay/rust-toolchain CI pin from v2.7.5 to latest SHA.
  • Bump actions/attest-build-provenance from v2 to v4.
  • Bump actions/configure-pages to latest SHA (d5606572…).
  • Bump actions/upload-pages-artifact from v3.0.1 to v4.0.0.

0.2.0 — 2026-03-07

Added

  • validate::description_yml module — parse and validate a complete description.yml metadata file end-to-end. Includes:

    • DescriptionYml struct — structured representation of all required and optional fields
    • parse_description_yml(content: &str) — parse and validate in one step
    • validate_description_yml_str(content: &str) — pass/fail validation
    • validate_rust_extension(desc: &DescriptionYml) — enforce Rust-specific fields (language: Rust, build: cargo, requires_toolchains includes rust)
    • 25+ unit tests covering all required fields, optional fields, error paths, and edge cases
  • prelude module — ergonomic glob-import for the most commonly used items. use quack_rs::prelude::*; brings in all builder types, state traits, vector helpers, types, error handling, and the API version constant. Reduces boilerplate for extension authors.

  • Scaffold: extension_config.cmake generation — the scaffold generator now produces extension_config.cmake, which is referenced by the EXT_CONFIG variable in the Makefile and required by extension-ci-tools for CI integration.

  • Scaffold: SQLLogicTest skeleton — generate_scaffold now produces test/sql/{name}.test, a ready-to-fill SQLLogicTest file with require directive, format comments, and example query/result blocks. E2E tests are required for community extension submission (Pitfall P3).

  • Scaffold: GitHub Actions CI workflow — generate_scaffold now produces .github/workflows/extension-ci.yml, a complete cross-platform CI workflow that builds and tests the extension on Linux, macOS, and Windows against a real DuckDB binary.

  • validate::validate_excluded_platforms_str — validates the excluded_platforms field from description.yml as a semicolon-delimited string (e.g., "wasm_mvp;wasm_eh;wasm_threads"). Splits on ; and validates each token. An empty string is valid (no exclusions).

  • validate::validate_excluded_platforms — re-exported at the validate module level (previously only accessible as validate::platform::validate_excluded_platforms).

  • validate::semver::classify_extension_version — returns ExtensionStability (Unstable/PreRelease/Stable) classifying the tier a version falls into.

  • validate::semver::ExtensionStability — enum for DuckDB extension version stability tiers (Unstable, PreRelease, Stable) with Display implementation.

  • scalar module — ScalarFunctionBuilder for registering scalar functions with the DuckDB C Extension API. Includes try_new with name validation, param, returns, function setters, and register. Full unit tests included.

  • entry_point! macro — generates the required #[no_mangle] extern "C" entry point with zero boilerplate from an identifier and registration closure.

  • VectorWriter::write_varchar — writes VARCHAR string values to output vectors using duckdb_vector_assign_string_element_len (handles both inline and pointer formats).

  • VectorWriter::write_bool — writes BOOLEAN values as a single byte.

  • VectorWriter::write_u16 — writes USMALLINT values.

  • VectorWriter::write_i16 — writes SMALLINT values.

  • VectorReader::read_interval — reads INTERVAL values from input vectors via the correct 16-byte layout helper.

  • CI: Windows testing — the CI matrix now includes windows-latest in the test job, covering all three major platforms (Linux, macOS, Windows).

  • CI: example-check job — CI now checks, lints, and tests examples/hello-ext as part of every PR, ensuring the example extension always compiles and its tests pass.

  • validate::validate_release_profile — checks Cargo release profile settings for loadable-extension correctness. Validates panic, lto, opt-level, and codegen-units.

Fixed

  • MSRV documentation now consistently states 1.84.1 across README.md, CONTRIBUTING.md, and Cargo.toml (previously README.md stated 1.80).

0.1.0 — 2025-05-01

Added

  • Initial release
  • entry_point module: init_extension helper for correct extension initialization
  • aggregate module: AggregateFunctionBuilder, AggregateFunctionSetBuilder
  • aggregate::state module: AggregateState trait, FfiState<T> wrapper
  • aggregate::callbacks module: type aliases for all 6 callback signatures
  • vector module: VectorReader, VectorWriter, ValidityBitmap, DuckStringView
  • types module: TypeId enum, LogicalType RAII wrapper
  • interval module: DuckInterval, interval_to_micros, read_interval_at
  • error module: ExtensionError, ExtResult<T>
  • testing module: AggregateTestHarness<S> for pure-Rust aggregate testing
  • validate module: validate_extension_name, validate_function_name, validate_semver, validate_extension_version, validate_spdx_license, validate_platform, validate_release_profile
  • scaffold module: generate_scaffold for generating complete extension projects
  • sql_macro module: SqlMacro for registering SQL macros without FFI callbacks
  • Complete hello-ext example extension
  • Documentation of all 15 DuckDB Rust FFI pitfalls (LESSONS.md)
  • CI pipeline: check, test, clippy, fmt, doc, MSRV, bench-compile
  • SECURITY.md vulnerability disclosure policy

FAQ

Frequently asked questions about quack-rs and building DuckDB extensions in Rust.


General

What is quack-rs?

quack-rs is a Rust SDK for building DuckDB loadable extensions using DuckDB's pure C Extension API. It provides safe, ergonomic builders for registering scalar functions, aggregate functions, table functions, cast functions, replacement scans, SQL macros, and copy functions (via the duckdb-1-5 feature), along with helpers for reading and writing DuckDB vectors, and utilities for publishing community extensions.

Why does this exist?

Building a DuckDB extension in Rust means solving a set of undocumented FFI problems that each developer otherwise discovers independently. quack-rs documents all 31 known pitfalls, and its API prevents most of them. See the Pitfall Catalog.

What DuckDB version does quack-rs target?

quack-rs requires libduckdb-sys = ">=1.4.4, <2" and supports DuckDB 1.4.x and 1.5.x. CI loads a built extension into DuckDB v1.4.4, v1.5.0, v1.5.5 and the latest release.

The C API version passed to the dispatch-table initializer is "v1.2.0", available as quack_rs::DUCKDB_API_VERSION. Every DuckDB 1.4.x and 1.5.x release loads extensions built for it (1.5.6 declares C API v1.5.6, but still accepts v1.2.0). The C API version is a separate identifier from the DuckDB release and from the libduckdb-sys crate version (e.g. 1.10505.0 for DuckDB 1.5.5).

What is the minimum supported Rust version (MSRV)?

Rust 1.86.0 or later. This is enforced in Cargo.toml with rust-version = "1.86.0".

Is quack-rs production-ready?

It is pre-1.0, and you should judge it against your own requirements rather than take a yes. What the record shows:

  • The API still changes. Minor releases before 1.0 can break it; each one lists its breaking changes, with migration notes, in the changelog.
  • Audits keep finding real defects. The review released as 0.16.0 fixed 24 defects in earlier releases, including two heap-corruption paths. The 0.18.0 audits fixed further soundness holes, process aborts and wrong answers, and found one serious defect in unreleased code before it was published: aggregate functions returned wrong results in release builds with Cargo's default profile. AUDIT.md in the repository records each audit, what it found, and whether each fix was reproduced against a real DuckDB or derived from DuckDB's source.
  • Some limits are DuckDB's. The C API has defects quack-rs can only document or work around; see Known Limitations.

It was extracted from duckdb-behavioral, a DuckDB community extension, where the first 16 of the pitfalls it now documents were discovered. If you ship an extension built on it, run end-to-end tests that load the extension into each DuckDB release you support (see the Testing Guide).


Functions

Can I expose SQL macros as an extension?

Yes, without any C++ wrapper code. Use quack_rs::sql_macro::SqlMacro:

#![allow(unused)]
fn main() {
use libduckdb_sys::duckdb_connection;
fn demo(con: duckdb_connection) -> Result<(), quack_rs::error::ExtensionError> {
use quack_rs::sql_macro::SqlMacro;

// Scalar macro
let m = SqlMacro::scalar("double_it", &["x"], "x * 2")?;
unsafe { m.register(con) }?;

// Table macro
let m = SqlMacro::table("recent_events", &["n"],
    "SELECT * FROM events ORDER BY ts DESC LIMIT n")?;
unsafe { m.register(con) }?;
Ok(())
}
}

Register them in your registration closure (the one passed to entry_point! or init_extension) alongside your other functions. A table macro's body is bound when it is created, so the events table must already exist or register returns an error. See SQL Macros.

Can I register multiple overloads of the same function?

Yes, using AggregateFunctionSetBuilder (for aggregates) or ScalarFunctionSetBuilder (for scalars). Both support complex parameter types via param_logical(LogicalType) and complex return types via returns_logical(LogicalType), and in both, each overload may return a different type — DuckDB resolves an overload from its parameter types and arity alone. See Overloading with Function Sets.

Can I register multiple functions in one extension?

Yes. The registration closure receives a duckdb_connection and can register as many functions as needed:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_connection, duckdb_extension_access, duckdb_extension_info};
use quack_rs::error::ExtensionError;
use quack_rs::sql_macro::SqlMacro;
use quack_rs::DUCKDB_API_VERSION;
unsafe fn register_word_count(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
unsafe fn register_sentence_count(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
unsafe fn demo(info: duckdb_extension_info, access: *const duckdb_extension_access) -> bool {
quack_rs::entry_point::init_extension(info, access, DUCKDB_API_VERSION, |con| {
    unsafe { register_word_count(con) }?;
    unsafe { register_sentence_count(con) }?;
    unsafe {
        SqlMacro::scalar("double_it", &["x"], "x * 2")?
            .register(con)?;
    }
    Ok(())
})
}
}

Can I use the duckdb crate instead of libduckdb-sys?

No. The duckdb crate's bundled feature embeds its own copy of DuckDB. A loadable extension must link against the DuckDB that loads it, not bundle a separate copy. Use libduckdb-sys with the loadable-extension feature.

Can I have a scalar function with no parameters?

Yes. Just do not call param:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::prelude::*;
unsafe extern "C" fn quack_callback(_: duckdb_function_info, _: duckdb_data_chunk, _: duckdb_vector) {}
fn demo(con: duckdb_connection) -> Result<(), ExtensionError> {
unsafe {
    ScalarFunctionBuilder::new("current_quack")
        .returns(TypeId::Varchar)
        .function(quack_callback)
        .register(con)?;
}
Ok(())
}
}

Testing

Do I need a DuckDB instance to run unit tests?

No. AggregateTestHarness simulates the aggregate lifecycle in pure Rust without any DuckDB dependency. You can run cargo test without loading a DuckDB binary.

My unit tests all pass but the extension crashes. Why?

Unit tests cannot detect FFI wiring bugs. See Pitfall P3 and the Testing Guide. Always run end-to-end tests that load the packaged extension into a real DuckDB process.

How do I test SQL macros?

SqlMacro::to_sql() is pure Rust and requires no DuckDB connection:

#![allow(unused)]
fn main() {
use quack_rs::sql_macro::SqlMacro;
let m = SqlMacro::scalar("triple", &["x"], "x * 3").unwrap();
assert_eq!(m.to_sql(), r#"CREATE OR REPLACE MACRO "triple"("x") AS (x * 3)"#);
}

For an end-to-end test, call the macro from your SQLLogicTest file:

query I
SELECT double_it(21);
----
42

Publishing

How do I publish to the DuckDB community extensions registry?

  1. Scaffold your project with generate_scaffold
  2. Push to GitHub
  3. Submit a pull request to the community-extensions repo with your description.yml

See Community Extensions for the full workflow.

My extension name is taken. What should I do?

Use a vendor-prefixed name: myorg_analytics instead of analytics. Extension names must be unique among DuckDB community extensions; check community-extensions.duckdb.org before you pick one.

Do I need to set up CI manually?

No. generate_scaffold produces .github/workflows/extension-ci.yml which builds and tests your extension on Linux, macOS, and Windows automatically.

Can my extension be installed with INSTALL ... FROM community?

Yes, once your pull request is merged into the community-extensions repository. Until then, users load the .duckdb_extension file directly, starting DuckDB with -unsigned because a local build is not signed:

LOAD './path/to/my_extension.duckdb_extension';

Troubleshooting

My aggregate returns wrong results with no error.

The most common cause is Pitfall L1: your combine callback is not propagating all configuration fields. See Pitfall L1 and test with AggregateTestHarness::combine.

The NULLs my function writes come back as values.

You are likely calling duckdb_vector_get_validity without first calling duckdb_vector_ensure_validity_writable. A vector that has never held a NULL has no validity mask, so duckdb_vector_get_validity returns a null pointer and duckdb_validity_set_row_invalid silently does nothing (dereferencing that pointer yourself crashes). Use VectorWriter::set_null instead. See Pitfall L4.

My function is not found in SQL after LOAD.

If LOAD succeeded, the entry point ran, so check the name you gave the builder. If you build a function set through the raw C API instead of quack-rs's set builders, every member needs its own name, or DuckDB silently skips it (Pitfall L6).

If LOAD itself fails, check the entry point symbol. DuckDB takes the file's base name, lowercases it and calls <base name>_init_c_api, so my_extension.duckdb_extension needs my_extension_init_c_api (Pitfall P1).

make configure fails with a missing file error.

The extension-ci-tools submodule is missing. In a new project, such as one fresh from the scaffold, add it once (git submodule update --init does nothing until it has been added):

git submodule add https://github.com/duckdb/extension-ci-tools.git extension-ci-tools

In a clone of a repository that already has the submodule:

git submodule update --init --recursive

See Pitfall P4.

My SQLLogicTest fails in CI but passes locally.

SQLLogicTest does exact string matching. The most common issue is a difference in NULL representation, decimal places, or line endings. Run the query in the same DuckDB version used by CI and copy the output verbatim.

How do I read a VARCHAR that is longer than 12 bytes?

VectorReader::read_str handles both the inline (≤ 12 bytes) and pointer (> 12 bytes) formats automatically. No special handling needed.

What happens if I read from a NULL row?

You get garbage data from the vector's data buffer — and for VARCHAR or BLOB, possibly a stale pointer into freed memory. Always check is_valid before reading. See NULL Handling & Strings.


Architecture

Why use libduckdb-sys with loadable-extension instead of the duckdb crate?

The duckdb crate is designed for embedding DuckDB, not for extending it. Its bundled feature includes a statically linked DuckDB binary, which conflicts with the DuckDB runtime that loads your extension. libduckdb-sys with loadable-extension provides lazy-initialized function pointers that are populated by DuckDB at extension load time.

Why not use duckdb-loadable-macros?

duckdb-loadable-macros relies on extract_raw_connection which uses the internal Rc<RefCell<InnerConnection>> layout. This is fragile and causes SEGFAULTs when the layout changes between duckdb crate versions. init_extension uses the correct C API entry sequence directly.

Why must the release profile use panic = "unwind"?

quack-rs wraps every entry point and every callback it generates in catch_unwind and reports a panic as an ordinary SQL error (a raw extern "C" callback you write yourself needs a *_callback! macro or catch_ffi_panic; see Installation). catch_unwind cannot catch anything under panic = "abort": the process terminates at the panic site, taking the user's DuckDB session with it. validate_release_profile rejects abort for this reason. Returning Result and using ? is still the right style — the guards are a safety net, not a substitute.

Can I use async Rust in my extension?

Not directly in FFI callbacks. DuckDB's callbacks are synchronous C functions. You can run an async runtime such as Tokio and block on async tasks inside callbacks (for example with Runtime::block_on), but the callbacks themselves must return synchronously.

How does FfiState<T> prevent double-free?

Each state slot starts with a tag that init_callback writes once the T is in place. destroy_callback drops the T (freeing its box, when T is too large to store inline) only when the tag matches, and clears the tag first. A second call to destroy_callback on the same state finds no tag and does nothing.

Contributing

This page describes how to build, test and contribute to quack-rs: the toolchain, the quality gates every pull request must pass, the test strategy, code standards, the repository layout and the release policy. Bug reports, documentation fixes, newly discovered pitfalls and code are all welcome.


Development prerequisites

ToolVersionPurpose
Rust≥ 1.86.0 (MSRV)Compiler
rustfmtstableFormatting
clippystableLinting
cargo-msrvlatestMSRV verification

Install the Rust toolchain via rustup.rs.


Building

# Build the library
cargo build

# Build in release mode (enables LTO + strip)
cargo build --release

# Build the hello-ext example extension
cargo build --release --manifest-path examples/hello-ext/Cargo.toml

Quality gates

All of the following must pass before merging any pull request:

# Tests — zero failures, zero ignored
cargo test

# Integration tests
cargo test --test integration_test

# Doctests — `--all-targets` does not include them
cargo test --doc --features duckdb-1-5-4

# Linting — zero warnings (warnings are errors)
cargo clippy --all-targets -- -D warnings

# Formatting
cargo fmt -- --check

# Documentation — zero broken links or missing docs
RUSTDOCFLAGS="-D warnings" cargo doc --no-deps

# MSRV — must compile on Rust 1.86.0 (excludes benches; matches CI)
cargo +1.86.0 check

These same checks run in CI on every push and pull request.


Test strategy

Unit tests

Unit tests live in #[cfg(test)] modules alongside the code. They test pure-Rust logic that does not require a live DuckDB instance.

Important constraint: libduckdb-sys with features = ["loadable-extension"] makes all DuckDB C API functions go through lazy AtomicPtr dispatch. These pointers are only populated when duckdb_rs_extension_api_init is called from within a real DuckDB extension load — or by testing::InMemoryDb::open() under the bundled-test / bundled-test-prebuilt features. Without one of those, calling any duckdb_* function in a unit test panics ("DuckDB API not initialized or DuckDB feature omitted"). Put such tests in tests/ffi_roundtrip.rs or its submodules under tests/ffi_roundtrip/, which open an InMemoryDb.

Integration tests

tests/integration_test.rs contains pure-Rust tests that cross module boundaries — testing interval with AggregateTestHarness, verifying FfiState lifecycle, and so on. These still cannot call duckdb_* functions.

Property-based tests

Selected modules include proptest-based tests:

  • interval.rs — overflow edge cases across the full i32/i64 range
  • testing/harness.rs — sum associativity, identity element for AggregateState

Example-extension tests

examples/hello-ext/ contains #[cfg(test)] unit tests for its pure-Rust logic. CI also tests it end to end: it builds the cdylib, appends the metadata footer with append_metadata, and loads it into DuckDB 1.4.4, 1.5.0, 1.5.5 and the latest release. See CONTRIBUTING.md for the exact commands.


Code standards

Safety documentation

Every unsafe block must have a // SAFETY: comment explaining:

  1. Which invariant the caller guarantees
  2. Why the operation is valid given that invariant

clippy::undocumented_unsafe_blocks enforces this in library code (CI treats it as an error); test code is exempt.

#![allow(unused)]
fn main() {
struct Ffi { inner: *mut u64 }
let ffi = Ffi { inner: Box::into_raw(Box::new(0_u64)) };
// SAFETY: `ffi.inner` came from `Box::into_raw` and has not been freed;
// nothing else holds it, so reclaiming and dropping the box is sound.
unsafe { drop(Box::from_raw(ffi.inner)) };
}

No panics across FFI

A panic must never escape a function DuckDB calls (callbacks and entry points): every callback kind runs under catch_unwind. Beyond that, library code uses Option/Result and ?, and panics only where its # Panics section says so — VectorWriter::write_varchar on a string over 4 GiB, or a builder's new(name) on an interior NUL, which has try_new(name) beside it.

Clippy lint policy

The crate enables the all, pedantic, nursery and cargo lint groups, plus undocumented_unsafe_blocks. All warnings are treated as errors in CI. Lints are suppressed only where they produce false positives for SDK API patterns:

[lints.clippy]
module_name_repetitions = "allow"  # e.g., AggregateFunctionBuilder
must_use_candidate = "allow"       # builder methods
missing_errors_doc = "allow"       # unsafe extern "C" callbacks
return_self_not_must_use = "allow" # builder pattern

Documentation

Every public item must have a doc comment. Follow these conventions:

  • First line: a one-sentence summary, ending with a period
  • # Safety: mandatory on every unsafe fn
  • # Panics: mandatory if the function can panic
  • # Errors: mandatory on functions returning Result
  • # Example: encouraged on public types and key methods

Repository structure

quack-rs/
├── src/
│   ├── abi.rs                         # `DuckDB` C Extension API ABI compatibility checking
│   ├── appender.rs                    # Bulk data appending
│   ├── arrow.rs                       # Arrow C Data Interface bridge (`duckdb-1-5-4` feature; floor set by the `libduckdb-sys` 1.10504.0 bindings)
│   ├── callback.rs                    # Panic-safe callback wrapper macros for `DuckDB` extension callbacks
│   ├── catalog.rs                     # Catalog entry lookup (`DuckDB` 1.5.0+)
│   ├── chunk_writer.rs                # Auto-sizing chunk writer for table function scan callbacks
│   ├── client_context.rs              # Client context access (`DuckDB` 1.5.0+)
│   ├── config.rs                      # RAII wrapper for `DuckDB` database configuration
│   ├── config_option.rs               # Extension-defined configuration options (`DuckDB` 1.5.0+)
│   ├── connection.rs                  # [`Connection`] — version-agnostic extension registration facade
│   ├── data_chunk.rs                  # Ergonomic wrapper around `DuckDB` data chunks
│   ├── debug_repr.rs                  # Internal helpers for the crate's `Debug` implementations
│   ├── entry_point.rs                 # Extension entry point helper
│   ├── error.rs                       # Error types for `DuckDB` extension FFI error propagation
│   ├── error_data.rs                  # Structured error data (`DuckDB` 1.5.0+)
│   ├── expression.rs                  # Bound expressions (`DuckDB` 1.5.0+)
│   ├── extra_info.rs                  # Ownership of a function's `extra_info` allocation until `DuckDB` takes it
│   ├── file_system.rs                 # File system access (`DuckDB` 1.5.0+)
│   ├── instance_cache.rs              # Database instance cache (`DuckDB` 1.5.0+)
│   ├── interval.rs                    # `DuckDB` `INTERVAL` type conversion utilities
│   ├── lib.rs                         # Crate root: module declarations and crate-level documentation
│   ├── prelude.rs                     # Convenience re-exports for the most commonly used `quack-rs` items
│   ├── query.rs                       # Running SQL from inside an extension
│   ├── secrets.rs                     # Credential handling for extensions
│   ├── selection_vector.rs            # Selection vectors (`DuckDB` 1.5.0+)
│   ├── sql_macro.rs                   # SQL macro registration for `DuckDB` extensions
│   ├── table_description.rs           # Table description metadata
│   ├── tls.rs                         # Type-erased TLS configuration provider for HTTP-capable extensions
│   ├── value.rs                       # RAII wrapper around `DuckDB` values (`duckdb_value`)
│   ├── warning.rs                     # Structured security warning API for extensions
│   ├── aggregate/
│   │   ├── callbacks.rs               # Type aliases for the five required `DuckDB` aggregate callback signatures
│   │   ├── info.rs                    # Ergonomic wrapper around `duckdb_function_info` for aggregate function callbacks
│   │   ├── mod.rs                     # Builders for registering `DuckDB` aggregate functions
│   │   ├── state.rs                   # Generic `FfiState<T>` wrapper for safe aggregate state management
│   │   └── builder/
│   │       ├── mod.rs                 # Builder types for registering `DuckDB` aggregate functions
│   │       ├── overload.rs            # One overload within an [`AggregateFunctionSetBuilder`]
│   │       ├── set.rs                 # Builder for registering a `DuckDB` aggregate function set (multiple overloads)
│   │       ├── single/
│   │       │   └── register.rs        # `AggregateFunctionBuilder::register`
│   │       ├── single.rs              # Builder for registering a single-signature `DuckDB` aggregate function
│   │       └── tests.rs               # Unit tests
│   ├── appender/
│   │   ├── chunk.rs                   # Chunk-at-a-time appends: handing the [`Appender`] a whole [`DataChunk`]
│   │   ├── construct.rs               # Creating an [`Appender`] and choosing the columns it appends to
│   │   ├── lifecycle.rs               # Flushing, closing and (with `duckdb-1-5`) clearing an [`Appender`]
│   │   ├── rows.rs                    # Row-at-a-time appends: `row`, `end_row` and the non-numeric `append_*` methods
│   │   └── scalars.rs                 # The fixed-width numeric `append_*` methods, `append_bool` through `append_u128`
│   ├── arrow/
│   │   ├── array.rs                   # `ArrowArray` — an owned Arrow C Data Interface array
│   │   ├── convert.rs                 # The four conversions between `DuckDB` data chunks and the Arrow C Data Interface
│   │   ├── converted.rs               # `ArrowConvertedSchema` — an Arrow schema translated into `DuckDB`'s own type descriptors
│   │   ├── export_check.rs            # Values `DuckDB` would export as different values, found before the export
│   │   ├── import_check.rs            # Structural checks on an imported array, and flattening what the readers cannot index
│   │   ├── import_layout/
│   │   │   └── tests.rs               # Unit tests
│   │   ├── import_layout.rs           # Arrow layouts `DuckDB` imports wrongly, found by walking the array with its schema
│   │   ├── options.rs                 # `ArrowOptions` — the Arrow production settings of a connection or a result
│   │   ├── schema.rs                  # `ArrowSchema` — an owned Arrow C Data Interface schema
│   │   └── tests.rs                   # Unit tests
│   ├── bin/
│   │   └── append_metadata/
│   │       ├── cli.rs                 # Command-line parsing and validation (std-only, no clap)
│   │       ├── footer.rs              # The 512-byte `DuckDB` extension footer, and the optional 22-byte WebAssembly custom-section header that precedes it
│   │       ├── main.rs                # Append a DuckDB extension metadata block to a compiled .so / .dylib / .dll file,
│   │       └── tests.rs               # Unit tests
│   ├── callback/
│   │   └── payload.rs                 # Disposing of a caught panic payload without re-entering the unwinder
│   ├── cast/
│   │   ├── builder.rs                 # Builder for registering custom `DuckDB` cast functions
│   │   └── mod.rs                     # Builder for registering `DuckDB` custom cast functions
│   ├── copy_function/
│   │   ├── info.rs                    # Callback info wrappers for copy function callbacks
│   │   └── mod.rs                     # Copy function registration (`DuckDB` 1.5.0+)
│   ├── datetime/
│   │   ├── checks.rs                  # Pure-Rust mirrors of the checks `DuckDB` makes before it throws
│   │   ├── mod.rs                     # Calendar conversions for `DuckDB`'s temporal types
│   │   └── tests.rs                   # Unit tests
│   ├── query/
│   │   ├── bind.rs                    # Binding a [`PreparedStatement`]'s parameters: the typed `bind_*` methods and `bind_value`
│   │   ├── chunk.rs                   # [`OwnedDataChunk`]: a `duckdb_data_chunk` destroyed on drop
│   │   ├── connection.rs              # [`OwnedConnection`] and its cross-thread [`InterruptHandle`]
│   │   ├── cstr.rs                    # The two C-string conversions the `query` module runs everything through
│   │   ├── live_tests.rs              # Tests that need a live `DuckDB`
│   │   ├── prepared.rs                # Inspecting and executing a [`PreparedStatement`]; `bind.rs` binds its parameters
│   │   └── result.rs                  # Reading a [`QueryResult`]: its columns, its chunks and what kind of outcome it is
│   ├── replacement_scan/
│   │   └── mod.rs                     # Builder for registering `DuckDB` replacement scans
│   ├── scaffold/
│   │   ├── escape.rs                  # Quoting configured free text for YAML and Rust doc comments
│   │   ├── mod.rs                     # Project scaffolding for `DuckDB` Rust extensions
│   │   ├── templates.rs               # Template generators for scaffold file content
│   │   ├── tests.rs                   # Unit tests
│   │   ├── tests_escaping.rs          # Free text reaches the generated files intact
│   │   └── tests_generated.rs         # Unit tests
│   ├── scalar/
│   │   ├── info.rs                    # Ergonomic wrapper around `duckdb_function_info` for scalar function callbacks
│   │   ├── mod.rs                     # Builder for registering `DuckDB` scalar functions
│   │   ├── state.rs                   # Typed bind data and per-thread local state for scalar functions (`DuckDB` 1.5.0+)
│   │   ├── typed.rs                   # Scalar functions written as ordinary Rust closures
│   │   ├── typed_builder.rs           # The builder the closure-based scalar constructors return, and the one `extern "C"` trampoline they all share
│   │   └── builder/
│   │       ├── collision.rs           # Refusing a scalar signature that would replace an existing one or make calls ambiguous
│   │       ├── mod.rs                 # Builder for registering `DuckDB` scalar functions
│   │       ├── overload.rs            # One overload within a [`ScalarFunctionSetBuilder`]
│   │       ├── set.rs                 # Builder for registering a `DuckDB` scalar function set (multiple overloads)
│   │       ├── signature.rs           # Detecting overloads that accept the same call
│   │       ├── single.rs              # Builder for registering a single-signature `DuckDB` scalar function
│   │       └── tests.rs               # Unit tests
│   ├── table/
│   │   ├── bind_data.rs               # Type-safe bind data management for table functions
│   │   ├── builder.rs                 # Builder for registering `DuckDB` table functions
│   │   ├── cstr.rs                    # Panic-free `&str` → `CString` conversion for the callback info wrappers
│   │   ├── info.rs                    # Ergonomic wrappers around `DuckDB` callback info handles
│   │   ├── init_data.rs               # Type-safe init data management for table functions
│   │   ├── mod.rs                     # Builder for registering `DuckDB` table functions
│   │   ├── type_check.rs              # Detects logical types that `DuckDB` refuses without saying so
│   │   ├── typed.rs                   # Closure-based table functions with typed scan state
│   │   └── typed/
│   │       └── trampolines.rs         # `extern "C"` trampolines behind [`TypedTableFunctionBuilder`][super::TypedTableFunctionBuilder]
│   ├── testing/
│   │   ├── bundled_api_init.cpp       # Compiled only when the `bundled-test` Cargo feature is active
│   │   ├── harness.rs                 # [`AggregateTestHarness`] — test aggregate logic without `DuckDB`
│   │   ├── in_memory_db.rs            # In-memory `DuckDB` helper for integration tests
│   │   ├── mock_registrar.rs          # [`MockRegistrar`] — a [`Registrar`] implementation for testing
│   │   ├── mock_vector.rs             # In-memory mock types for `DuckDB` vectors
│   │   ├── mod.rs                     # Test utilities for `DuckDB` extension development
│   │   └── mock_vector/
│   │       ├── reader.rs              # `MockVectorReader` — an in-memory mock input vector
│   │       ├── tests.rs               # Unit tests
│   │       └── writer.rs              # `MockVectorWriter` — an in-memory mock output vector
│   ├── types/
│   │   ├── logical_type.rs            # RAII wrapper for `duckdb_logical_type`
│   │   ├── mod.rs                     # `DuckDB` type system wrappers
│   │   ├── null_handling.rs           # NULL propagation behaviour for `DuckDB` functions
│   │   ├── type_id.rs                 # Ergonomic enum of all `DuckDB` column types
│   │   └── logical_type/
│   │       └── construct.rs           # Every `LogicalType` constructor, as `try_*` plus a panicking wrapper
│   ├── validate/
│   │   ├── extension_name.rs          # Extension name validation per `DuckDB` community extension rules
│   │   ├── function_name.rs           # SQL function name validation for `DuckDB` extensions
│   │   ├── mod.rs                     # Validation utilities for `DuckDB` community extension compliance
│   │   ├── platform.rs                # `DuckDB` build platform validation
│   │   ├── release_profile.rs         # Release profile validation for `DuckDB` loadable extensions
│   │   ├── semver.rs                  # Semantic versioning validation for `DuckDB` community extensions
│   │   ├── spdx.rs                    # SPDX license identifier validation for `DuckDB` community extensions
│   │   ├── spdx_exceptions.rs         # The SPDX license-exception identifiers accepted after `WITH`
│   │   └── description_yml/
│   │       ├── mod.rs                 # Validation of `DuckDB` community extension `description.yml` files
│   │       ├── model.rs               # A validated representation of a `DuckDB` community extension `description.yml`
│   │       ├── parser.rs              # Parses and validates a `description.yml` string
│   │       ├── tests.rs               # Unit tests
│   │       ├── tests_corpus.rs        # Unit tests
│   │       ├── tests_yaml.rs          # Unit tests
│   │       ├── validator.rs           # Validates a `description.yml` string and returns `Ok(())` if it passes all checks
│   │       ├── yaml.rs                # A reader for the subset of YAML that `description.yml` files use
│   │       ├── yaml11.rs              # Which plain scalars PyYAML (YAML 1.1) reads as booleans, numbers, dates or null
│   │       └── yaml/
│   │           └── scalar.rs          # Scalar-level pieces of the `description.yml` YAML reader: decoding plain, quoted, block and flow values, and recognising keys and comments
│   ├── value/
│   │   ├── blob.rs                    # `Value::as_blob` — `BLOB` extraction
│   │   ├── checks.rs                  # Pure-Rust preconditions checked before a `Value` call reaches `DuckDB`
│   │   ├── composite.rs               # Composite constructors: `STRUCT`, `LIST`, `ARRAY`, `ENUM`, `MAP`, `UNION`
│   │   ├── defaults.rs                # The defaulting accessors — `Value::as_*_or`
│   │   ├── getters.rs                 # The typed scalar accessors — `Value::as_i64`, `as_timestamp`, `as_decimal`, …
│   │   ├── hugeint.rs                 # Conversions between Rust's 128-bit integers and `DuckDB`'s split-word `HUGEINT` / `UHUGEINT` records
│   │   ├── nested.rs                  # Reading nested values: `LIST` elements, `STRUCT` fields, `MAP` entries
│   │   ├── render_guard.rs            # Refusing to render a value `DuckDB` would throw on (an out-of-range timestamp from SQL)
│   │   ├── scalars.rs                 # The non-temporal scalar constructors
│   │   ├── temporal.rs                # Temporal constructors, validated against `DuckDB`'s ranges
│   │   └── temporal_checks.rs         # Pure-Rust range checks for the temporal types, derived from `DuckDB`'s source
│   └── vector/
│       ├── complex.rs                 # Complex type vector operations: STRUCT fields, LIST elements, MAP entries
│       ├── list_builder.rs            # Safe construction of `LIST` and `MAP` output vectors
│       ├── mod.rs                     # Safe helpers for reading from and writing to `DuckDB` data vectors
│       ├── nested_null.rs             # The child validity masks a NULL in a nested vector must also clear
│       ├── ops.rs                     # Whole-vector operations (`DuckDB` 1.5.0+)
│       ├── reader.rs                  # Safe typed reading from `DuckDB` data vectors
│       ├── string.rs                  # `DuckDB` `VARCHAR` and `BLOB` (`duckdb_string_t`) reading utilities
│       ├── struct_reader.rs           # Batched, typed reader for STRUCT input vectors
│       ├── struct_writer.rs           # Batched, typed writer for STRUCT output vectors
│       ├── uuid.rs                    # Converting between a `UUID`'s textual bits and `DuckDB`'s vector storage
│       ├── validity.rs                # Validity bitmap helpers for `DuckDB` NULL tracking
│       └── writer.rs                  # Safe typed writing to `DuckDB` result vectors
├── tests/
│   ├── aggregate_leaks.rs             # Aggregate states `DuckDB` never destroys leak no Rust heap
│   ├── append_metadata_cli.rs         # The `append_metadata` binary run end to end: exit status, output and the file it writes
│   ├── ffi_roundtrip.rs               # End-to-end FFI round-trips against a real `DuckDB`
│   ├── file_handle_close.rs           # `FileHandle`'s `Drop` when the close fails (stubbed C API)
│   ├── handle_leaks.rs                # Every RAII handle frees what `DuckDB` allocated for it (glibc)
│   ├── integration_test.rs            # Integration tests for `quack-rs`
│   ├── secret_zeroize.rs              # `SecretEntry` never frees a buffer that still holds a secret
│   └── ffi_roundtrip/
│       ├── agg_states.rs              # Every aggregate state is dropped, including the ones `DuckDB` moves
│       ├── agg_window.rs              # Aggregates in the running-window and sorted-aggregate paths
│       ├── appender_api.rs            # `Appender` methods a no-op replacement survived: schemas, column types, `clear_columns`, defaults
│       ├── appender_rows.rs           # What happens to buffered rows when an append fails mid-row
│       ├── arrow_export.rs            # `arrow::data_chunk_to_arrow` refuses values `DuckDB` would export wrongly
│       ├── arrow_import.rs            # `arrow::data_chunk_from_arrow` checks against a live `DuckDB`
│       ├── arrow_layout.rs            # Valid Arrow layouts `DuckDB` imports wrongly, refused, and their correct neighbours
│       ├── bind_expressions.rs        # What a bind callback learns about its arguments from `Expression`
│       ├── chunk_writer.rs            # `ChunkWriter` against a chunk `DuckDB` allocated
│       ├── collision.rs               # The scalar signature-collision check, held to `DuckDB`'s own binder
│       ├── copy_from_columns.rs       # A typed `COPY … FROM` reader that declares a column is refused
│       ├── file_errors.rs             # `FileHandle` reports write, sync and seek failures, and refuses a seek past `i64::MAX`
│       ├── handles_api.rs             # `StructWriter` child handles and `InMemoryDb::execute`'s row count
│       ├── lifecycle.rs               # Aggregate NULL rows, name collisions, overload builders, bind-data sharing
│       ├── list_limits.rs             # `ListBuilder` stops at `DuckDB`'s byte ceiling, not an element count
│       ├── mock_parity.rs             # `MockRegistrar` refuses a bad type with the live registration's message
│       ├── nested_reserve.rs          # which nested buffers a `LIST` reserve moves (the writer contracts)
│       ├── nested_validity.rs         # `VectorWriter::set_valid` on nested rows, against a live `DuckDB`
│       ├── panic_guards.rs            # The panic-guard macros and `set_error` methods, against a live `DuckDB`
│       ├── query_docs.rs              # Pins the documented behaviour of `query`, `PreparedStatement`, `DbConfig`
│       ├── query_stream.rs            # A streaming result that stops early must not look like a finished one
│       ├── scalar_agg.rs              # Scalar and aggregate builder regressions
│       ├── table_cast.rs              # Table function, cast, replacement scan, SQL macro and COPY regressions
│       ├── table_description.rs       # `TableDescription` accessors on an index `DuckDB` cannot hold
│       ├── temporal_binds.rs          # Temporal and over-4-GiB values refused by `PreparedStatement` binds and the `Appender`
│       ├── tooling.rs                 # Checks of quack-rs's tooling tables against the linked `DuckDB`
│       ├── value_getters.rs           # Each typed `Value` getter at its own type; out-of-range TIMETZ / TIME_NS are not cast
│       ├── value_nested.rs            # Nested `Value` construction and inspection against a live `DuckDB`
│       ├── value_query.rs             # `Value` getters, DECIMAL binding, `Expression::fold`
│       ├── value_render.rs            # `Value` rendering of values SQL builds and `DuckDB` cannot render
│       ├── value_temporal.rs          # Every `Value` getter against every temporal source type, at every edge
│       └── vector_dt.rs               # NULLs in nested output vectors; selection vectors
├── benches/
│   └── interval_bench.rs          # Criterion benchmarks
├── examples/
│   └── hello-ext/                 # Reference extension: aggregates, scalars, table functions, casts
├── book/                          # mdBook documentation source
│   ├── src/                       # Markdown pages (this site)
│   └── theme/custom.css
├── .github/workflows/ci.yml       # CI pipeline
├── .github/workflows/docs.yml     # GitHub Pages deployment
├── CONTRIBUTING.md
├── LESSONS.md                     # The DuckDB Rust FFI pitfalls (L1–L19, P1–P12)
├── CHANGELOG.md
└── README.md

Releasing

quack-rs uses libduckdb-sys = ">=1.4.4, <2" — a bounded range covering DuckDB 1.4.x and 1.5.x, every one of which loads C API v1.2.0 extensions. The <2 upper bound prevents silent adoption of a future major release that may change the C API. Before broadening the range to a new major band:

  1. Read the DuckDB changelog for C API changes
  2. Check the new C API version string (used in duckdb_rs_extension_api_init)
  3. Update DUCKDB_API_VERSION in src/lib.rs if the C API version changed
  4. Audit all callback signatures against the new libduckdb-sys bindings
  5. Update the libduckdb-sys and duckdb version requirements in Cargo.toml

Versions follow Semantic Versioning as Cargo applies it. While the crate is pre-1.0, a breaking change to the public API bumps the minor version (0.17.x → 0.18.0) and is marked Breaking: in CHANGELOG.md; see the semantic versioning policy in RELEASING.md.


Reporting issues

Use GitHub Issues. For security vulnerabilities, see SECURITY.md for responsible disclosure policy.


License

quack-rs is licensed under the MIT License. Contributions are accepted under the same license. By submitting a pull request, you agree to license your contribution under MIT.