Your First Extension

This page builds a DuckDB extension in Rust step by step, using hello-ext, the example extension bundled with quack-rs. The table lists four of the functions it registers, one of each major kind (the full list is in its README); this page walks through the aggregate and the scalar. The table function and the cast are covered in Table Functions and Cast Functions.

SQLKindSignature
word_count(text)AggregateVARCHAR → BIGINT
first_word(text)ScalarVARCHAR → VARCHAR
generate_series_ext(n)TableBIGINT → TABLE(value BIGINT)
CAST(VARCHAR AS INTEGER)CastVARCHAR → INTEGER

Full source: examples/hello-ext/src/lib.rs


Build and try it

cargo build --release --manifest-path examples/hello-ext/Cargo.toml

# DuckDB only loads files ending in `.duckdb_extension` that carry its
# 512-byte metadata footer; a bare `.so` is refused. Append the footer:
cargo run --bin append_metadata -- \
    examples/hello-ext/target/release/libhello_ext.so \
    hello_ext.duckdb_extension \
    --abi-type C_STRUCT --extension-version v0.1.0 \
    --duckdb-version v1.2.0 --platform linux_amd64

# An unsigned local build needs -unsigned (it must be a startup flag).
duckdb -unsigned

Then in the DuckDB CLI:

LOAD './hello_ext.duckdb_extension';

-- Aggregate: total words across all rows
SELECT word_count(sentence) FROM (
    VALUES ('hello world'), ('one two three'), (NULL)
) t(sentence);
-- → 5  (2 + 3; NULL contributes 0)

-- Scalar: first word of each row
SELECT first_word(sentence) FROM (
    VALUES ('hello world'), ('  padded  '), (''), (NULL)
) t(sentence);
-- → 'hello', 'padded', '', NULL

Overview

An extension has four parts:

  1. State struct — holds data accumulated during aggregation (aggregate only)
  2. Callbacks — update, combine, finalize (aggregate; ffi_state::<T>() supplies state_size, state_init and state_destroy) or a single function callback (scalar)
  3. Registration — wire callbacks to DuckDB via AggregateFunctionBuilder / ScalarFunctionBuilder
  4. Entry point — DuckDB's initialization hook, generated by entry_point!

Part 1 — Aggregate function: word_count

An aggregate function accumulates state across many rows and emits one result per group.

1a. The state struct

#![allow(unused)]
fn main() {
use quack_rs::prelude::*;
#[derive(Default, Debug)]
struct WordCountState {
    count: i64,
}

impl AggregateState for WordCountState {}
}

AggregateState is a marker trait with no methods; its supertraits require the state to be Default + Send + Sync + 'static. FfiState<WordCountState> stores it in the bytes DuckDB allocates for each group (boxing it only when it is large or over-aligned) and manages its lifecycle (size, init, destroy).

1b. state_size, state_init and state_destroy

These three callbacks are always identical boilerplate, so you do not write them: .ffi_state::<WordCountState>() on the builder (see Part 3) installs FfiState::<WordCountState>::size_callback, init_callback and destroy_callback together. Installing all three for one type in a single call means DuckDB never allocates space sized for one state and has it initialised as another.

size_callback returns FfiState::<WordCountState>::size() — one usize tag word followed by the state. A T aligned no more strictly than usize and at most 256 bytes (like WordCountState) is stored inline in the bytes DuckDB allocates per group; a larger or more strictly aligned T is boxed, and the slot holds the Box pointer. init_callback writes WordCountState::default() into the slot (or boxes it), then sets the tag — a value salted per state type, so a destructor only drops states that were initialised for its own T.

destroy_callback drops the T in each state (and its box, if the state is too large to be stored inline), clearing the state's tag first so a second call is a no-op.

1c. update — accumulate one batch

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 }
unsafe extern "C" fn wc_update(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    states: *mut duckdb_aggregate_state,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let row_count = reader.row_count();

    for row in 0..row_count {
        if !unsafe { reader.is_valid(row) } {
            continue; // NULL input → skip (contributes 0 words)
        }
        let s = unsafe { reader.read_str(row) };
        let words = count_words(s);

        let state_ptr = unsafe { *states.add(row) };
        if let Some(st) = unsafe { FfiState::<WordCountState>::with_state_mut(state_ptr) } {
            st.count += words;
        }
    }
}
}

Key points:

  • Check is_valid(row) before reading — never dereference an invalid (NULL) row
  • VectorReader::new(chunk, col) gives column col from the chunk
  • count_words is pure Rust — no unsafe, easy to unit-test separately

1d. combine — merge parallel results

Pitfall L1: DuckDB creates fresh target states before calling combine, set up by state_init (here WordCountState::default()), not copies of the source. You must copy all fields — not just the result field. In an aggregate with config fields (e.g., a histogram with a bin_width) you must also copy those, or results will be silently corrupted.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
unsafe extern "C" fn wc_combine(
    _info: duckdb_function_info,
    source: *mut duckdb_aggregate_state,
    target: *mut duckdb_aggregate_state,
    count: idx_t,
) {
    for i in 0..count as usize {
        let src_ptr = unsafe { *source.add(i) };
        let tgt_ptr = unsafe { *target.add(i) };
        let src = unsafe { FfiState::<WordCountState>::with_state(src_ptr) };
        let tgt = unsafe { FfiState::<WordCountState>::with_state_mut(tgt_ptr) };
        if let (Some(s), Some(t)) = (src, tgt) {
            t.count += s.count;
            // If you add fields to WordCountState, combine them here too.
        }
    }
}
}

1e. finalize — write output

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
unsafe extern "C" fn wc_finalize(
    _info: duckdb_function_info,
    source: *mut duckdb_aggregate_state,
    result: duckdb_vector,
    count: idx_t,
    offset: idx_t,
) {
    let mut writer = unsafe { VectorWriter::new(result) };

    for i in 0..count as usize {
        let state_ptr = unsafe { *source.add(i) };
        match unsafe { FfiState::<WordCountState>::with_state(state_ptr) } {
            Some(st) => unsafe { writer.write_i64(offset as usize + i, st.count) },
            None     => unsafe { writer.set_null(offset as usize + i) },
        }
    }
}
}

offset is DuckDB's output row offset — always use offset as usize + i, not just i.


Part 2 — Scalar function: first_word

A scalar function processes one data chunk and returns one output value per row. The callback receives the full chunk and an output vector (not per-row state pointers).

Key rule: always propagate NULL

If the input row is NULL, write NULL to output — never read from an invalid row.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") }
unsafe extern "C" fn first_word_scalar(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };
    let row_count = reader.row_count();

    for row in 0..row_count {
        if !unsafe { reader.is_valid(row) } {
            unsafe { writer.set_null(row) }; // NULL in → NULL out
            continue;
        }
        let s = unsafe { reader.read_str(row) };
        unsafe { writer.write_varchar(row, first_word(s)) };
    }
}
}

The pure logic:

#![allow(unused)]
fn main() {
pub fn first_word(s: &str) -> &str {
    s.split_whitespace().next().unwrap_or("")
}
}

Note: set_null internally calls duckdb_vector_ensure_validity_writable before writing the null flag — this is required by DuckDB and handled for you by VectorWriter.


Part 3 — Registration

The snippets below register through the raw builders and entry_point!, which take a duckdb_connection. hello-ext itself uses the equivalent entry_point_v2!, whose closure receives a &Connection and registers each builder through the Registrar trait (con.register_aggregate(...), con.register_scalar(...)); see The Entry Point.

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 }
fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") }
unsafe extern "C" fn wc_update(_info: duckdb_function_info, input: duckdb_data_chunk,
    states: *mut duckdb_aggregate_state) {
    let reader = unsafe { VectorReader::new(input, 0) };
    for row in 0..reader.row_count() {
        if !unsafe { reader.is_valid(row) } { continue; }
        let words = count_words(unsafe { reader.read_str(row) });
        if let Some(st) = unsafe { FfiState::<WordCountState>::with_state_mut(*states.add(row)) } {
            st.count += words;
        }
    }
}
unsafe extern "C" fn wc_combine(_info: duckdb_function_info, source: *mut duckdb_aggregate_state,
    target: *mut duckdb_aggregate_state, count: idx_t) {
    for i in 0..count as usize {
        let src = unsafe { FfiState::<WordCountState>::with_state(*source.add(i)) };
        let tgt = unsafe { FfiState::<WordCountState>::with_state_mut(*target.add(i)) };
        if let (Some(s), Some(t)) = (src, tgt) { t.count += s.count; }
    }
}
unsafe extern "C" fn wc_finalize(_info: duckdb_function_info, source: *mut duckdb_aggregate_state,
    result: duckdb_vector, count: idx_t, offset: idx_t) {
    let mut writer = unsafe { VectorWriter::new(result) };
    for i in 0..count as usize {
        match unsafe { FfiState::<WordCountState>::with_state(*source.add(i)) } {
            Some(st) => unsafe { writer.write_i64(offset as usize + i, st.count) },
            None => unsafe { writer.set_null(offset as usize + i) },
        }
    }
}
unsafe extern "C" fn first_word_scalar(_info: duckdb_function_info, input: duckdb_data_chunk,
    output: duckdb_vector) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };
    for row in 0..reader.row_count() {
        if !unsafe { reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; }
        unsafe { writer.write_varchar(row, first_word(reader.read_str(row))) };
    }
}
fn live_connection() -> libduckdb_sys::duckdb_connection {
    std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap());
    let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut());
    unsafe {
        assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess);
        assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess);
    }
    con
}
/// First column of the first row, as BIGINT; `None` for NULL.
fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> {
    let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap();
    let chunk = result.next_chunk().unwrap().unwrap();
    let reader = unsafe { chunk.reader(0) };
    unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) }
}
unsafe fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> {
    unsafe {
        AggregateFunctionBuilder::new("word_count")
            .param(TypeId::Varchar)
            .returns(TypeId::BigInt)
            .ffi_state::<WordCountState>() // state_size + init + destructor
            .update(wc_update)
            .combine(wc_combine)
            .finalize(wc_finalize)
            .register(con)?;

        ScalarFunctionBuilder::new("first_word")
            .param(TypeId::Varchar)
            .returns(TypeId::Varchar)
            .function(first_word_scalar)
            .register(con)?;
    }
    Ok(())
}
let con = live_connection();
unsafe { register(con) }.unwrap();
assert_eq!(query_i64(con, "SELECT word_count(s) FROM (VALUES ('hello world'), (NULL), ('one two three')) t(s)"), Some(5));
assert_eq!(query_i64(con, "SELECT length(first_word('  quack rs'))::BIGINT"), Some(5));
assert_eq!(query_i64(con, "SELECT count(*) FROM (SELECT first_word(NULL) AS w) WHERE w IS NULL"), Some(1));
// Many groups, so DuckDB runs combine on partial states.
assert_eq!(query_i64(con, "SELECT sum(c)::BIGINT FROM (SELECT word_count('a b') AS c FROM range(100000) GROUP BY range % 997)"), Some(200000));
}

Both builders call the DuckDB C API internally. register returns Err if DuckDB reports a failure — this propagates to the entry point and is surfaced to the user.


Part 4 — Entry point

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection,
    duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t};
use quack_rs::prelude::*;
unsafe fn register(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
quack_rs::entry_point!(hello_ext_init_c_api, |con| unsafe { register(con) });
}

This one line expands to the equivalent of:

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_connection, duckdb_extension_access, duckdb_extension_info};
use quack_rs::error::ExtensionError;
unsafe fn register(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) }
#[no_mangle]
pub unsafe extern "C" fn hello_ext_init_c_api(
    info: duckdb_extension_info,
    access: *const duckdb_extension_access,
) -> bool {
    unsafe {
        quack_rs::entry_point::init_extension_with_policy(
            info, access, quack_rs::DUCKDB_API_VERSION,
            quack_rs::abi::AbiPolicy::Strict,
            |con| unsafe { register(con) },
        )
    }
}
}

Pass the full symbol name — hello_ext_init_c_api here. DuckDB looks up this exact symbol when loading the extension. See The Entry Point for the full initialization sequence.


Unit tests (no DuckDB process needed)

Test pure logic directly:

#![allow(unused)]
fn main() {
fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 }
fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") }
#[test]
fn count_words_whitespace_variants() {
    assert_eq!(count_words("  hello  world  "), 2);
    assert_eq!(count_words("\t\nhello\tworld\n"), 2);
    assert_eq!(count_words("   "), 0); // all whitespace → 0
}

#[test]
fn first_word_empty_and_whitespace() {
    assert_eq!(first_word(""), "");
    assert_eq!(first_word("   "), "");
}
}

Test aggregate state with AggregateTestHarness:

#![allow(unused)]
fn main() {
use quack_rs::prelude::*;
use quack_rs::testing::AggregateTestHarness;
#[derive(Default, Debug)] struct WordCountState { count: i64 }
impl AggregateState for WordCountState {}
fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 }
#[test]
fn word_count_null_rows_are_skipped() {
    // DuckDB passes NULL rows to `update`; the callback's `is_valid` check
    // skips them, so they never reach the state.
    let mut h = AggregateTestHarness::<WordCountState>::new();
    h.update(|s| s.count += count_words("hello"));
    // NULL row omitted — models the callback's `is_valid` skip
    h.update(|s| s.count += count_words("world"));
    assert_eq!(h.finalize().count, 2);
}

#[test]
fn word_count_combine() {
    let mut h1 = AggregateTestHarness::<WordCountState>::new();
    h1.update(|s| s.count += count_words("hello world")); // 2

    let mut h2 = AggregateTestHarness::<WordCountState>::new();
    h2.update(|s| s.count += count_words("one two three four")); // 4

    h2.combine(&h1, |src, tgt| tgt.count += src.count);
    assert_eq!(h2.finalize().count, 6);
}
}

Run all tests with:

cargo test --manifest-path examples/hello-ext/Cargo.toml

See the Testing Guide for the full test strategy.