Your First Extension
This page builds a DuckDB extension in Rust step by step, using hello-ext, the example
extension bundled with quack-rs. The table lists four of the functions it registers, one of
each major kind (the full list is in its
README); this page walks
through the aggregate and the scalar. The table function and the cast are covered in
Table Functions and
Cast Functions.
| SQL | Kind | Signature |
|---|---|---|
word_count(text) | Aggregate | VARCHAR → BIGINT |
first_word(text) | Scalar | VARCHAR → VARCHAR |
generate_series_ext(n) | Table | BIGINT → TABLE(value BIGINT) |
CAST(VARCHAR AS INTEGER) | Cast | VARCHAR → INTEGER |
Full source: examples/hello-ext/src/lib.rs
Build and try it
cargo build --release --manifest-path examples/hello-ext/Cargo.toml
# DuckDB only loads files ending in `.duckdb_extension` that carry its
# 512-byte metadata footer; a bare `.so` is refused. Append the footer:
cargo run --bin append_metadata -- \
examples/hello-ext/target/release/libhello_ext.so \
hello_ext.duckdb_extension \
--abi-type C_STRUCT --extension-version v0.1.0 \
--duckdb-version v1.2.0 --platform linux_amd64
# An unsigned local build needs -unsigned (it must be a startup flag).
duckdb -unsigned
Then in the DuckDB CLI:
LOAD './hello_ext.duckdb_extension';
-- Aggregate: total words across all rows
SELECT word_count(sentence) FROM (
VALUES ('hello world'), ('one two three'), (NULL)
) t(sentence);
-- → 5 (2 + 3; NULL contributes 0)
-- Scalar: first word of each row
SELECT first_word(sentence) FROM (
VALUES ('hello world'), (' padded '), (''), (NULL)
) t(sentence);
-- → 'hello', 'padded', '', NULL
Overview
An extension has four parts:
- State struct — holds data accumulated during aggregation (aggregate only)
- Callbacks —
update,combine,finalize(aggregate;ffi_state::<T>()suppliesstate_size,state_initandstate_destroy) or a single function callback (scalar) - Registration — wire callbacks to DuckDB via
AggregateFunctionBuilder/ScalarFunctionBuilder - Entry point — DuckDB's initialization hook, generated by
entry_point!
Part 1 — Aggregate function: word_count
An aggregate function accumulates state across many rows and emits one result per group.
1a. The state struct
#![allow(unused)] fn main() { use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64, } impl AggregateState for WordCountState {} }
AggregateState is a marker trait with no methods; its supertraits require the state to be
Default + Send + Sync + 'static.
FfiState<WordCountState> stores it in the bytes DuckDB allocates for each group
(boxing it only when it is large or over-aligned) and manages its lifecycle
(size, init, destroy).
1b. state_size, state_init and state_destroy
These three callbacks are always identical boilerplate, so you do not write them:
.ffi_state::<WordCountState>() on the builder (see Part 3)
installs FfiState::<WordCountState>::size_callback, init_callback and
destroy_callback together. Installing all three for one type in a single call means
DuckDB never allocates space sized for one state and has it initialised as another.
size_callback returns FfiState::<WordCountState>::size() — one usize tag word followed
by the state. A T aligned no more strictly than usize and at most 256 bytes (like
WordCountState) is stored inline in the bytes DuckDB allocates per group; a larger or more
strictly aligned T is boxed, and the slot holds the Box pointer. init_callback writes
WordCountState::default() into the slot (or boxes it), then sets the tag — a value salted per
state type, so a destructor only drops states that were initialised for its own T.
destroy_callback drops the T in each state (and its box, if the state is too
large to be stored inline), clearing the state's tag first so a second call is a
no-op.
1c. update — accumulate one batch
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 } unsafe extern "C" fn wc_update( _info: duckdb_function_info, input: duckdb_data_chunk, states: *mut duckdb_aggregate_state, ) { let reader = unsafe { VectorReader::new(input, 0) }; let row_count = reader.row_count(); for row in 0..row_count { if !unsafe { reader.is_valid(row) } { continue; // NULL input → skip (contributes 0 words) } let s = unsafe { reader.read_str(row) }; let words = count_words(s); let state_ptr = unsafe { *states.add(row) }; if let Some(st) = unsafe { FfiState::<WordCountState>::with_state_mut(state_ptr) } { st.count += words; } } } }
Key points:
- Check
is_valid(row)before reading — never dereference an invalid (NULL) row VectorReader::new(chunk, col)gives columncolfrom the chunkcount_wordsis pure Rust — no unsafe, easy to unit-test separately
1d. combine — merge parallel results
Pitfall L1: DuckDB creates fresh target states before calling
combine, set up bystate_init(hereWordCountState::default()), not copies of the source. You must copy all fields — not just the result field. In an aggregate with config fields (e.g., a histogram with abin_width) you must also copy those, or results will be silently corrupted.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} unsafe extern "C" fn wc_combine( _info: duckdb_function_info, source: *mut duckdb_aggregate_state, target: *mut duckdb_aggregate_state, count: idx_t, ) { for i in 0..count as usize { let src_ptr = unsafe { *source.add(i) }; let tgt_ptr = unsafe { *target.add(i) }; let src = unsafe { FfiState::<WordCountState>::with_state(src_ptr) }; let tgt = unsafe { FfiState::<WordCountState>::with_state_mut(tgt_ptr) }; if let (Some(s), Some(t)) = (src, tgt) { t.count += s.count; // If you add fields to WordCountState, combine them here too. } } } }
1e. finalize — write output
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} unsafe extern "C" fn wc_finalize( _info: duckdb_function_info, source: *mut duckdb_aggregate_state, result: duckdb_vector, count: idx_t, offset: idx_t, ) { let mut writer = unsafe { VectorWriter::new(result) }; for i in 0..count as usize { let state_ptr = unsafe { *source.add(i) }; match unsafe { FfiState::<WordCountState>::with_state(state_ptr) } { Some(st) => unsafe { writer.write_i64(offset as usize + i, st.count) }, None => unsafe { writer.set_null(offset as usize + i) }, } } } }
offset is DuckDB's output row offset — always use offset as usize + i, not just i.
Part 2 — Scalar function: first_word
A scalar function processes one data chunk and returns one output value per row. The callback receives the full chunk and an output vector (not per-row state pointers).
Key rule: always propagate NULL
If the input row is NULL, write NULL to output — never read from an invalid row.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } unsafe extern "C" fn first_word_scalar( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; let row_count = reader.row_count(); for row in 0..row_count { if !unsafe { reader.is_valid(row) } { unsafe { writer.set_null(row) }; // NULL in → NULL out continue; } let s = unsafe { reader.read_str(row) }; unsafe { writer.write_varchar(row, first_word(s)) }; } } }
The pure logic:
#![allow(unused)] fn main() { pub fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } }
Note: set_null internally calls duckdb_vector_ensure_validity_writable before writing
the null flag — this is required by DuckDB and handled for you by VectorWriter.
Part 3 — Registration
The snippets below register through the raw builders and entry_point!, which take a
duckdb_connection. hello-ext itself uses the equivalent entry_point_v2!, whose closure
receives a &Connection and registers each builder through the Registrar trait
(con.register_aggregate(...), con.register_scalar(...)); see
The Entry Point.
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 } fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } unsafe extern "C" fn wc_update(_info: duckdb_function_info, input: duckdb_data_chunk, states: *mut duckdb_aggregate_state) { let reader = unsafe { VectorReader::new(input, 0) }; for row in 0..reader.row_count() { if !unsafe { reader.is_valid(row) } { continue; } let words = count_words(unsafe { reader.read_str(row) }); if let Some(st) = unsafe { FfiState::<WordCountState>::with_state_mut(*states.add(row)) } { st.count += words; } } } unsafe extern "C" fn wc_combine(_info: duckdb_function_info, source: *mut duckdb_aggregate_state, target: *mut duckdb_aggregate_state, count: idx_t) { for i in 0..count as usize { let src = unsafe { FfiState::<WordCountState>::with_state(*source.add(i)) }; let tgt = unsafe { FfiState::<WordCountState>::with_state_mut(*target.add(i)) }; if let (Some(s), Some(t)) = (src, tgt) { t.count += s.count; } } } unsafe extern "C" fn wc_finalize(_info: duckdb_function_info, source: *mut duckdb_aggregate_state, result: duckdb_vector, count: idx_t, offset: idx_t) { let mut writer = unsafe { VectorWriter::new(result) }; for i in 0..count as usize { match unsafe { FfiState::<WordCountState>::with_state(*source.add(i)) } { Some(st) => unsafe { writer.write_i64(offset as usize + i, st.count) }, None => unsafe { writer.set_null(offset as usize + i) }, } } } unsafe extern "C" fn first_word_scalar(_info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector) { let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..reader.row_count() { if !unsafe { reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } unsafe { writer.write_varchar(row, first_word(reader.read_str(row))) }; } } fn live_connection() -> libduckdb_sys::duckdb_connection { std::mem::forget(quack_rs::testing::InMemoryDb::open().unwrap()); let (mut db, mut con) = (std::ptr::null_mut(), std::ptr::null_mut()); unsafe { assert_eq!(libduckdb_sys::duckdb_open(std::ptr::null(), &mut db), libduckdb_sys::DuckDBSuccess); assert_eq!(libduckdb_sys::duckdb_connect(db, &mut con), libduckdb_sys::DuckDBSuccess); } con } /// First column of the first row, as BIGINT; `None` for NULL. fn query_i64(con: libduckdb_sys::duckdb_connection, sql: &str) -> Option<i64> { let mut result = unsafe { quack_rs::query::query(con, sql) }.unwrap(); let chunk = result.next_chunk().unwrap().unwrap(); let reader = unsafe { chunk.reader(0) }; unsafe { reader.is_valid(0).then(|| reader.read_i64(0)) } } unsafe fn register(con: libduckdb_sys::duckdb_connection) -> Result<(), ExtensionError> { unsafe { AggregateFunctionBuilder::new("word_count") .param(TypeId::Varchar) .returns(TypeId::BigInt) .ffi_state::<WordCountState>() // state_size + init + destructor .update(wc_update) .combine(wc_combine) .finalize(wc_finalize) .register(con)?; ScalarFunctionBuilder::new("first_word") .param(TypeId::Varchar) .returns(TypeId::Varchar) .function(first_word_scalar) .register(con)?; } Ok(()) } let con = live_connection(); unsafe { register(con) }.unwrap(); assert_eq!(query_i64(con, "SELECT word_count(s) FROM (VALUES ('hello world'), (NULL), ('one two three')) t(s)"), Some(5)); assert_eq!(query_i64(con, "SELECT length(first_word(' quack rs'))::BIGINT"), Some(5)); assert_eq!(query_i64(con, "SELECT count(*) FROM (SELECT first_word(NULL) AS w) WHERE w IS NULL"), Some(1)); // Many groups, so DuckDB runs combine on partial states. assert_eq!(query_i64(con, "SELECT sum(c)::BIGINT FROM (SELECT word_count('a b') AS c FROM range(100000) GROUP BY range % 997)"), Some(200000)); }
Both builders call the DuckDB C API internally. register returns Err if DuckDB reports
a failure — this propagates to the entry point and is surfaced to the user.
Part 4 — Entry point
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_aggregate_state, duckdb_bind_info, duckdb_connection, duckdb_data_chunk, duckdb_function_info, duckdb_init_info, duckdb_vector, idx_t}; use quack_rs::prelude::*; unsafe fn register(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } quack_rs::entry_point!(hello_ext_init_c_api, |con| unsafe { register(con) }); }
This one line expands to the equivalent of:
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_connection, duckdb_extension_access, duckdb_extension_info}; use quack_rs::error::ExtensionError; unsafe fn register(_: duckdb_connection) -> Result<(), ExtensionError> { Ok(()) } #[no_mangle] pub unsafe extern "C" fn hello_ext_init_c_api( info: duckdb_extension_info, access: *const duckdb_extension_access, ) -> bool { unsafe { quack_rs::entry_point::init_extension_with_policy( info, access, quack_rs::DUCKDB_API_VERSION, quack_rs::abi::AbiPolicy::Strict, |con| unsafe { register(con) }, ) } } }
Pass the full symbol name — hello_ext_init_c_api here. DuckDB looks up this exact
symbol when loading the extension. See The Entry Point for
the full initialization sequence.
Unit tests (no DuckDB process needed)
Test pure logic directly:
#![allow(unused)] fn main() { fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 } fn first_word(s: &str) -> &str { s.split_whitespace().next().unwrap_or("") } #[test] fn count_words_whitespace_variants() { assert_eq!(count_words(" hello world "), 2); assert_eq!(count_words("\t\nhello\tworld\n"), 2); assert_eq!(count_words(" "), 0); // all whitespace → 0 } #[test] fn first_word_empty_and_whitespace() { assert_eq!(first_word(""), ""); assert_eq!(first_word(" "), ""); } }
Test aggregate state with AggregateTestHarness:
#![allow(unused)] fn main() { use quack_rs::prelude::*; use quack_rs::testing::AggregateTestHarness; #[derive(Default, Debug)] struct WordCountState { count: i64 } impl AggregateState for WordCountState {} fn count_words(s: &str) -> i64 { s.split_whitespace().count() as i64 } #[test] fn word_count_null_rows_are_skipped() { // DuckDB passes NULL rows to `update`; the callback's `is_valid` check // skips them, so they never reach the state. let mut h = AggregateTestHarness::<WordCountState>::new(); h.update(|s| s.count += count_words("hello")); // NULL row omitted — models the callback's `is_valid` skip h.update(|s| s.count += count_words("world")); assert_eq!(h.finalize().count, 2); } #[test] fn word_count_combine() { let mut h1 = AggregateTestHarness::<WordCountState>::new(); h1.update(|s| s.count += count_words("hello world")); // 2 let mut h2 = AggregateTestHarness::<WordCountState>::new(); h2.update(|s| s.count += count_words("one two three four")); // 4 h2.combine(&h1, |src, tgt| tgt.count += src.count); assert_eq!(h2.finalize().count, 6); } }
Run all tests with:
cargo test --manifest-path examples/hello-ext/Cargo.toml
See the Testing Guide for the full test strategy.