NULL Handling & Strings

This page covers checking for NULL before reading a DuckDB vector, writing NULL output, and reading and writing VARCHAR and BLOB values. The two topics belong together: reading a string from a NULL row is undefined behaviour, not just a wrong value.


NULL checks

Every row in a DuckDB vector may be NULL. Always check validity before reading:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, writer: &mut VectorWriter) {
for row in 0..reader.row_count() {
    if unsafe { !reader.is_valid(row) } {
        // Propagate NULL to output
        unsafe { writer.set_null(row) };
        continue;
    }
    // Safe to read
    let value = unsafe { reader.read_str(row) };
}
}
}

Reading a fixed-width value from a NULL row returns garbage. The vector's data buffer is not zeroed at NULL positions, and no error is raised: you get whatever bytes the buffer holds at that position.

For VARCHAR and BLOB it is worse than garbage. A NULL row's 16-byte entry is left as it was, and DuckDB reuses vector buffers between chunks, so the entry can still be the pointer-format record of a string from an earlier chunk whose memory may since have been freed or reused. read_str / read_blob on such a row follow that pointer: undefined behaviour, not merely a wrong answer. Their # Safety sections require the row to be valid.

Writing NULL

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(writer: &mut VectorWriter, row: usize) {
unsafe { writer.set_null(row) };
}
}

Pitfall L4: VectorWriter::set_null calls duckdb_vector_ensure_validity_writable before accessing the validity bitmap. Without it, a vector that has no mask yet makes duckdb_vector_get_validity return NULL, and the NULL you write is silently dropped. Never write NULL manually; always use set_null — which, for a STRUCT or ARRAY output, also nulls the fields / elements of that row the way DuckDB expects. See Pitfall L4.

Clearing NULL (v0.11.0+)

To mark a row as valid after a previous set_null:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(writer: &mut VectorWriter, row: usize) {
unsafe { writer.set_valid(row) };
}
}

VARCHAR reading

Read VARCHAR columns with VectorReader::read_str:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, row: usize) {
let s: &str = unsafe { reader.read_str(row) };
}
}

The returned &str borrows from the DuckDB vector — it must not outlive the callback. Do not store it in a struct; clone it to a String if you need to keep it.

The duckdb_string_t format

Pitfall P7: the Rust bindings do not document the layout of duckdb_string_t. quack-rs decodes it for you; the details below are for reference. See Pitfall P7.

DuckDB stores VARCHAR values in a 16-byte duckdb_string_t struct with two representations, selected at runtime based on string length:

FormatConditionLayout
Inlinelength ≤ 12[len: u32][data: [u8; 12]]
Pointerlength > 12[len: u32][prefix: [u8; 4]][ptr: *const u8]

On a 32-bit target (DuckDB-WASM) the pointer is 4 bytes and the last 4 bytes are unused. The length and the pointer are in the target's byte order.

VectorReader::read_str and the underlying read_duck_string function handle both formats, so you do not need to inspect the raw struct. A value that is not valid UTF-8 is returned as ""; use read_blob to get its bytes.

Empty strings vs NULL

An empty string ("") and NULL are distinct values:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, row: usize) {
// NULL: is_valid returns false
// Empty string: is_valid returns true, read_str returns ""
if unsafe { !reader.is_valid(row) } {
    // This is NULL
} else {
    let s = unsafe { reader.read_str(row) };
    if s.is_empty() {
        // This is an empty string, not NULL
    }
}
}
}

Writing VARCHAR

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(writer: &mut VectorWriter, row: usize, my_str: &str) {
unsafe { writer.write_varchar(row, my_str) };  // &str
}
}

write_varchar copies the string bytes into DuckDB's managed storage, so the &str need not outlive the call. write_blob does the same for &[u8].

Both panic for a value longer than MAX_STRING_LEN (u32::MAX bytes, the most a duckdb_string_t length field can hold) instead of storing a truncated value; inside scalar_callback! and the typed scalar constructors the panic becomes a SQL error. try_write_varchar and try_write_blob return Result<(), ExtensionError> instead and write nothing on error:

#![allow(unused)]
fn main() {
use quack_rs::vector::VectorWriter;
use quack_rs::error::ExtensionError;
fn demo(writer: &mut VectorWriter, row: usize, my_str: &str) -> Result<(), ExtensionError> {
unsafe { writer.try_write_varchar(row, my_str)? };
Ok(())
}
}

Reading BLOB values

BLOB uses the same inline/pointer layout as VARCHAR, but may contain any bytes. Use read_blob so the data is not interpreted as UTF-8:

#![allow(unused)]
fn main() {
use quack_rs::vector::{VectorReader, VectorWriter};
fn demo(reader: &VectorReader, row: usize) {
let bytes: &[u8] = unsafe { reader.read_blob(row) };
}
}

Like read_str, the returned slice borrows from the DuckDB vector and must not outlive the callback.


Complete NULL-safe VARCHAR pattern

#![allow(unused)]
fn main() {
use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector};
use quack_rs::vector::{VectorReader, VectorWriter};
unsafe extern "C" fn my_scalar(
    _info: duckdb_function_info,
    input: duckdb_data_chunk,
    output: duckdb_vector,
) {
    let reader = unsafe { VectorReader::new(input, 0) };
    let mut writer = unsafe { VectorWriter::new(output) };

    for row in 0..reader.row_count() {
        if unsafe { !reader.is_valid(row) } {
            unsafe { writer.set_null(row) };
            continue;
        }
        let s = unsafe { reader.read_str(row) };
        let upper = s.to_uppercase();
        unsafe { writer.write_varchar(row, &upper) };
    }
}
}

DuckStringView

For advanced use cases where you need access to the raw string bytes or the inline/pointer distinction, quack_rs::vector::string::DuckStringView is available:

#![allow(unused)]
fn main() {
fn demo(data: *const u8, idx: usize) {
use quack_rs::vector::string::{DuckStringView, DUCK_STRING_SIZE};

// From raw 16-byte data (inside a vector callback). `from_raw` is unsafe because it
// follows the pointer of a string longer than 12 bytes; untrusted bytes go through
// `DuckStringView::inline_from_bytes`, which refuses that format instead.
let raw: &[u8; 16] = unsafe { &*data.add(idx * DUCK_STRING_SIZE).cast() };
let view = unsafe { DuckStringView::from_raw(raw) };

println!("length: {}", view.len());
println!("is_empty: {}", view.is_empty());
if let Some(s) = view.as_str() {
    println!("content: {s}");
}
}
}

In practice, prefer reader.read_str(row). DuckStringView is needed only when you have a raw data pointer rather than a VectorReader. Unlike read_str, its as_str returns None, not "", for a value that is not valid UTF-8.


Constants

ConstantValueMeaning
DUCK_STRING_SIZE16Size of one duckdb_string_t in bytes
DUCK_STRING_INLINE_MAX_LEN12Longest value stored inline (no heap pointer), in bytes
MAX_STRING_LENu32::MAX (4,294,967,295)Longest VARCHAR or BLOB value DuckDB can store, in bytes

All three are in quack_rs::vector::string.