NULL Handling & Strings
This page covers checking for NULL before reading a DuckDB vector, writing NULL
output, and reading and writing VARCHAR and BLOB values. The two topics
belong together: reading a string from a NULL row is undefined behaviour, not
just a wrong value.
NULL checks
Every row in a DuckDB vector may be NULL. Always check validity before reading:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, writer: &mut VectorWriter) { for row in 0..reader.row_count() { if unsafe { !reader.is_valid(row) } { // Propagate NULL to output unsafe { writer.set_null(row) }; continue; } // Safe to read let value = unsafe { reader.read_str(row) }; } } }
Reading a fixed-width value from a NULL row returns garbage. The vector's data buffer is not zeroed at NULL positions, and no error is raised: you get whatever bytes the buffer holds at that position.
For VARCHAR and BLOB it is worse than garbage. A NULL row's 16-byte entry
is left as it was, and DuckDB reuses vector buffers between chunks, so the
entry can still be the pointer-format record of a string from an earlier chunk
whose memory may since have been freed or reused. read_str / read_blob on such a row
follow that pointer: undefined behaviour, not merely a wrong answer. Their
# Safety sections require the row to be valid.
Writing NULL
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(writer: &mut VectorWriter, row: usize) { unsafe { writer.set_null(row) }; } }
Pitfall L4:
VectorWriter::set_nullcallsduckdb_vector_ensure_validity_writablebefore accessing the validity bitmap. Without it, a vector that has no mask yet makesduckdb_vector_get_validityreturn NULL, and the NULL you write is silently dropped. Never write NULL manually; always useset_null— which, for aSTRUCTorARRAYoutput, also nulls the fields / elements of that row the way DuckDB expects. See Pitfall L4.
Clearing NULL (v0.11.0+)
To mark a row as valid after a previous set_null:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(writer: &mut VectorWriter, row: usize) { unsafe { writer.set_valid(row) }; } }
VARCHAR reading
Read VARCHAR columns with VectorReader::read_str:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, row: usize) { let s: &str = unsafe { reader.read_str(row) }; } }
The returned &str borrows from the DuckDB vector — it must not outlive the
callback. Do not store it in a struct; clone it to a String if you need to
keep it.
The duckdb_string_t format
Pitfall P7: the Rust bindings do not document the layout of
duckdb_string_t.quack-rsdecodes it for you; the details below are for reference. See Pitfall P7.
DuckDB stores VARCHAR values in a 16-byte duckdb_string_t struct with two
representations, selected at runtime based on string length:
| Format | Condition | Layout |
|---|---|---|
| Inline | length ≤ 12 | [len: u32][data: [u8; 12]] |
| Pointer | length > 12 | [len: u32][prefix: [u8; 4]][ptr: *const u8] |
On a 32-bit target (DuckDB-WASM) the pointer is 4 bytes and the last 4 bytes are unused. The length and the pointer are in the target's byte order.
VectorReader::read_str and the underlying read_duck_string function handle
both formats, so you do not need to inspect the raw struct. A value that is not
valid UTF-8 is returned as ""; use read_blob to get its bytes.
Empty strings vs NULL
An empty string ("") and NULL are distinct values:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, row: usize) { // NULL: is_valid returns false // Empty string: is_valid returns true, read_str returns "" if unsafe { !reader.is_valid(row) } { // This is NULL } else { let s = unsafe { reader.read_str(row) }; if s.is_empty() { // This is an empty string, not NULL } } } }
Writing VARCHAR
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(writer: &mut VectorWriter, row: usize, my_str: &str) { unsafe { writer.write_varchar(row, my_str) }; // &str } }
write_varchar copies the string bytes into DuckDB's managed storage, so the
&str need not outlive the call. write_blob does the same for &[u8].
Both panic for a value longer than MAX_STRING_LEN (u32::MAX bytes, the
most a duckdb_string_t length field can hold) instead of storing a truncated
value; inside scalar_callback! and the typed scalar constructors the panic
becomes a SQL error. try_write_varchar and try_write_blob return
Result<(), ExtensionError> instead and write nothing on error:
#![allow(unused)] fn main() { use quack_rs::vector::VectorWriter; use quack_rs::error::ExtensionError; fn demo(writer: &mut VectorWriter, row: usize, my_str: &str) -> Result<(), ExtensionError> { unsafe { writer.try_write_varchar(row, my_str)? }; Ok(()) } }
Reading BLOB values
BLOB uses the same inline/pointer layout as VARCHAR, but may contain any
bytes. Use read_blob so the data is not interpreted as UTF-8:
#![allow(unused)] fn main() { use quack_rs::vector::{VectorReader, VectorWriter}; fn demo(reader: &VectorReader, row: usize) { let bytes: &[u8] = unsafe { reader.read_blob(row) }; } }
Like read_str, the returned slice borrows from the DuckDB vector and must not
outlive the callback.
Complete NULL-safe VARCHAR pattern
#![allow(unused)] fn main() { use libduckdb_sys::{duckdb_data_chunk, duckdb_function_info, duckdb_vector}; use quack_rs::vector::{VectorReader, VectorWriter}; unsafe extern "C" fn my_scalar( _info: duckdb_function_info, input: duckdb_data_chunk, output: duckdb_vector, ) { let reader = unsafe { VectorReader::new(input, 0) }; let mut writer = unsafe { VectorWriter::new(output) }; for row in 0..reader.row_count() { if unsafe { !reader.is_valid(row) } { unsafe { writer.set_null(row) }; continue; } let s = unsafe { reader.read_str(row) }; let upper = s.to_uppercase(); unsafe { writer.write_varchar(row, &upper) }; } } }
DuckStringView
For advanced use cases where you need access to the raw string bytes or the
inline/pointer distinction, quack_rs::vector::string::DuckStringView is
available:
#![allow(unused)] fn main() { fn demo(data: *const u8, idx: usize) { use quack_rs::vector::string::{DuckStringView, DUCK_STRING_SIZE}; // From raw 16-byte data (inside a vector callback). `from_raw` is unsafe because it // follows the pointer of a string longer than 12 bytes; untrusted bytes go through // `DuckStringView::inline_from_bytes`, which refuses that format instead. let raw: &[u8; 16] = unsafe { &*data.add(idx * DUCK_STRING_SIZE).cast() }; let view = unsafe { DuckStringView::from_raw(raw) }; println!("length: {}", view.len()); println!("is_empty: {}", view.is_empty()); if let Some(s) = view.as_str() { println!("content: {s}"); } } }
In practice, prefer reader.read_str(row). DuckStringView is needed only when
you have a raw data pointer rather than a VectorReader. Unlike read_str, its
as_str returns None, not "", for a value that is not valid UTF-8.
Constants
| Constant | Value | Meaning |
|---|---|---|
DUCK_STRING_SIZE | 16 | Size of one duckdb_string_t in bytes |
DUCK_STRING_INLINE_MAX_LEN | 12 | Longest value stored inline (no heap pointer), in bytes |
MAX_STRING_LEN | u32::MAX (4,294,967,295) | Longest VARCHAR or BLOB value DuckDB can store, in bytes |
All three are in quack_rs::vector::string.