Arrow Interop
Requires the
duckdb-1-5-4feature flag.
DuckDB's C API has a family of conversion functions (present since 1.4.4) that
move data directly between a duckdb_data_chunk and the
Arrow C Data Interface,
without a query result in between. The quack_rs::arrow module wraps all eight
of them.
No arrow crate dependency
The Arrow C Data Interface is an ABI, not a library: ArrowSchema and
ArrowArray are plain #[repr(C)] records with a release callback.
libduckdb-sys defines them directly — and asserts in its own test suite that
they match arrow-rs's FFI_ArrowSchema / FFI_ArrowArray field for field — so
quack-rs speaks Arrow without pulling in the arrow crate, and an extension
that does use arrow-rs bridges across with a pointer cast.
Why the feature is duckdb-1-5-4 and not duckdb-1-5
All eight C functions are already in duckdb_ext_api_v1 in DuckDB 1.4.4
(extension_api.hpp at v1.4.4, slots 410 to 434; they moved to 411 to 509 in
1.5.0). The floor comes from the bindings: libduckdb-sys declared both records
as opaque zero-sized bindgen placeholders (_unused: [u8; 0]) until
1.10504.0, and the caller-allocated structs these functions need cannot be
created from a zero-sized type. src/arrow.rs carries a const assertion that
fails with exactly that message against an older binding. Because the feature
implies duckdb-1-5, an extension built with it also needs a DuckDB 1.5.0+
engine.
The types
| Type | Wraps | Freed by |
|---|---|---|
ArrowOptions | duckdb_arrow_options | duckdb_destroy_arrow_options |
ArrowSchema | the ArrowSchema ABI record | release(schema) |
ArrowArray | the ArrowArray ABI record | release(array) |
ArrowConvertedSchema | duckdb_arrow_converted_schema | duckdb_destroy_arrow_converted_schema |
RawArrowSchema and RawArrowArray are the ABI records themselves, re-exported
for code that already speaks the raw interface.
Exporting a chunk
ArrowOptions carries the settings DuckDB renders Arrow with — the timezone for
TIMESTAMPTZ, the string offset width, registered extension types. Take them
from the result whose chunks you are exporting, so schema and data agree:
#![allow(unused)] fn main() { use quack_rs::arrow::{data_chunk_to_arrow, to_arrow_schema}; use quack_rs::query::QueryResult; fn demo(result: &mut QueryResult) -> Result<(), Box<dyn std::error::Error>> { // SAFETY: the connection that ran this query stays open while `options` is // in use — the options point at that connection's client context. let options = unsafe { result.arrow_options() }?; let columns: Vec<(String, quack_rs::types::LogicalType)> = (0..result.column_count()) .filter_map(|i| Some((result.column_name(i)?, result.column_logical_type(i)?))) .collect(); let pairs: Vec<(&str, &quack_rs::types::LogicalType)> = columns.iter().map(|(n, t)| (n.as_str(), t)).collect(); let schema = to_arrow_schema(&options, &pairs)?; assert_eq!(schema.format(), Some("+s")); // a record batch is a struct while let Some(chunk) = result.next_chunk()? { let array = data_chunk_to_arrow(&options, &chunk)?; // hand `array` (plus `schema`) to any Arrow consumer let _ = array; } Ok(()) } }
ArrowOptions must not outlive its connection
The options hold a raw pointer to the connection's client context, and the
conversion functions dereference it. Used after the connection is closed they
read freed memory. ArrowOptions<'conn> carries that lifetime:
ArrowOptions::from_connection(&con)is safe. It borrows theOwnedConnection, so the compiler rejects adrop(con)while the options are still in use.ArrowOptions::from_raw_connection(raw)(for a rawduckdb_connection) andQueryResult::arrow_options()/ArrowOptions::from_resultareunsafe. AQueryResultdoes not borrow the connection that ran it, so nothing checks that the connection is still open. The caller must keep it open for as long as the options are used.
Importing an array
Going the other way needs the Arrow schema translated into DuckDB's own type descriptors first. That translation is reusable — do it once, not per batch:
#![allow(unused)] fn main() { use quack_rs::arrow::{data_chunk_from_arrow, schema_from_arrow, ArrowArray, ArrowSchema}; fn demo( con: libduckdb_sys::duckdb_connection, schema: &mut ArrowSchema, array: ArrowArray, ) -> Result<(), Box<dyn std::error::Error>> { // SAFETY: `con` is a live DuckDB connection. let converted = unsafe { schema_from_arrow(con, schema) }?; // SAFETY: same connection; `array` was built against `schema`. let chunk = unsafe { data_chunk_from_arrow(con, array, &converted) }?; let _ = chunk.size(); Ok(()) } }
Note the asymmetry, which mirrors what DuckDB actually does:
schema_from_arrowborrows the schema. You still own it and it is released when itsArrowSchemadrops.data_chunk_from_arrowtakes the array by value. DuckDB setsarrow_array->release = nullptrbefore the conversion loop body — so it claims the array even when the conversion then fails. The by-value binding is still dropped on the way out, which releases the array in the one case where DuckDB does not claim it (a zero-column schema, where the loop never runs).
The resulting chunk keeps the Arrow buffers alive, so most columns share the data rather than copy it. Dictionary-encoded columns are the exception: the wrapper copies them into flat vectors.
What the wrapper refuses that DuckDB would not
duckdb_data_chunk_from_arrow indexes arrow_array->children[i] once per column
in the converted schema with no bounds check, dereferences each child without a
null check, reads offset + length rows from each child without comparing its
length, and dereferences an array without checking whether it was already
released. Those are segfaults or out-of-bounds reads, not errors.
data_chunk_from_arrow checks them first — which is why ArrowConvertedSchema
remembers the column count of the schema it was built from — and returns an
InvalidInput error instead.
It also refuses a zero-row array. DuckDB passes arrow_array->length
through as the chunk's capacity, and a capacity of zero trips a
D_ASSERT(size > 0) that aborts a debug build of DuckDB, while a release build
carries on. Skip empty batches, or create the empty chunk directly with
duckdb_create_data_chunk.
What it cannot check, and what data_chunk_from_arrow's # Safety section
therefore makes the caller's job:
- The array must conform to the converted schema. Nothing in an Arrow
array records its type, so DuckDB reads each child's buffers as the format
the schema declares. An
int32child imported under autf8schema has its values read as string offsets into a buffer that does not exist. Arrays exported withdata_chunk_to_arrowunder the schema you converted conform. - A fixed-width dictionary whose indices can be NULL needs one element of
padding past its values: DuckDB points NULL indices at an entry there,
and the copy
data_chunk_from_arrowmakes of every dictionary column reads it (docs/upstream-duckdb-reports.md, item 29). - The buffers must be as long as the lengths say, and
lengthmust be the true row count. DuckDB allocates the chunk forlengthrows before its error handling starts, so an absurd length is an allocation failure that aborts the process.
It also refuses valid Arrow layouts that DuckDB imports wrongly: it walks the
array alongside its schema and returns InvalidInput, naming the node, for
- an offset below the top level that DuckDB applies to the wrong rows: a
struct inside an offset struct or a list, a union's members, a run-end-encoded
array's value validity (
docs/upstream-duckdb-reports.md, item 24); - a dictionary with NULLs under a list that starts past element 0, or with more than 2048 rows and NULLs of its own or an enclosing struct's, which DuckDB copies past a 2048-row heap mask (items 9 and 25);
- a dictionary whose values are themselves dictionary-encoded (item 26);
- list views that overlap or leave gaps (item 27);
- a sparse union whose
+us:type codes are not0, 1, …(item 28), or whosenull_countis not 0 (item 32); - a dictionary whose
null_countis -1 ("not computed", item 31); - a
geoarrow.wkbcolumn read as more than 2048 rows (item 33); - a run-end-encoded array where DuckDB reads a plain one: a fixed-size list's child, or another run-end array's values (item 24).
Arrays that data_chunk_to_arrow produced, paired with the schema they were
produced with, never take these shapes. Arrays from other producers can: one
that slices a nested array without copying it may leave offsets below the top
level. Copying the slice before export avoids them. Every error DuckDB
reports from the conversion arrives as InvalidInput.
Round trips are not always exact
Two types come back different from an Arrow round trip through DuckDB's own converters:
TIMETZcomes back asTIMEwith the offset dropped:01:02:03+05:30returns as01:02:03.BITcomes back asBLOB.
Check the converted types (ArrowConvertedSchema) when a round trip must be
lossless.
DuckDB would export three kinds of value wrongly, with no error (checked on
1.4.4, 1.5.0 and 1.5.5), so data_chunk_to_arrow checks the chunk first and
refuses one that holds such a value, at any nesting depth:
- An
INTERVALwhose microseconds exceed about ±106,751 days (2,562,047 hours) would wrap, because Arrow counts nanoseconds in ani64and DuckDB multiplies by 1000 unchecked:INTERVAL 2562048 HOURwould export as a negative interval. - A 39-digit
UHUGEINTwould export as adecimal128(38, 0)it does not fit; from 2^127 it comes out negative (2^128 - 1becomes-1). - A 39-digit
HUGEINTwould export as adecimal128(38, 0)it does not fit, unlessarrow_lossless_conversionis set (then it exports as a 16-byte fixed-size binary and is not refused).
After the export, the array is also checked against the schema DuckDB
declares for the chunk's types. Before 1.5.5, BIGNUM (and, from 1.5.0,
GEOMETRY) exported under arrow_output_version = '1.4' is written as
binary views while the schema declares plain binary, which a consumer reads as
offsets (docs/upstream-duckdb-reports.md, item 34); such an export is
refused. Set arrow_output_version = '1.0' on those releases.
Bridging to arrow-rs
This sketch uses the arrow crate's FFI_ArrowArray, which quack-rs does not
depend on, so it is not compiled with the book:
// quack-rs -> arrow-rs
let ffi: FFI_ArrowArray = unsafe { std::mem::transmute(array.into_raw()) };
// arrow-rs -> quack-rs, neutralising the source so only one side releases
let array = unsafe { ArrowArray::take_from(std::ptr::from_mut(&mut ffi).cast()) };
take_from moves the record out and writes a released placeholder back, so the
foreign wrapper's own Drop becomes a no-op instead of a double free.
Thread safety
None of these types are Send or Sync. The Arrow C Data Interface says nothing
about which thread may call release, and duckdb_arrow_options wraps a
ClientProperties tied to the connection's client context.