Skip to content

feat: add JSON lines, Arrow IPC stream, and Parquet output formats - #26

Open
esadek wants to merge 7 commits into
mainfrom
output-formats
Open

esadek wants to merge 7 commits into
mainfrom
output-formats

Conversation

@esadek

@esadek esadek commented Jul 14, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #5
Closes #12

@esadek
esadek requested a review from ianmcook August 5, 2026 23:56
Comment thread docs/reference.md Outdated
Comment thread docs/tutorial.md Outdated
Comment thread README.md Outdated
Comment thread docs/index.md Outdated
@ianmcook

ianmcook commented Aug 6, 2026

Copy link
Copy Markdown
Member

If there's an unrecognized file extension, databow connects to the database before erroring. Is it possible to error earlier?

@esadek
esadek requested a review from ianmcook October 1, 2026 22:57
@ianmcook

ianmcook commented Oct 7, 2026

Copy link
Copy Markdown
Member

Note

I reviewed this PR and wrote this comment with help from Claude Opus 5.5.

I built this locally and tested the new output formats with DuckDB 1.5 and DataFusion 54.1. The .jsonl, .arrows, and .parquet files are valid and read back correctly in both engines. There are two issues I'd like to see addressed (here or in a follow-up), plus a few suggestions.

What I tested

  • 32 DuckDB column types written to all six formats and read back with DuckDB and DataFusion. Types: integers, floats (incl. NaN/Inf), decimals, strings with commas/quotes/newlines/unicode, dates, all timestamp units, timestamptz, time, interval, blob, UUID, lists, fixed-size arrays, structs, maps, enums, and an all-null column.
  • DataFusion as the source driver, to cover Arrow types DuckDB doesn't produce: Utf8View, BinaryView, LargeUtf8, Decimal256, Date64, Duration, LargeList, FixedSizeList, Dictionary.
  • 3M rows (multiple batches): row counts and sums match in every format and both engines.
    • Parquet is SNAPPY-compressed with 3 row groups and embeds ARROW:schema.
    • .arrows has correct stream framing (no ARROW1 magic, has the end-of-stream marker).
    • .jsonl ends with a newline.
Format DuckDB DataFusion Result
.jsonl ✅ read_json / auto-detect ✅ STORED AS JSON Same encoding as .json
.arrows ✅ read_arrow (arrow community extension) ✅ STORED AS ARROW Identical to .arrow for all types
.parquet ✅ ✅ Matches .arrow except the items below

Issues

1. A failed write leaves a corrupt file and clobbers any existing file

File::create truncates the target before writing starts (src/output.rs#L46). DuckDB INTERVAL columns arrive as Interval(MonthDayNano), which the parquet crate can't write, so:

databow --driver duckdb --query "SELECT 42 AS x" --output keep.parquet
databow --driver duckdb --query "SELECT INTERVAL 1 DAY AS iv" --output keep.parquet
# Failed to write output file: Parquet argument error: NYI: Attempting to write an Arrow interval type MonthDayNano to parquet that is not yet implemented
# keep.parquet is now 4 bytes ("PAR1"); the previous good file is gone

CSV with nested columns already does this, but Parquet with intervals makes it much more likely. Writing to a temp file in the same directory and renaming on success would fix it.

2. Zero-row results write no file and leave stale files in place

write_batches_to_file returns early when there are no batches (src/output.rs#L41-L43), so a zero-row query exits 0 without writing anything:

databow --driver duckdb --query "SELECT 'old' AS v" --output out.parquet
databow --driver duckdb --query "SELECT 'new' AS v WHERE false" --output out.parquet
# exit 0, out.parquet still contains 'old'

This also predates the PR, but it matters more now that three formats can store a schema with zero rows. The fix would be for execute_query (src/database.rs#L491) to return the reader's schema with the batches, so the writers can emit schema-only files.

Relatedly, the if batches.is_empty() guards in write_arrow_stream and write_parquet can never run. test_write_arrow_stream_empty_batches_direct and test_write_parquet_empty_batches_direct test them anyway, and they accept a 0-byte file as success.

Suggestions

  • Parquet: Date64 columns. These are written as plain INT64 with no logical type, so DuckDB and pyarrow read them as BIGINT. Adding .set_coerce_types(true) to the writer properties makes DuckDB read DATE. I checked that lists, maps, and nested structs still round-trip in both engines with it on.
  • Parquet: Timestamp(Second) columns (e.g. DuckDB TIMESTAMP_S). These are always written as bare INT64, since Parquet has no seconds unit, so DuckDB reads BIGINT. Casting to milliseconds before writing would fix that.
  • Docs. Write files with extension .arrows in Arrow IPC stream format #5 asked for an explanation of IPC file vs. stream format; right now there's only the table row. It would also help to note:
    • DataFusion doesn't recognize the .arrows/.jsonl extensions; use STORED AS ARROW / STORED AS JSON.
    • DuckDB needs the arrow community extension for both .arrow and .arrows.
  • Nits:
    • Accept .ndjson as an alias for .jsonl.
    • Match extensions case-insensitively (out.PARQUET is currently rejected).
    • The format is inferred from the extension twice (parse_output_path and write_batches_to_file); it could be stored once in AppConfig.
  • Binary size. The parquet dependency grows the stripped release binary from 6.9 MB to 9.0 MB (+30%). Probably fine, just noting it.
Other things I noticed (pre-existing or upstream, not caused by this PR)
  • JSON and JSONL omit null fields, so an all-null column disappears from the output entirely. They also write NaN/Inf as null. This is existing .json behavior that JSONL inherits; WriterBuilder::with_explicit_nulls(true) would fix the first part.
  • DuckDB misreads Parquet decimals with precision > 38: a Decimal256(50, 2) value of 12.34 comes back as 207030845.44. A pyarrow-written file does the same, and DataFusion and pyarrow read databow's file correctly, so it's a DuckDB bug. Relevant for e.g. BigQuery BIGNUMERIC.
  • DuckDB's arrow extension can't read dictionary-encoded (ENUM) or string-view columns from .arrow or .arrows.
  • DataFusion can't read the JSON-array .json output.
  • cargo clippy --all-targets -- -D warnings fails on items_after_test_module in src/database.rs. It fails on main too, and CI doesn't use those flags.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support JSONL output Write files with extension .arrows in Arrow IPC stream format

2 participants