Repository navigation
Conversation
|
If there's an unrecognized file extension, databow connects to the database before erroring. Is it possible to error earlier? |
Co-authored-by: Ian Cook <ianmcook@gmail.com>
Co-authored-by: Ian Cook <ianmcook@gmail.com>
Co-authored-by: Ian Cook <ianmcook@gmail.com>
Co-authored-by: Ian Cook <ianmcook@gmail.com>
|
Note I reviewed this PR and wrote this comment with help from Claude Opus 5.5. I built this locally and tested the new output formats with DuckDB 1.5 and DataFusion 54.1. The What I tested
Issues1. A failed write leaves a corrupt file and clobbers any existing file
databow --driver duckdb --query "SELECT 42 AS x" --output keep.parquet
databow --driver duckdb --query "SELECT INTERVAL 1 DAY AS iv" --output keep.parquet
# Failed to write output file: Parquet argument error: NYI: Attempting to write an Arrow interval type MonthDayNano to parquet that is not yet implemented
# keep.parquet is now 4 bytes ("PAR1"); the previous good file is goneCSV with nested columns already does this, but Parquet with intervals makes it much more likely. Writing to a temp file in the same directory and renaming on success would fix it. 2. Zero-row results write no file and leave stale files in place
databow --driver duckdb --query "SELECT 'old' AS v" --output out.parquet
databow --driver duckdb --query "SELECT 'new' AS v WHERE false" --output out.parquet
# exit 0, out.parquet still contains 'old'This also predates the PR, but it matters more now that three formats can store a schema with zero rows. The fix would be for Relatedly, the Suggestions
Other things I noticed (pre-existing or upstream, not caused by this PR)
|
Closes #5
Closes #12