Repository navigation
Extract only certain columns of a csv line #166
Description
Activity
Imagine csv lines with like 150 columns in a row. But you are only interested in 20 of those values. On large data sets, this makes a huge difference to not parse just all fields of a line into the
CsvRecord.Regardless of the number of fields you are interested in, the parser always has to parse the entire record. It needs to do so to find the end of the current record and the beginning of the next one. As line breaks and delimiters can be part of quoted field values, the parser cannot simply skip parsing fields.
So, the only optimization possible is basically to avoid adding those unneeded field values to the resulting CsvRecord object.
Performance tests in the past have shown negligible impact of such an optimization. But you certainly will find use cases where it makes a difference.Having that said, you can already do what you want with a custom callback handler that only collects the desired fields into a custom record object. The example https://fastcsv.org/guides/examples/custom-callback-handler/ demonstrates how to implement such a callback handler and also show some performance measurements.
Here's a simpler example that fits your use case:
import java.io.IOException; import java.time.Instant; import de.siegmar.fastcsv.reader.AbstractBaseCsvCallbackHandler; import de.siegmar.fastcsv.reader.CsvReader; void main() throws IOException { // Some CSV data with departure and arrival timestamps in fields 5 and 8 // and irrelevant data in other fields. var csv = """ foo,,"foo, bar",,,2026-01-01T10:00:00Z,foo,"multi line\r\ndata",2026-01-01T12:00:00Z,foo ,,,,,2026-02-01T04:45:00Z,,,2026-02-01T06:15:00Z, """; try (var reader = CsvReader.builder().build(new TimetableCallbackHandler(), csv)) { for (Timetable timetable : reader) { IO.println(timetable); } } } static class TimetableCallbackHandler extends AbstractBaseCsvCallbackHandler<Timetable> { private Instant departure; private Instant arrival; @Override protected void handleBegin(long startingLineNumber) { departure = arrival = null; } @Override protected void handleField(int fieldIdx, char[] buf, int offset, int len, boolean quoted) { switch (fieldIdx) { case 5 -> departure = Instant.parse(new String(buf, offset, len)); case 8 -> arrival = Instant.parse(new String(buf, offset, len)); } } @Override protected Timetable buildRecord() { return new Timetable(departure, arrival); } } record Timetable(Instant departure, Instant arrival) { }
Implementing such a feature directly to FastCSV would probably add methods like
selectFields(int... fieldIndices)andselectFields(String... headerNames)to the existing callback handlers. While handy in some cases, I currently don't see a way to implement this without causing performance overhead in the general case.I leave this open for now to see if there are more requests for such a feature.
Reacted by kodyrecordsThanks for the insight! I adapted to your example code which works for my usecase.
You're probably right that further optimization, apart from the above, would not have enough impact on performance to justify a feature for this.
I would be great to configure the reader to only read certain columns (indexes) of a csv file.
Imagine csv lines with like 150 columns in a row. But you are only interested in 20 of those values.
On large data sets, this makes a huge difference to not parse just all fields of a line into the
CsvRecord.Example of extracting just 2 values of each row.
It would be nice to provide some kind of "tokens-map" for the reader, like:
CsvFieldcould best be an Enum, or as an alternative static String constants.The reader would then have to keep those values in some kind of a reverse map, like
Map<CsvField, Object>.So the user is then able to fetch the desired values using the
CsvFieldas key (instead of indexes access).Usage: