Skip to content

Extract only certain columns of a csv line #166

Description

@kodyrecords

I would be great to configure the reader to only read certain columns (indexes) of a csv file.

Imagine csv lines with like 150 columns in a row. But you are only interested in 20 of those values.
On large data sets, this makes a huge difference to not parse just all fields of a line into the CsvRecord.

Example of extracting just 2 values of each row.
It would be nice to provide some kind of "tokens-map" for the reader, like:

Map<Integer, CsvField> tokens = Map.of(
    5, CsvField.DEPARTURE
    8, CsvField.ARRIVAL
);

CsvField could best be an Enum, or as an alternative static String constants.

The reader would then have to keep those values in some kind of a reverse map, like Map<CsvField, Object>.
So the user is then able to fetch the desired values using the CsvField as key (instead of indexes access).

Usage:

                CsvReader<CsvRecord> reader = CsvReader.builder()
                                 .withTokens(tokens) //or similar
                                 .ofCsvRecord(file);

                reader.stream().forEach(record -> {
                       //consume the results by enums or static strings
                       record.getField(CsvField.DEPARTURE); //value at initial index 5
                       record.getField(CsvField.ARRIVAL); //value at initial index 8
                });

Activity

  1. osiegmar commented on Jan 10, 2026

    @osiegmar
    Owner

    Imagine csv lines with like 150 columns in a row. But you are only interested in 20 of those values. On large data sets, this makes a huge difference to not parse just all fields of a line into the CsvRecord.

    Regardless of the number of fields you are interested in, the parser always has to parse the entire record. It needs to do so to find the end of the current record and the beginning of the next one. As line breaks and delimiters can be part of quoted field values, the parser cannot simply skip parsing fields.

    So, the only optimization possible is basically to avoid adding those unneeded field values to the resulting CsvRecord object.
    Performance tests in the past have shown negligible impact of such an optimization. But you certainly will find use cases where it makes a difference.

    Having that said, you can already do what you want with a custom callback handler that only collects the desired fields into a custom record object. The example https://fastcsv.org/guides/examples/custom-callback-handler/ demonstrates how to implement such a callback handler and also show some performance measurements.

    Here's a simpler example that fits your use case:

    import java.io.IOException;
    import java.time.Instant;
    
    import de.siegmar.fastcsv.reader.AbstractBaseCsvCallbackHandler;
    import de.siegmar.fastcsv.reader.CsvReader;
    
    void main() throws IOException {
    
        // Some CSV data with departure and arrival timestamps in fields 5 and 8
        // and irrelevant data in other fields.
        var csv = """
            foo,,"foo, bar",,,2026-01-01T10:00:00Z,foo,"multi line\r\ndata",2026-01-01T12:00:00Z,foo
            ,,,,,2026-02-01T04:45:00Z,,,2026-02-01T06:15:00Z,
            """;
    
        try (var reader = CsvReader.builder().build(new TimetableCallbackHandler(), csv)) {
            for (Timetable timetable : reader) {
                IO.println(timetable);
            }
        }
    
    }
    
    static class TimetableCallbackHandler extends AbstractBaseCsvCallbackHandler<Timetable> {
    
        private Instant departure;
        private Instant arrival;
    
        @Override
        protected void handleBegin(long startingLineNumber) {
            departure = arrival = null;
        }
    
        @Override
        protected void handleField(int fieldIdx, char[] buf, int offset, int len, boolean quoted) {
            switch (fieldIdx) {
                case 5 -> departure = Instant.parse(new String(buf, offset, len));
                case 8 -> arrival = Instant.parse(new String(buf, offset, len));
            }
        }
    
        @Override
        protected Timetable buildRecord() {
            return new Timetable(departure, arrival);
        }
    
    }
    
    record Timetable(Instant departure, Instant arrival) {
    }

    Implementing such a feature directly to FastCSV would probably add methods like selectFields(int... fieldIndices) and selectFields(String... headerNames) to the existing callback handlers. While handy in some cases, I currently don't see a way to implement this without causing performance overhead in the general case.

    I leave this open for now to see if there are more requests for such a feature.

  2. kodyrecords commented on Jan 11, 2026

    @kodyrecords
    Author

    Thanks for the insight! I adapted to your example code which works for my usecase.

    You're probably right that further optimization, apart from the above, would not have enough impact on performance to justify a feature for this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions