Skip to content

[Feature Request] Add pyiceberg.catalog.hadoop.HadoopCatalog (filesystem-only catalog) #3897

Description

@fightBoxing

Is your feature request related to a problem? Please describe.

Java Iceberg ships a filesystem-only HadoopCatalog / HadoopTables, where table metadata lives under <warehouse>/<db>.db/<table>/metadata/ with no external metastore. PyIceberg currently has no equivalent — the available catalog types are rest / hive / glue / dynamodb / sql / in-memory / bigquery.

This gap matters in two ways:

  1. Interop with Java-side HadoopCatalog tables. Tables created by Java HadoopCatalog (common in lightweight deployments without a metastore) cannot be opened through any supported PyIceberg catalog. Users must fall back to StaticTable.from_metadata and resolve the latest metadata.json themselves, which loses catalog semantics (no namespace listing, no create/commit).

  2. Downstream projects already assume the module exists. Daft's Gravitino integration imports from pyiceberg.catalog.hadoop import HadoopCatalog and calls HadoopCatalog("gravitino_reader", props).load_table(table_dir) to open a table from a storage location (daft/catalog/__gravitino/_catalog.py). Against PyIceberg 0.11.x this raises TypeError: HadoopCatalog.__init__() takes 2 positional arguments but 3 were given, and against versions without the module it fails at import time.

Describe the solution you'd like

A pyiceberg.catalog.hadoop.HadoopCatalog (subclassing MetastoreCatalog) implementing Java HadoopCatalog semantics:

  • warehouse property as the root location
  • table dir = <warehouse>/<namespace>/<table>
  • metadata at <table_dir>/metadata/v{n}.metadata.json plus a version-hint.text holding the current version
  • latest-version resolution: read version-hint.text, fall back to scanning metadata/ for the max v{n} (matching Java behavior)
  • namespace/table create/list/commit driven purely by the warehouse filesystem (no metastore calls)

Additional context / pitfalls observed while prototyping

Happy to contribute a PR if this is in scope. A few notes from an internal prototype:

  1. Metadata file naming. Java HadoopCatalog uses v{n}.metadata.json + version-hint.text, but tables created by JDBC/REST catalogs use 00000-<uuid>.metadata.json with no version-hint. To open those as well, the scan fallback should accept both patterns (v(\d+)\.metadata\.json and \d{5}-.*\.metadata\.json), or at least document the limitation.

  2. Filesystem abstraction. __init__ should derive the filesystem from the catalog's FileIO (PyArrowFileIO) instead of hardcoding pyarrow.fs.HadoopFileSystem.from_uri(warehouse). The JVM-backed HadoopFileSystem only supports hdfs:// and fails for object-store schemes (s3://, and custom schemes), so routing through FileIO keeps it scheme-agnostic.

  3. Atomicity. create_table / commit_table use create-if-absent on v{n}.metadata.json for optimistic concurrency — safe on HDFS but not atomic on plain object stores (S3 has no create-if-absent guarantee). Java has the same caveat; worth documenting or using a conditional-write primitive where available.

Adapting a custom storage scheme (Tencent Cloud TBDSFS as a concrete case)

A related gap surfaced while prototyping against Tencent Cloud TBDS's distributed filesystem scheme tbdsfs://<cluster>/<path> (exposed by a Python client, plus a JVM fs.tbdsfs.impl):

  • PyArrowFileIO._initialize_fs(scheme, netloc) only understands a fixed set of schemes (hdfs / s3 / gs / file / abfs / ...), so any tbdsfs://... location raises ValueError: Unrecognized filesystem type in URI: tbdsfs.

  • There is no public, documented way to plug in a custom filesystem. The only workaround today is monkey-patching a private method:

    1. Implement a pyarrow.fs.FileSystemHandler subclass wrapping the TBDSFS Python client, wrap it in pyarrow.fs.PyFileSystem, then patch PyArrowFileIO._initialize_fs to return that filesystem when scheme == "tbdsfs".
    2. pyarrow 21's PyFileSystem callback also has non-obvious contracts any custom handler must satisfy: single paths arrive as one-element lists; get_file_info must return a one-element list; and for scheme'd URIs the netloc is prepended into the path (e.g. internal/usr/..., without a leading slash).

This works, but it depends on patching a private API (_initialize_fs), which is brittle across PyIceberg releases.

Suggested improvement: a documented, public extension point for registering an arbitrary pyarrow.fs.FileSystem (or a custom FileIO) per scheme — e.g. a register_file_system(scheme, factory) helper, or a scheme → FileSystem mapping read from FileIO/catalog properties — so non-standard object stores and filesystems can be integrated without touching internals. This would also naturally address the HadoopCatalog.__init__ filesystem-abstraction point above.

References

  • Java: org.apache.iceberg.hadoop.HadoopCatalog / HadoopTables
  • Downstream usage that currently breaks: daft/catalog/__gravitino/_catalog.py_open_iceberg_table

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions