Skip to main content

easyfabric.maintenance

optimize_by_folder​

def optimize_by_folder(table_folder: str,
config_manager: ConfigManager = None,
except_folders: list[str] = None,
except_files: list[str] = None,
layers: list[str] = None,
max_workers: int = 10,
skip_missing_tables: bool = True) -> list[str]

Optimizes tables across specified lakehouseA place where you store both "raw" data (like files) and "organized" data (like tables). It combines the best of a File Cabinet and a Database. layers (Bronze, Silver, Gold) with proper existence checks and concurrent execution.

Arguments:

  • table_folder str - Path to folder with YAMLA simple way to write configurations. It's basically a list that computers can read easily. table metadata.
  • config_manager ConfigManager - Config manager instance.
  • except_folders List[str] - Folders to exclude.
  • except_files List[str] - Files to exclude.
  • layers list[str] - List like ["Bronze", "Silver", "Gold"].
  • max_workers int - Max concurrent optimizations.
  • skip_missing_tables bool - If True, skip tables that don't exist (log warning). If False, raise exception on missing table.

Returns:

  • list[str] - List of successfully optimized table full names.

Raises:

  • ValueError - Invalid layers or parameters.
  • Exception - If a table is missing and skip_missing_tables=False.

count_by_folder​

def count_by_folder(table_folder: str,
config_manager: ConfigManager = None,
except_folders: list[str] = None,
except_files: list[str] = None,
layers: list[str] = None,
schemas: list[str] = None,
max_workers: int = 10,
skip_missing_tables: bool = True) -> list[dict]

Counts rows for every folder-defined table across the requested bronze/silver layers and schemas, emitting one MSG_MAINT_ROWCOUNT event row per table. The row's objectname is the naked table name and its log context holds every count as {"objectname": ..., "row_counts": [{"layer": ..., <schema>: row_count}]}, queryable downstream.

Both the current-state (dbo) and history (his) tables are counted by default; the history count is the interesting one for change-tracked tables. History is skipped for tables without keephistory. Counting uses Spark when a session is active and delta-rs otherwise, so it also runs on Spark-free Fabric compute.

Gold is model-based, not folder-based — use count_by_model for gold.

The run identifiers (run_id / activity_id / parent_run_id) are read from the logging config globals populated by set_run_id / init_logging, so each emitted row is joinable downstream on run_id without threading the ids through this call.

Arguments:

  • table_folder str - Path to folder with YAML table metadata.
  • config_manager ConfigManager - Config manager instance.
  • except_folders List[str] - Folders to exclude.
  • except_files List[str] - Files to exclude.
  • layers list[str] - List like ["Bronze", "Silver"].
  • schemas list[str] - Schemas to count: ["dbo", "his"] (default: both).
  • max_workers int - Max concurrent counts.
  • skip_missing_tables bool - If True, skip tables that don't exist (log warning). If False, raise exception on missing table.

Returns:

  • list[dict] - One entry per counted table with object, layer, schema, table and row_count.

Raises:

  • ValueError - Invalid layers, schemas or parameters.

count_by_model​

def count_by_model(model_name: str,
config_manager: ConfigManager = None,
max_workers: int = 10,
skip_missing_tables: bool = True) -> list[dict]

Counts rows for every gold table defined by a model, emitting one MSG_MAINT_ROWCOUNT event row per table (naked table name as objectname, count in the log context under {"objectname": ..., "row_counts": [{"layer": "Gold", <schema>: row_count}]}).

Gold tables are enumerated from the model definition rather than an object folder, since gold is model-based. Counting uses Spark when a session is active and delta-rs otherwise.

Arguments:

  • model_name str - Name of the model to count gold tables for.
  • config_manager ConfigManager - Config manager instance.
  • max_workers int - Max concurrent counts.
  • skip_missing_tables bool - If True, skip tables that don't exist (log warning). If False, raise exception on missing table.

Returns:

  • list[dict] - One entry per counted table with object, layer, schema, table and row_count.