easyfabric.maintenance
optimize_by_folder
def optimize_by_folder(table_folder: str,
config_manager: ConfigManager = None,
except_folders: list[str] = None,
except_files: list[str] = None,
layers: list[str] = None,
max_workers: int = 10,
skip_missing_tables: bool = True) -> list[str]
Optimizes tables across specified lakehouseA place where you store both "raw" data (like files) and "organized" data (like tables). It combines the best of a File Cabinet and a Database. layers (Bronze, Silver, Gold) with proper existence checks and concurrent execution.
Arguments:
table_folderstr - Path to folder with YAMLA simple way to write configurations. It's basically a list that computers can read easily. table metadata.config_managerConfigManager - Config manager instance.except_foldersList[str] - Folders to exclude.except_filesList[str] - Files to exclude.layerslist[str] - List like ["Bronze", "Silver", "Gold"].max_workersint - Max concurrent optimizations.skip_missing_tablesbool - If True, skip tables that don't exist (log warning). If False, raise exception on missing table.
Returns:
list[str]- List of successfully optimized table full names.
Raises:
ValueError- Invalid layers or parameters.Exception- If a table is missing and skip_missing_tables=False.
count_by_folder
def count_by_folder(table_folder: str,
config_manager: ConfigManager = None,
except_folders: list[str] = None,
except_files: list[str] = None,
layers: list[str] = None,
schemas: list[str] = None,
max_workers: int = 10,
skip_missing_tables: bool = True) -> list[dict]
Counts rows for every folder-defined table across the requested bronze/silver
layers and schemas, emitting one MSG_MAINT_ROWCOUNT event row per table. The
row's objectname is the naked table name and its log context holds every
count as {"objectname": ..., "row_counts": [{"layer": ..., <schema>: row_count}]}, queryable downstream.
Both the current-state (dbo) and history (his) tables are counted by
default; the history count is the interesting one for change-tracked tables.
History is skipped for tables without keephistory. Counting uses Spark
when a session is active and delta-rs otherwise, so it also runs on
Spark-free Fabric compute.
Gold is model-based, not folder-based — use count_by_model for gold.
The run identifiers (run_id / activity_id / parent_run_id) are read from the logging config globals populated by set_run_id / init_logging, so each emitted row is joinable downstream on run_id without threading the ids through this call.
Arguments:
table_folderstr - Path to folder with YAML table metadata.config_managerConfigManager - Config manager instance.except_foldersList[str] - Folders to exclude.except_filesList[str] - Files to exclude.layerslist[str] - List like ["Bronze", "Silver"].schemaslist[str] - Schemas to count: ["dbo", "his"] (default: both).max_workersint - Max concurrent counts.skip_missing_tablesbool - If True, skip tables that don't exist (log warning). If False, raise exception on missing table.
Returns:
list[dict]- One entry per counted table with object, layer, schema, table and row_count.
Raises:
ValueError- Invalid layers, schemas or parameters.
count_by_model
def count_by_model(model_name: str,
config_manager: ConfigManager = None,
max_workers: int = 10,
skip_missing_tables: bool = True) -> list[dict]
Counts rows for every gold table defined by a model, emitting one
MSG_MAINT_ROWCOUNT event row per table (naked table name as objectname, count
in the log context under {"objectname": ..., "row_counts": [{"layer": "Gold", <schema>: row_count}]}).
Gold tables are enumerated from the model definition rather than an object folder, since gold is model-based. Counting uses Spark when a session is active and delta-rs otherwise.
Arguments:
model_namestr - Name of the model to count gold tables for.config_managerConfigManager - Config manager instance.max_workersint - Max concurrent counts.skip_missing_tablesbool - If True, skip tables that don't exist (log warning). If False, raise exception on missing table.
Returns:
list[dict]- One entry per counted table with object, layer, schema, table and row_count.