Load Bronze
The load_data_bronze module is the entry point for extracting data from source systems and landing it in the Bronze layer. It handles multiple file formats (CSV, JSON, XML, etc.) and ensures that raw data is captured with minimal transformation while managing pre- and post-processing notebook hooks. The module includes intelligent file change detection to skip unnecessary loads and validation checks to ensure data quality.
A Load Bronze run completing — the activity panel reports 5 files loaded for table adv_products in 1 min 39 sec.
Overview
The bronze loader performs the following workflow:
- Source Check (optional): For Azure Blob sources, checks blob metadata without downloading to determine if any files have changed
- File Pulling: Downloads files from the configured source system
- Freshness Validation: Verifies that source files are not stale based on configurable age thresholds
- Change Detection (optional): Compares current files against a snapshot to skip loading if nothing has changed
- Data Loading: Processes files according to their file type and loads into Bronze tables
- History Tracking (optional): Maintains history records and validates that file changes correlate with data changes
Key Features
Skip-if-Unchanged (bronzeskipifsourceunchanged)
This feature prevents unnecessary bronze loads when source files have not changed. It works in two stages:
Stage 1: Azure Blob Pre-Check (for Azure Blob sources)
- Compares blob metadata (MD5 hash and size) without downloading files
- If all blobs match previous snapshot, entire load is skipped
- Significantly reduces bandwidth and processing time
- The freshness check (Validation 1) still runs against the listed metadata before skipping, so an unchanged-but-stale source — a supplier that stopped delivering, leaving an old file in place — is surfaced rather than silently skipped
Stage 2: File Change Detection (after file pull)
- Compares pulled files against previously saved snapshot
- Uses MD5 hash or file size for comparison
- Skips load if no changes detected
Enable this feature by setting bronzeskipifsourceunchanged: true in your table configuration.
What both stages log. A skipped load is only trustworthy if you can see what it was based on, so each stage reports how many files the sourcefilter matched, how the comparison came out, and how fresh the source is:
Source pre-check: 8 file(s) match "transactions_*" in container "sourcefiles" folder "exports/current".
Compared 8 file(s) against tracker: 8 unchanged, 0 changed, 0 new, 0 no longer delivered — basis: MD5 8, size 0, not comparable 0.
Freshness: newest transactions_eu.csv 3.2h, oldest transactions_apac.csv 26.5h (max 48h).
Source files unchanged for dp_transactions_csv — skipping bronze load.
File names are listed only for the files that deviate — changed, new, no longer delivered, or not comparable — capped at ten names with a remainder count, so a batch of hundreds stays readable:
Tracker delta — changed: transactions_eu.csv; new: transactions_latam.csv; no longer delivered: transactions_us.csv; not comparable: none.
No longer delivered means the tracker still knows a file that the current sourcefilter no longer matches. Not comparable means neither MD5 nor size was available on both sides; such a file counts as unchanged and never triggers a reload on its own. When the filter matches nothing at all, the pre-check logs a warning and falls back to a full pull.
The basis matters. Azure only exposes a blob's Content-MD5 when the uploader set that header. A supplier that uploads without it leaves both the listing and the tracker with sizes only, and the comparison silently falls back to size. The skip reason recorded on the segment names what the verdict rested on, so this is visible per run rather than assumed: source_unchanged_md5, source_unchanged_size, source_unchanged_mtime or source_unchanged_none for the pre-check, joined with + when one run took more than one route (source_unchanged_size+mtime), and the matching files_unchanged_* for the post-download check.
Size and MD5 are not a weak and a strong version of one thing — they measure differently. Size is a coarse checksum that ignores row order, so a supplier who re-exports the same rows in a different order every day keeps reporting unchanged, which is usually what you want. MD5 over those same deliveries changes daily and would reload every run, defeating the skip entirely. Whether Content-MD5 is an upgrade therefore depends on the source: ask for it when deliveries are byte-stable, leave it off when row order drifts.
What size cannot see is a change that leaves the byte length intact — a corrected value of equal length, a row swapped for another of the same width. Added or removed rows, changed field widths and different column sets all change the size.
Bounding the skip (bronzemaxunchangedhours)
bronzemaxunchangedhours caps how long the source may stay unchanged before one full load runs regardless. It turns an open-ended blind spot into one that is at most one window wide, at the cost of one full load per window:
bronzeskipifsourceunchanged: true
bronzemaxunchangedhours: 24 # after a day without a real change, load anyway
bronzeunchangedseverity: warning
Set per table, falling back to the connection. Unset — the default — leaves skipping unbounded, so existing behaviour is unchanged until you opt in. The window is measured from the last run that actually loaded, not the last run: skipped runs return before the tracker is written, so the newest captured_at in the tracker is exactly that timestamp. A tracker without a usable timestamp counts as expired, so the next run writes one and the bound applies from then on. When the window has passed the run says so and goes straight to the pull, skipping the pre-check altogether:
Source of dp_transactions_csv has not changed for 26.4h, past bronzemaxunchangedhours (24h) — loading in full instead of skipping.
That line is reported at bronzeunchangedseverity, on the same four rungs as the other checks, so a monitor can route it rather than having to spot it in the log. none silences the report without switching the window off — the full load still happens.
Worth setting whenever the reason reads source_unchanged_size, where byte length is all that was compared and a changed delivery of identical length is invisible; less pressing on source_unchanged_mtime, where the modification time had to match as well, and unnecessary on source_unchanged_md5.
bronzemaxunchangedhours replaces maxskiphours, which named what the loader did (stop skipping) rather than what is measured, and whose expiry was only logged in passing. The old key is no longer read: rename it.
File Freshness Validation (Validation 1)
Ensures source files are recent enough for processing:
- Checks the
last_modifiedtimestamp for each file. Files that have nolast_modifiedtimestamp are exempt and always pass the check. - Compares against
bronzemaxfileagehours(default: 2 hours) - Runs both after the file pull and on the Azure Blob unchanged fast-path, so a stale source is caught even when the load would otherwise be skipped as unchanged
- The response is governed by
bronzeloadviolationaction(table, falling back to the connection):continuelogs a warning and proceeds (or skips, if unchanged),stopraisesFabricStaleFileError,silentproceeds without logging anything - Bypass per-run with the
skip_stale_checkoverride
The check runs on the whole batch before the Bronze table is truncated and before any file is loaded, and it is all-or-nothing:
- Under the default
stop, the first stale file in the batch raisesFabricStaleFileErrorand aborts the entire Bronze run for that table. Nothing is written — not the stale file, and not the fresh files pulled alongside it. Because the abort happens before the truncate, the existing Bronze table is left untouched. - Under
continue, every stale file is logged as a warning and the run proceeds normally, loading all files (the stale one included). Stale files are never skipped individually.
How loudly it is reported (bronzeloadseverity)
bronzeloadviolationaction decides whether the load goes on. It cannot say how
much the result matters, and that is a separate question: a stale transaction
file is urgent, a stale product table for one day is not, and both may want
continue. bronzeloadseverity carries that second answer.
bronzeloadviolationaction: continue # the load proceeds
bronzeloadseverity: critical # but wake someone
Four rungs: none, info, warning (the default) and critical. Set per
table, falling back to the connection. The rung lands in the log_level column
of Meta.dbo.logging: info as INFO, warning as WARNING, critical as
CRITICAL. A monitor filters on log_level IN ('WARNING', 'CRITICAL'), or on
'CRITICAL' alone for the urgent ones. This applies to continue; stop always
logs ERROR and skip always logs INFO, because there the action decides.
none writes no log line; the result is still recorded on the file tracker. It is
the successor to bronzeloadviolationaction: silent, which does the same and keeps
working with a deprecation notice. To keep a record in the log without alerting
anyone, use info.
Validation 2 has its own rung, bronzechecksummismatchseverity, on the same
four values. Sources whose row order drifts trip Validation 2 daily without
anything being wrong, so this is the setting that turns that noise down to
info for one object without silencing it everywhere.
Bounding how long a warning may repeat (bronzemaxfileagecriticalhours)
A warning that comes back every morning for three weeks stops being read. A second age threshold turns that stretch into an escalation:
bronzemaxfileagehours: 2 # stale from here — reported at the base rung
bronzemaxfileagecriticalhours: 72 # still stale here — reported as critical
bronzeloadviolationaction: continue
Past the second threshold the same file reports as critical instead of its
base rung. Escalation moves the severity and never the action: a load configured
to continue keeps continuing, however long the delivery has been missing.
Unset — the default — never escalates, so a file stale for three weeks reports
exactly as loudly as one stale for three hours.
Only one line per file: past the critical threshold you get the critical message
instead of the base one, not both. bronzemaxfileagecriticalhours must be above
bronzemaxfileagehours, otherwise the base rung is unreachable — set both on the same
object and the config is refused at load; split across a connection and an
object, where neither half is wrong on its own, the run reports it instead.
History Validation (Validation 2)
When history is enabled (keephistory: true) and bronzeskipifsourceunchanged is on, validates data integrity per source file:
- Tracks MD5 hash and file size for each source file
- Counts the rows this run added to the history table (
Bronze.his) per source file, grouped on theSYSTEMSOURCETAGsystem column (which carries the file's abfs path) - Flags every file whose MD5 (or, as fallback, size) differs from its previous tracker entry while its own new-row count is 0. Because the count is per file, one file that did bring new rows in a multi-file batch does not silence the warning for a file that did not
- Files without a previous tracker entry (first delivery) are never flagged
- It is a heuristic (checksum/size diff without a matching row-count change), not a guaranteed data problem, since dedup/upsert logic in the history write can legitimately produce zero new rows for a byte-different-but-semantically-identical file — a re-delivery with the same rows in a different order, for instance
- Logs detailed change information for debugging and writes the per-file outcome (
changed,new_history_rows) to the file tracker - The response is governed by its own
bronzechecksummismatchviolationaction(default:continue, i.e. one warning per flagged file;stopraises on the first flagged file;silentdoes neither, useful when this heuristic is known to be noisy for a given source), independent ofbronzeloadviolationaction— so astopfreshness policy for Validation 1 doesn't also turn this softer signal into a hardFabricChecksumMismatchError
File Tracker
The file tracker maintains snapshots of source files in the Bronze layer:
- Stores metadata (MD5, size, timestamp) in NDJSON format — one entry per file per run, appended oldest-first
- Located at
<bronze table folder>/_tracking/tracker.json(seeget_tracker_file_path) - Used for change detection without re-downloading files
- Automatically updated after successful loads
One tracker file can hold the entries of more than one object: for fabricfiles the path is derived from the Bronze folder and the source folder, so two objects reading the same source folder share it. Every entry therefore records the object that wrote it in shortcode, and a load reads back only its own entries — otherwise the second object would compare against the first object's snapshot, conclude the source was unchanged and skip its own load. Entries written before the wheel filtered on shortcode cannot be attributed and are ignored, so the first run after upgrading reloads once.
Entries are keyed on partial_filename: the file's name in the source — the basename of the blob for azblob, the file name for fabricfiles. It is that source name, not the Bronze file name, because azblob prefixes every downloaded file with a run-specific {hhmm}_{rand}_. Matching a delivery against its previous one, and pattern rules such as startswith("orders_2026") below, therefore work on the name the supplier delivers.
Every entry also records what the two validations concluded for that file in that run:
| Field | Type | Meaning |
|---|---|---|
stale | bool or null | Validation 1: true when the file's last_modified is older than bronzemaxfileagehours, false when it is fresh. null when the file has no last_modified or the stale check was skipped (skip_stale_check) |
changed | bool or null | Whether the file's MD5 (or, as fallback, size) differs from its previous tracker entry. null when the file has no previous entry (first delivery) or nothing could be compared |
new_history_rows | int or null | Rows this run added to Bronze.his for this file (grouped on SYSTEMSOURCETAG). null when Validation 2 did not run — keephistory: false, bronzeskipifsourceunchanged off, or the count failed |
The tracker is stateless: each entry compares this run with the previous delivery of the same file only. No verdict is carried forward, and the ViolationAction settings decide what the loader itself does with these facts. Which combination is a problem for a given source — a current file that did not change versus a three-year-old file that is expected to stay unchanged — is a customer rule, which is what the fields are exposed for.
Using the tracker from a post-bronze notebook
A postbronzenotebook runs after the tracker was written, so it can read the entries of the run that just finished and apply its own severity rules. All entries of one run share the same captured_at:
from easyfabric.loaders import get_tracker_file_path, load_previous_snapshot
entries = load_previous_snapshot(get_tracker_file_path(bronze_table_folder), dataplatformobjectname)
this_run = max(e["captured_at"] for e in entries)
for entry in (e for e in entries if e["captured_at"] == this_run):
unchanged_current_year = entry["changed"] is False and entry["partial_filename"].startswith("orders_2026")
if unchanged_current_year or entry["stale"]:
raise Exception(f"{entry['partial_filename']}: changed={entry['changed']} stale={entry['stale']} new_history_rows={entry['new_history_rows']}")
bronze_table_folder is the folder the loader computed for the table: <Bronze abfs path>/Files/<bronzefolder>/<sourcetable> for fabricfiles, <Bronze abfs path>/Files/<bronzefolder>/<dataplatformobjectname> for azblob. dataplatformobjectname is the object's own name, which selects its entries from a tracker that may be shared with another object. Test on is False rather than falsiness, since null means "no previous delivery", not "unchanged".
A complete worked example is the Check_Tracker notebook in the DemoPlatform (Generator/Dataplatform/DP/Notebooks/Bronze/Check_Tracker): it treats files whose name matches Param001 as urgent and logs an ERROR when such a file is stale, re-delivered unchanged, or changed without new history rows, and a WARNING for the other files. Both go through the wheel's LogMessage with category="Validation" and log_type="VIOLATION", so they land in Meta.dbo.logging next to the loader's own validations and the load itself stays successful; a notebook that should fail the load instead exits with a string starting with error. Object vt_orders_c wires it up and variant scenario a19_postbronze_tracker_regels exercises both outcomes.
Function Reference
run
def run(tablefile: str, config_manager: ConfigManager = None,
overrides: LoadOverrides = None) -> str | None
Runs the bronze loader process for a specified table configuration and pulls files from the source, processes them, and loads them into the bronze layer.
Workflow:
- Validates table configuration and layer settings
- Computes stable Bronze folder path for file tracker
- Loads previous file snapshot from tracker (if exists)
- Optionally performs Azure Blob pre-check for unchanged sources (runs the freshness check before skipping, so a stale-but-unchanged source is still surfaced)
- Executes pre-bronze notebook (if configured)
- Pulls files from source system
- Validates file freshness (Validation 1) — under the default
stop, a single stale file aborts the run here, before any file is loaded - Checks for file changes using tracker snapshot (if
bronzeskipifsourceunchangedenabled) - Truncates Bronze table and loads file data by type
- Executes mid-bronze notebook and reloads data (if configured)
- Loads history records and validates per source file that a changed file brought new history rows (Validation 2)
- Saves file snapshot for next run, including each file's
stale,changedandnew_history_rows - Executes post-bronze notebook (if configured)
Arguments:
tablefilestr - Path to the YAMLA simple way to write configurations. It's basically a list that computers can read easily. file representing a table's configuration.config_managerConfigManager - An instance of ConfigManager used for accessing the application's configuration settings. Defaults to global config if not provided.overridesLoadOverrides - Optional per-run overrides that flip default validation/gating behaviour for this call only — skip notebook hooks, skip the stale-file check, or force a reload of unchanged source. Defaults to no overrides (use YAML / ConfigManager defaults).
Returns:
str- A message indicating the outcome, such as file count, skip reason, or error details. ReturnsNoneif table is inactive or skipped.
Raises:
Exception- If ConfigManager is not initialized, table config is invalid, or filetype is unsupported.
Configuration Options:
bronzeskipifsourceunchanged(bool) - Enable skip-if-unchanged detectionbronzeloadskip(bool) - Skip entire bronze load for this tablekeephistory(bool) - Maintain history records and validationprebronzenotebook(str) - Path to notebook to run before loadingmidbronzenotebook(str) - Path to notebook to run between load and historypostbronzenotebook(str) - Path to notebook to run after loadingsourceorder(bool) - Sort files by sourceorder instead of namebronzefolder(str) - Override Bronze folder from connectionbronzemaxfileagehours(int) - Maximum file age for freshness validation (per table, falling back to the connection)bronzemaxunchangedhours(int) - Longest stretch the source may stay unchanged before one full load runs anyway and the expiry is reported (per table, falling back to the connection; unset leaves skipping unbounded)bronzeunchangedseverity(str) - How loudly an expired unchanged-window is reported, on the same four rungs (default warning)bronzemaxfileagecriticalhours(int) - Age past which a stale file is reported ascriticalinstead of its base severity; escalation moves the severity, not the action (per table, falling back to the connection; unset never escalates)bronzeloadseverity(str) - How loudly a stale source file is reported:none,info,warning,critical(default warning)bronzechecksummismatchseverity(str) - How loudly a Validation 2 checksum/row-count mismatch is reported, on the same four rungs (default warning)
dataframeloader
def dataframeloader(data_frame: DataFrame, load_config: LoadConfig,
table_config: TableConfig,
config_manager: ConfigManager = None) -> str | None
Loads a DataFrame into a specified data platform table using the provided configuration and manager.
This function handles the loading operation by using detailed configurations for the DataFrame, table, and the application configuration manager. It sets up logging, ensures required parameters are initialized, and supports specific settings for different layers (e.g., bronze layer). The function handles exception logging and provides mechanisms to stop processing upon encountering errors based on configuration settings.
Arguments:
data_frameDataFrame - The data to be loaded into the specified table.load_configLoadConfig - Contains configuration for the loading process, including destination table.table_configTableConfig - Holds table-specific settings, e.g., table name identifiers and layers.config_managerConfigManager - Manages and validates application-level configurations.
Returns:
str- Message indicating the result of the DataFrame loading process, including the target table name and error details if applicable.
Raises:
Exception- If the destination table name is missing from LoadConfig.Exception- If the ConfigManager is not properly initialized.
Supported File Types
The bronze loader supports the following file formats:
- CSV - Comma-separated values with automatic schema detection
- JSON - JSON objects and arrays
- XML - Extensible Markup Language files
- Parquet - Apache Parquet columnar format
- XLSX - Excel spreadsheets
- Notebook - Databricks notebooks for custom processing
Error Handling
The module provides comprehensive error handling:
- Pre-check failures: Logs warnings and continues with regular load if pre-check fails
- File change detection failures: Proceeds with load if change detection fails
- Stale files: Behavior depends on
bronzeloadviolationaction. Under the defaultstop, the first stale file raisesFabricStaleFileErrorbefore any file loads, aborting the whole Bronze run for that table (no files written, including fresh ones). Undercontinue, old files are logged as warnings and processing continues. Undersilent, processing continues with no log entry - Checksum/no-new-rows mismatch: Evaluated per source file. Behavior depends on
bronzechecksummismatchviolationaction(defaultcontinue— logs a warning for each flagged file). Set it tostopto raiseFabricChecksumMismatchErroron the first flagged file instead, orsilentto suppress the warning entirely. This is a separate knob frombronzeloadviolationaction - Configuration errors: Raises exceptions for missing or invalid configurations
- Stop at error: Respects the
stop_at_errorconfig to halt on first error or continue with warnings
Example Usage
from easyfabric.load_data_bronze import run
from easyfabric.data import ConfigManager
# Initialize config manager
config_manager = ConfigManager.initialize()
# Run bronze loader for a specific table
result = run(
tablefile="/path/to/table_config.yaml",
config_manager=config_manager
)
print(result)
# Output: "5 files loaded for table: my_table" or "Bronze: my_table skipped (source unchanged)"
Best Practices
- Enable skip-if-unchanged for tables with infrequent source changes to save processing time and bandwidth
- Monitor validation warnings in logs to identify data quality issues early
- Configure bronzemaxfileagehours appropriately for your data refresh schedule
- Use keephistory for critical tables to maintain an audit trail and detect data anomalies
- Set sourceorder if file processing order matters for your use case
- Pre- and post-notebooks for custom transformations and cleanup before/after loading