Skip to content

Cleaning

finml_core.processing.cleaning

DataCleaner(ticker_level, date_level, method='ffill')

Module responsible for data sanitization and stability.

Standardizes the handling of infinities (infs) and null values (NaNs) to ensure the numerical stability of the pipeline before ML ingestion.

Key Features
  1. Multi-Asset Safety: Applies operations per-ticker to prevent data leakage.
  2. Look-ahead Bias Prevention: Uses conservative filling methods (ffill limit=1).
  3. Strict Validation: Raises errors if data quality standards are not met.

Parameters:

Name Type Description Default
ticker_level str

Name of the MultiIndex level containing the asset identifiers (e.g., 'Ticker').

required
date_level str

Name of the MultiIndex level containing the timestamps (e.g., 'Date'). Required for re-indexing during edge trimming.

required
method str

Strategy for handling NaNs. for now only 'ffill' (Forward Fill) is supported as it is standard in in financial time series to handle minor gaps (e.g., holidays) without look-ahead bias.

'ffill'

validate_and_clean(df)

Execution Protocol: Strict Sanitization.

Runs the data through a 4-stage quality gate to ensure numerical integrity.

Protocol Steps
  1. Trim Edges: Removes leading/trailing NaNs (warm-up periods).
  2. Fill Gaps: Applies limited forward fill for internal gaps.
  3. NaN Validation: checks for remaining Nulls; raises Error if found.
  4. Inf Validation: checks for Infinite values; raises Error if found.

Parameters:

Name Type Description Default
df DataFrame

Input DataFrame (Raw or Feature Matrix). Must have a MultiIndex (Date, Ticker).

required

Returns:

Type Description
DataFrame

A clean, numerically stable DataFrame ready for modeling.

Raises:

Type Description
ValueError

If NaNs persist after cleaning or if Infs are detected.