Skip to content

Data Factory

finml_core.pipelines.data_factory

DatasetGenerator(data_source=None, custom_provider=None, feature_config=None, feature_selection=None, label_config=None)

Master Orchestrator for Financial Machine Learning Pipelines.

This class centralizes the end-to-end workflow of a quantitative trading strategy
  1. Data Acquisition: Either via automated download or manual injection.
  2. Cleaning: Handles MultiIndex alignment, NaNs, and infinite values.
  3. Feature Engineering: Generates technical indicators and temporal lags.
  4. Labeling: Executes the Triple Barrier Method for target definition.
  5. Consolidation: Aligns \(X\) and \(y\) by removing inconsistent 'warm-up' periods.

Operation Modes

The generator adapts to two primary user workflows:

  • Automated Mode (Beginner-Friendly): Uses pre-configured mappings from the internal registry. You only need to provide a data_source (e.g., 'yfinance') and the list of tickers in the run method.
  • Custom Mode (Pro/Injected): Designed for users with proprietary datasets or unsupported APIs. By providing a custom_provider dictionary, the orchestrator switches to an injection-only logic, requiring a pre-loaded DataFrame in the run method.

Attributes:

Name Type Description
analysis_data DataFrame

Full dataset available after run(), containing original OHLCV, all indicators, and labeling details.

Pipeline Configuration & Engine Setup

Instantiating this class initializes the core modular engines required for the full data lifecycle. This constructor acts as a factory, setting up the internal instances of MarketDataLoader, FeatureGenerator, TripleBarrierLabeling, and DataCleaner based on the selected mode.

It determines whether the pipeline will operate in 'Automated' mode (leveraging predefined provider registries) or 'Custom' mode (handling user-injected datasets and coordinate maps).

Parameters:

Name Type Description Default
data_source str

Name of the provider in PROVIDERS registry. Defaults to 'yfinance'.

None
custom_provider Dict[str, Any]

Custom coordinate map. If provided, ignores data_source. Must contain col_map, ticker_level, and date_level.

Example:

{
    'col_map': {
        'open': 'Open', 'high': 'High', 'low': 'Low',
        'close': 'Close', 'volume': 'Volume'
    },
    'ticker_level': 'Ticker',
    'date_level': 'Date'
}
which is 'yfinance' mapping.

None
feature_config Dict[str, Dict[str, Any]]

Technical indicator parameters (e.g., rsi, bollinger). If None, uses compute_indicators defaults from FeatureGenerator class.

None
feature_selection Dict[str, List[int]]

The final 'filter' for the matrix. Defines which metrics and how many lags to include in \(X\). If None, uses construct_feature_matrix defaults from FeatureGenerator class.

None
label_config Dict[str, Any]

Hyperparameters for the Triple Barrier Method. If None, uses compute_outcomes defaults from TripleBarrierLabeling class.

None

run(tickers=None, etf_reference=None, start_date=None, end_date=None, df_input=None)

Executes the modeling pipeline to produce model-ready matrices.

This method orchestrates the full data lifecycle. It follows a strict Priority Logic:

  1. If custom_provider was used at __init__, it requires df_input.
  2. If data_source was used, it attempts to download data using tickers, etf_reference, and start_date.

Data Structure Requirement (MultiIndex)

Whether injected via df_input or downloaded, the internal engine requires a MultiIndex DataFrame with levels corresponding to ticker_level and date_level.

Standard format: [Date (DatetimeIndex), Ticker (str)].

Parameters:

Name Type Description Default
tickers List[str]

Asset symbols to fetch. Required if df_input is None and custom_provider was specified.

None
etf_reference str

Reference symbol (e.g., '^GSPC') to align market holidays.

None
start_date str

Start of the historical period (YYYY-MM-DD hh:mm:ss).

None
end_date str

End of the period. Defaults to current date.

None
df_input DataFrame

Pre-loaded MultiIndex DataFrame.

None

Returns:

Type Description
Tuple
  • X (pd.DataFrame): Final Feature Matrix with selected metrics and lags.
  • y (pd.DataFrame): Target labels (target_side) aligned with \(X\).

Index Alignment & Trimming:

The pipeline performs a bidirectional trim to ensure a clean, MultiIndex-aligned dataset. It removes the Warm-up Period at the beginning (caused by rolling metrics like volatility and indicators) and the Horizon Gap at the end (due to the forward-looking nature of triple-barrier labels). This ensures that \(X\) and \(y\) contain only fully realized, non-null observations ready for machine learning.