Data Factory
finml_core.pipelines.data_factory
DatasetGenerator(data_source=None, custom_provider=None, feature_config=None, feature_selection=None, label_config=None)
Master Orchestrator for Financial Machine Learning Pipelines.
This class centralizes the end-to-end workflow of a quantitative trading strategy
- Data Acquisition: Either via automated download or manual injection.
- Cleaning: Handles MultiIndex alignment, NaNs, and infinite values.
- Feature Engineering: Generates technical indicators and temporal lags.
- Labeling: Executes the Triple Barrier Method for target definition.
- Consolidation: Aligns \(X\) and \(y\) by removing inconsistent 'warm-up' periods.
Operation Modes
The generator adapts to two primary user workflows:
- Automated Mode (Beginner-Friendly): Uses pre-configured mappings from the
internal registry. You only need to provide a
data_source(e.g., 'yfinance') and the list of tickers in therunmethod. - Custom Mode (Pro/Injected): Designed for users with proprietary datasets
or unsupported APIs. By providing a
custom_providerdictionary, the orchestrator switches to an injection-only logic, requiring a pre-loaded DataFrame in therunmethod.
Attributes:
| Name | Type | Description |
|---|---|---|
analysis_data |
DataFrame
|
Full dataset available after |
Pipeline Configuration & Engine Setup
Instantiating this class initializes the core modular engines required for
the full data lifecycle. This constructor acts as a factory, setting up
the internal instances of MarketDataLoader, FeatureGenerator,
TripleBarrierLabeling, and DataCleaner based on the selected mode.
It determines whether the pipeline will operate in 'Automated' mode (leveraging predefined provider registries) or 'Custom' mode (handling user-injected datasets and coordinate maps).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data_source
|
str
|
Name of the provider in
|
None
|
custom_provider
|
Dict[str, Any]
|
Custom coordinate map. If provided, ignores Example: |
None
|
feature_config
|
Dict[str, Dict[str, Any]]
|
Technical indicator parameters (e.g., |
None
|
feature_selection
|
Dict[str, List[int]]
|
The final 'filter' for the matrix. Defines which metrics and
how many lags to include in \(X\). If None, uses
|
None
|
label_config
|
Dict[str, Any]
|
Hyperparameters for the Triple Barrier Method. If None, uses
|
None
|
run(tickers=None, etf_reference=None, start_date=None, end_date=None, df_input=None)
Executes the modeling pipeline to produce model-ready matrices.
This method orchestrates the full data lifecycle. It follows a strict Priority Logic:
- If
custom_providerwas used at__init__, it requiresdf_input. - If
data_sourcewas used, it attempts to download data usingtickers,etf_reference, andstart_date.
Data Structure Requirement (MultiIndex)
Whether injected via df_input or downloaded, the internal engine
requires a MultiIndex DataFrame with levels corresponding to
ticker_level and date_level.
Standard format: [Date (DatetimeIndex), Ticker (str)].
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tickers
|
List[str]
|
Asset symbols to fetch. Required if |
None
|
etf_reference
|
str
|
Reference symbol (e.g., '^GSPC') to align market holidays. |
None
|
start_date
|
str
|
Start of the historical period ( |
None
|
end_date
|
str
|
End of the period. Defaults to current date. |
None
|
df_input
|
DataFrame
|
Pre-loaded MultiIndex DataFrame. |
None
|
Returns:
| Type | Description |
|---|---|
Tuple
|
|
Index Alignment & Trimming:
The pipeline performs a bidirectional trim to ensure a clean, MultiIndex-aligned dataset. It removes the Warm-up Period at the beginning (caused by rolling metrics like volatility and indicators) and the Horizon Gap at the end (due to the forward-looking nature of triple-barrier labels). This ensures that \(X\) and \(y\) contain only fully realized, non-null observations ready for machine learning.