Features
finml_core.processing.features
FeatureGenerator(col_map, ticker_level, config=None)
Orchestrates feature engineering pipeline from raw data.
This class handles
- Column Mapping (Standardizing input names).
- Feature Calculation (Calling metrics module).
- Lag Generation (Per-feature customization).
- Multi-asset grouping (Stacking support).
Supported Metrics & Configuration Keys:
log_ret:- period (int): Return period (default: 1).
rsi:- period (int): Lookback period. (Default: 14).
bollinger:- window (int): Moving average window (default: 20).
- std (float): Standard deviations (default: 2.0).
macd:- fast (int): Fast EMA (default: 12).
- slow (int): Slow EMA (default: 26).
- signal (int): Signal EMA (default: 9).
rvol:- window (int): SMA window (default: 20).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
col_map
|
Dict
|
Maps standard names to dataframe columns. Example: |
required |
ticker_level
|
str
|
Name of the MultiIndex level containing the asset identifiers (e.g., 'Ticker'). |
required |
config
|
Dict
|
Configuration dictionary. Defaults to: |
None
|
Note
Your DataFrame must include 'Close' and 'Volume' columns to compute all supported metrics.
compute_indicators(df)
Phase 1: Feature Calculation Engine
This method orchestrates the vectorization of technical indicators across a MultiIndex DataFrame. It ensures that calculations are performed independently for each ticker to prevent look-ahead bias and cross-sectional data leakage.
The process follows a strict 3-step pipeline
- Validation: Verifies that the MultiIndex structure and required
columns (mapped via
col_map) exist before computation. - Grouped Transformation: Uses
groupby().transform()orapply()to isolate time-series calculations by asset. - Feature Expansion: Dynamically appends new analytical columns to the dataset while preserving the original index.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df
|
DataFrame
|
Input DataFrame with a MultiIndex (Date, Ticker) and raw OHLCV columns. |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
A copy of the input DataFrame enriched with:
|
Raises:
| Type | Description |
|---|---|
KeyError
|
If a configured indicator lacks its required raw column mapping (e.g., trying to calculate RSI without a 'close' column). |
ValueError
|
If the |
Note
This method is "non-destructive" regarding rows; it calculates metrics
for all available data. Handling of NaN values generated by rolling
windows is deferred to Phase 2 (Filtering).
construct_feature_matrix(df_metrics, selection=None)
Phase 2: Selection & Lagging
Refines the calculated metrics into a finalized feature matrix. This phase handles the dimensionality of the input space by selecting specific columns and generating lagged observations to capture autocorrelation and temporal dependencies.
Process
- Feature Pruning: Filters the dataset to keep only the metrics
specified in the
selectiondictionary. - MultiIndex Lagging: Generates \(t-n\) observations using
groupby().shift(). This ensures that lags are calculated per asset, preventing cross-contamination between different tickers. - Feature Expansion: Creates new columns with the
suffix
_lag_n.
Machine Learning Recommendations
For most ML models (Random Forest, XGBoost, etc.), it is highly recommended to
use Stationary Features. Using raw prices or moving averages (like bb_middle)
can lead to poor generalization due to unit roots.
Recommended Stationary Set (Default):
- Returns:
log_ret(captures percentage change). - Momentum:
rsi,rsi_diff(bounded between 0-100). - Volatility:
bb_pct_b,bb_width(normalized volatility). - Trend:
macd_rel_hist(price-normalized momentum). - Volume:
rvol(normalized activity).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
df_metrics
|
DataFrame
|
Enriched MultiIndex DataFrame
from Phase 1 ( |
required |
selection
|
Dict[str, List[int]]
|
Dictionary mapping feature names to a list of desired lags.
Defaults to: |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
The finalized Feature Matrix (X). Columns are ordered by feature
then by lag (e.g., |
Notes
- Temporal Dynamics: Adding lags is essential for non-sequential models (like Random Forest or XGBoost) to "see" the trend.
- Data Leakage: This method preserves the MultiIndex to ensure
that
shiftoperations never mix data from different tickers at the boundaries.