Skip to content

Features

finml_core.processing.features

FeatureGenerator(col_map, ticker_level, config=None)

Orchestrates feature engineering pipeline from raw data.

This class handles
  1. Column Mapping (Standardizing input names).
  2. Feature Calculation (Calling metrics module).
  3. Lag Generation (Per-feature customization).
  4. Multi-asset grouping (Stacking support).

Supported Metrics & Configuration Keys:

  • log_ret:
    • period (int): Return period (default: 1).
  • rsi:
    • period (int): Lookback period. (Default: 14).
  • bollinger:
    • window (int): Moving average window (default: 20).
    • std (float): Standard deviations (default: 2.0).
  • macd:
    • fast (int): Fast EMA (default: 12).
    • slow (int): Slow EMA (default: 26).
    • signal (int): Signal EMA (default: 9).
  • rvol:
    • window (int): SMA window (default: 20).

Parameters:

Name Type Description Default
col_map Dict

Maps standard names to dataframe columns.

Example:

    {
        'open': 'Open',
        'high': 'High',
        'low': 'Low',
        'close': 'Close',
        'volume': 'Volume'
    }
which is yfinance standard.

required
ticker_level str

Name of the MultiIndex level containing the asset identifiers (e.g., 'Ticker').

required
config Dict

Configuration dictionary.

Defaults to:

    {
        'log_ret': {'period': 1},
        'rsi': {'period': 14},
        'bollinger': {'window': 20, 'std': 2.0},
        'macd': {'fast': 12, 'slow': 26, 'signal': 9},
        'rvol': {'window': 20}
    }
which is the market standard configuration.

None
Note

Your DataFrame must include 'Close' and 'Volume' columns to compute all supported metrics.


compute_indicators(df)

Phase 1: Feature Calculation Engine

This method orchestrates the vectorization of technical indicators across a MultiIndex DataFrame. It ensures that calculations are performed independently for each ticker to prevent look-ahead bias and cross-sectional data leakage.

The process follows a strict 3-step pipeline
  1. Validation: Verifies that the MultiIndex structure and required columns (mapped via col_map) exist before computation.
  2. Grouped Transformation: Uses groupby().transform() or apply() to isolate time-series calculations by asset.
  3. Feature Expansion: Dynamically appends new analytical columns to the dataset while preserving the original index.

Parameters:

Name Type Description Default
df DataFrame

Input DataFrame with a MultiIndex (Date, Ticker) and raw OHLCV columns.

required

Returns:

Type Description
DataFrame

A copy of the input DataFrame enriched with:

  • Momentum: Log Returns, RSI, RSI Change.
  • Volatility: Bollinger Bands (Middle, Upper, Lower, %B, Bandwidth).
  • Trend: MACD (Line, Signal, Histogram, Relative Histogram).
  • Volume: Relative Volume (RVol).

Raises:

Type Description
KeyError

If a configured indicator lacks its required raw column mapping (e.g., trying to calculate RSI without a 'close' column).

ValueError

If the ticker_level is not correctly identified in the Index.

Note

This method is "non-destructive" regarding rows; it calculates metrics for all available data. Handling of NaN values generated by rolling windows is deferred to Phase 2 (Filtering).


construct_feature_matrix(df_metrics, selection=None)

Phase 2: Selection & Lagging

Refines the calculated metrics into a finalized feature matrix. This phase handles the dimensionality of the input space by selecting specific columns and generating lagged observations to capture autocorrelation and temporal dependencies.

Process
  1. Feature Pruning: Filters the dataset to keep only the metrics specified in the selection dictionary.
  2. MultiIndex Lagging: Generates \(t-n\) observations using groupby().shift(). This ensures that lags are calculated per asset, preventing cross-contamination between different tickers.
  3. Feature Expansion: Creates new columns with the suffix _lag_n.

Machine Learning Recommendations

For most ML models (Random Forest, XGBoost, etc.), it is highly recommended to use Stationary Features. Using raw prices or moving averages (like bb_middle) can lead to poor generalization due to unit roots.

Recommended Stationary Set (Default):

  • Returns: log_ret (captures percentage change).
  • Momentum: rsi, rsi_diff (bounded between 0-100).
  • Volatility: bb_pct_b, bb_width (normalized volatility).
  • Trend: macd_rel_hist (price-normalized momentum).
  • Volume: rvol (normalized activity).

Parameters:

Name Type Description Default
df_metrics DataFrame

Enriched MultiIndex DataFrame from Phase 1 (generate() function).

required
selection Dict[str, List[int]]

Dictionary mapping feature names to a list of desired lags.

  • Use [] for the current value (\(t\)) only.
  • Use [1, 2] for current value plus two previous steps.

Defaults to:

{
    'log_ret': [],
    'rsi': [],
    'rsi_diff': [],
    'bb_pct_b': [],
    'bb_width': [],
    'macd_rel_hist': [],
    'rvol': []
}
None

Returns:

Type Description
DataFrame

The finalized Feature Matrix (X). Columns are ordered by feature then by lag (e.g., rsi, rsi_lag_1, rsi_lag_2).

Notes
  • Temporal Dynamics: Adding lags is essential for non-sequential models (like Random Forest or XGBoost) to "see" the trend.
  • Data Leakage: This method preserves the MultiIndex to ensure that shift operations never mix data from different tickers at the boundaries.