Skip to content

Split

finml_core.model_selection.split

PurgedKFold(n_splits, t1, date_level, pct_embargo=0.01)

Cross-Validation with Purging and Embargo for Financial Time Series.

Implements the methodology proposed by Marcos Lรณpez de Prado in "Advances in Financial Machine Learning" (Chapter 7). Standard K-Fold CV assumes IID data, which is false in finance due to overlapping labels and serial correlation.

This class prevents data leakage through two mechanisms
  1. Purging: Removes training observations whose labels overlap in time with the test set.
  2. Embargo: Eliminates a period immediately following the test set to handle auto-correlated residuals.
Mathematical Overlap Condition

A training observation is purged if its interval \([t_{i,0}, t_{i,1}]\) overlaps with the test interval \([T_{j,0}, T_{j,1}]\): $$ (t_{i,0} \le T_{j,1}) \land (t_{i,1} \ge T_{j,0}) $$

Attributes:

Name Type Description
n_splits int

Number of folds.

t1 Series

End timestamps of labels (Vertical Barriers).

date_level str

MultiIndex level name for timestamps.

pct_embargo float

Percentage of total timeframe for the embargo.

Note

Sklearn Compatibility: Works as a drop-in replacement for KFold in GridSearchCV and cross_val_score.


split(X, y, groups=None)

Generates indices to split data into training and test sets.

Process
  1. Temporal Grouping: Groups data by unique dates to maintain chronological integrity.
  2. Interval Analysis: Evaluates the t1 (label span) of each observation.
  3. Purge & Embargo: Drops training indices that fail the non-overlap condition against the current test fold.

Parameters:

Name Type Description Default
X DataFrame

Features dataset with MultiIndex.

required
y Series

Target variable. Defaults to None.

required
groups

Compatibility placeholder.

None

Yields:

Type Description
ndarray

Integer indices for the training set.

ndarray

Integer indices for the test set.

purged_train_test_split(X, y, t1, date_level, test_size=0.2, pct_embargo=0.01)

Chronological split with financial purging and embargo.

Performs a non-shuffled temporal split (Past -> Train, Future -> Test) and applies a quarantine period to ensure the test set is strictly independent of the training data.

Parameters:

Name Type Description Default
X DataFrame

Features dataset.

required
y Series

Target labels.

required
t1 Series

Timestamps of label ends (Vertical Barriers).

required
date_level str

MultiIndex level name for timestamps.

required
test_size float

Proportion of data for testing. Defaults to 0.2.

0.2
pct_embargo float

Total timeline percentage for embargo. Defaults to 0.01 (1%).

0.01

Returns:

Type Description
Tuple
  • X_train (pd.DataFrame): Purged features.
  • y_train (pd.Series): Purged labels.
  • X_test (pd.DataFrame): Test features.
  • y_test (pd.Series): Test labels.