Split
finml_core.model_selection.split
PurgedKFold(n_splits, t1, date_level, pct_embargo=0.01)
Cross-Validation with Purging and Embargo for Financial Time Series.
Implements the methodology proposed by Marcos Lรณpez de Prado in "Advances in Financial Machine Learning" (Chapter 7). Standard K-Fold CV assumes IID data, which is false in finance due to overlapping labels and serial correlation.
This class prevents data leakage through two mechanisms
- Purging: Removes training observations whose labels overlap in time with the test set.
- Embargo: Eliminates a period immediately following the test set to handle auto-correlated residuals.
Mathematical Overlap Condition
A training observation is purged if its interval \([t_{i,0}, t_{i,1}]\) overlaps with the test interval \([T_{j,0}, T_{j,1}]\): $$ (t_{i,0} \le T_{j,1}) \land (t_{i,1} \ge T_{j,0}) $$
Attributes:
| Name | Type | Description |
|---|---|---|
n_splits |
int
|
Number of folds. |
t1 |
Series
|
End timestamps of labels (Vertical Barriers). |
date_level |
str
|
MultiIndex level name for timestamps. |
pct_embargo |
float
|
Percentage of total timeframe for the embargo. |
Note
Sklearn Compatibility: Works as a drop-in replacement for KFold in GridSearchCV and cross_val_score.
split(X, y, groups=None)
Generates indices to split data into training and test sets.
Process
- Temporal Grouping: Groups data by unique dates to maintain chronological integrity.
- Interval Analysis: Evaluates the
t1(label span) of each observation. - Purge & Embargo: Drops training indices that fail the non-overlap condition against the current test fold.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
DataFrame
|
Features dataset with MultiIndex. |
required |
y
|
Series
|
Target variable. Defaults to None. |
required |
groups
|
Compatibility placeholder. |
None
|
Yields:
| Type | Description |
|---|---|
ndarray
|
Integer indices for the training set. |
ndarray
|
Integer indices for the test set. |
purged_train_test_split(X, y, t1, date_level, test_size=0.2, pct_embargo=0.01)
Chronological split with financial purging and embargo.
Performs a non-shuffled temporal split (Past -> Train, Future -> Test) and applies a quarantine period to ensure the test set is strictly independent of the training data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
DataFrame
|
Features dataset. |
required |
y
|
Series
|
Target labels. |
required |
t1
|
Series
|
Timestamps of label ends (Vertical Barriers). |
required |
date_level
|
str
|
MultiIndex level name for timestamps. |
required |
test_size
|
float
|
Proportion of data for testing. Defaults to 0.2. |
0.2
|
pct_embargo
|
float
|
Total timeline percentage for embargo. Defaults to 0.01 (1%). |
0.01
|
Returns:
| Type | Description |
|---|---|
Tuple
|
|