What Is the Triple Barrier Method?
In February 2018, Marcos López de Prado published Advances in Financial Machine Learning. Chapter 3 set out a path-dependent way to label trading events using three exits: a profit-taking barrier, a stop-loss barrier and a maximum holding time. That is the triple barrier method.
The method does not discover alpha and it does not make labels objectively better. It answers a narrower question: given an event at , a trading side and a pre-specified exit policy, which outcome occurred first? That is useful when a fixed-horizon return does not match the trade you would actually execute. A five-day label, for example, ignores a stop that would have closed the position on day two.
The canonical source is López de Prado (2018), chapter 3. Hudson & Thames also published the book-aligned labelling source code, which is useful for checking details that abbreviated explanations often lose. For the modelling context, start with our machine learning for trading guide.
How the Three Barriers Work
Let be the entry price, a target measured using information available at the event time, and the profit-taking and stop-loss multipliers. For a long event, the horizontal return barriers are
The vertical barrier is a timestamp . We follow the path of the signed return
until it first reaches , reaches , or arrives at . Multiplying by the side means that a favourable move is positive for both longs and shorts.
For ordinary multi-class labelling, the final label is usually , where is the first-touch time. That detail matters: an event reaching the vertical barrier is not automatically class zero. If its terminal return is positive or negative, the standard binning step can assign or . A neutral class requires an explicit convention, such as a minimum-return band around zero.
Volatility scaling
Scaling the horizontal barriers by an ex-ante volatility target makes a 2% move harder to label as exceptional in a turbulent market than in a quiet one. It does not remove heteroskedasticity by magic. The estimator, lookback and lag are all modelling choices. A volatility estimate that includes returns after leaks the answer into the label.
Use a lagged EWMA, realised measure or forecast that could genuinely have been computed at the decision time. Keep the units consistent: a one-day volatility estimate does not become a five-day target unless you deliberately rescale it.
Event sampling and meta-labelling
Labelling every bar creates many near-duplicate events. López de Prado's workflow first samples event timestamps, often with a symmetric CUSUM filter, and drops events whose target is below a minimum return. Neither choice is part of the mathematical definition of the barriers, but both materially affect sample dependence and class balance.
With meta-labelling, a primary model supplies . The secondary model does not learn direction again; it learns whether to take the primary bet. After returns are multiplied by the proposed side, profitable outcomes receive label one and non-profitable outcomes receive zero. This separates side from size or participation. It does not guarantee improved precision, and the threshold still has to be selected out of sample.
For fixed-horizon forecasting, meanwhile, an ordinary forward return may be exactly the right target. Choose the label from the economic decision, not because one scheme sounds more advanced.
Worked Example and Python
Suppose an event begins at £100 with side , target volatility 1.5%, , and a five-session cap. The barriers are +3% and -1.5%, or £103 and £98.50. If closes are £100.40, £99.10, £98.40, £101 and £103.20, the stop is first touched on the third close. The label is -1 even though the fifth close is above the profit target. Path order is the point.
This implementation accepts one event per row. Each row contains a vertical-barrier timestamp in , a strictly positive target and, optionally, a side. It uses close-to-close returns, so it only claims to know about touches visible at the sampling frequency.
import numpy as np import pandas as pd def triple_barrier_labels( close: pd.Series, events: pd.DataFrame, pt_sl: tuple[float, float] = (2.0, 2.0), ) -> pd.DataFrame: rows = [] for t0, event in events.sort_index().iterrows(): path = close.loc[t0:event["t1"]] / close.at[t0] - 1.0 side = float(event.get("side", 1.0)) signed = side * path pt = signed[signed >= pt_sl[0] * event["target"]].index sl = signed[signed <= -pt_sl[1] * event["target"]].index candidates = { "pt": pt[0] if len(pt) else pd.NaT, "sl": sl[0] if len(sl) else pd.NaT, "vertical": path.index[-1], } barrier, touch = min( ((name, time) for name, time in candidates.items() if pd.notna(time)), key=lambda item: item[1], ) realised = float(signed.loc[touch]) label = int(realised > 0) if event.get("side") is not None else int(np.sign(realised)) rows.append((t0, touch, barrier, realised, label)) return pd.DataFrame( rows, columns=["t0", "touch", "barrier", "return", "label"] ).set_index("t0")
For the numbers above, barrier is sl, return is -0.016 and label is -1. If a side column is supplied for meta-labelling, the same loss becomes class zero.
Production code needs more guards than fit comfortably in an article: verify monotonic unique timestamps, reject missing or non-positive targets, decide what happens beyond the last available price and preserve transaction-cost assumptions alongside each event. The reference implementation also parallelises first-touch searches.
There is no honest way to infer within-bar order from OHLC data when both high and low cross their barriers in the same bar. Do not silently choose the stop, the target or whichever gives the prettier result. Mark the event ambiguous, use finer data or impose a documented conservative convention.
Label diagnostics before modelling
Treat the labelled event table as a dataset in its own right. Tabulate first-touch frequencies by year, asset and volatility decile. A target that is hit on 80% of low-volatility events but 5% of high-volatility events is not achieving regime comparability. Plot holding times and realised overshoots too. Large stop overshoots may reveal gaps or illiquid closes that a threshold-only backtest cannot trade.
Measure concurrency, the number of active event intervals at each timestamp. Ten labels can all depend on the same market move; they are not ten independent pieces of evidence. Sample weights based on uniqueness are one proposed response in López de Prado's framework, but weighted observations still do not create new information. Report the raw event count, average uniqueness and the number of non-overlapping time blocks.
Class balance needs an economic reading. If almost every event reaches the vertical barrier, the horizontal levels may be irrelevant at that horizon. If almost every event hits a horizontal barrier on the next bar, the target may be below the bid-ask spread or normal bar range. Neither outcome is fixed by resampling the minority class. Revisit the exit policy first.
Finally, freeze label generation before fitting a serious model. Save the event timestamp, target forecast, side, all three barrier values, first-touch timestamp, available fill and code version. That audit trail makes it possible to distinguish a changed signal from a changed definition of success.
Where Triple Barrier Labelling Goes Wrong
The common failure is not the loop above. It is allowing the research design to move information backwards through time.
Overlapping labels. Events at adjacent timestamps can share most of their return path. A random split may then place heavily overlapping event intervals on opposite sides of the train-test boundary. Use contiguous folds and purge any training interval that intersects a test interval. An embargo can add a further buffer. These controls address interval overlap; they do not repair features calculated with future data.
Post-selection barriers. Trying 50 combinations of targets and holding periods, then reporting the winner's cross-validated score, makes the label definition part of the hyperparameter search. Select it inside a nested validation loop or justify it from execution economics before modelling. The same warning applies to event filters and probability cut-offs. Our Sharpe ratio of pure noise shows why a large search can manufacture a convincing winner.
Execution mismatch. Close-only labels are inappropriate if the real order triggers intraday. Conversely, high-low labels may overstate fill quality because touching a quoted price does not guarantee an executable fill. Include bid-ask spread, slippage, gaps and the next tradable timestamp. A stop crossed on an overnight gap is filled at the available price, not at the barrier.
Changing universes. Survivorship-biased prices and revised corporate-action data contaminate labels even when the barrier code is perfect. Use point-in-time constituents and adjustments. Audit delistings, stale marks, limit moves and missing bars.
Finally, inspect the output. Barrier frequencies should be stable enough to explain, labels should not be dominated by one class and the average concurrency of events should be measured. If a tiny change in volatility lookback reverses the result, the model is learning the labelling policy rather than a durable market relation. Pair these checks with the tests in our mean-reversion guide where the underlying signal is a spread.
Frequently Asked Questions
What is the triple barrier method in one sentence?
A path-dependent event-labelling method that stops at the first profit-taking, stop-loss or time barrier and assigns a label from the signed realised return.
Where was the triple barrier method introduced?
In chapter 3 of López de Prado's Advances in Financial Machine Learning (2018).
Is a vertical-barrier outcome always label zero?
No. In the standard binning step, the return at the vertical barrier can be positive or negative and its sign becomes the label. Zero is a separate convention, not an automatic timeout class.
Why purge triple barrier labels in cross-validation?
Because two event labels may depend on the same future price path. Purging removes training events whose information intervals overlap a test interval; ordinary shuffled K-fold does not.
What is meta-labelling?
A primary model chooses long or short. A secondary classifier estimates whether that proposed side should be taken, usually with zero for a non-profitable signed outcome and one for a profitable one.
Can both horizontal barriers be touched in one bar?
Yes, if OHLC data show a high above the target and a low below the stop. The bar does not reveal which came first. Use finer observations or flag the event rather than inventing an ordering.
Skip the £25k programme - try the alternative
Master's programmes are slow and expensive. Quantt is a self-paced alternative. Start free with a real lesson and interview practice, then unlock 50+ courses and your personalised plan.
Free to start · No credit card required