Harsh Lessons I Learned About Pairs Trad ...

Harsh Lessons I Learned About Pairs Trading

Feb 20, 2025

A lot of folks talk about trying to profit from pairs trading. The idea is to look at two assets that are nearly identical to one another, and track the discrepancy between the two. When that discrepancy gets big, you bet on it eventually getting smaller. 

In 90% of cases, pairs trading strategies discussed on social media don't make sense. We'll look at an example that makes sense, an example that is nonsense, and some common issues when testing these strategies.

Example of a Plausible Pairs Trading Scenario (Stolen from @TheRobotJames)

One example might involve ADRs (American depositary receipts). For example, some German companies have stocks that trade on the Frankfurt Exchange but not on American stock exchanges - they are not dual listed. Sometimes, American investors can still own these stocks despite this; they can buy an ADR, a security that trades on an American exchange and that represents a share of a foreign stock. 

In theory, the price of the ADR should be pretty much identical to the price of the stock on the foreign exchange (after adjusting for differences in currency). 

In reality, there will be a discrepancy for two main reasons. 

  1. Someone trying to profit from the discrepancy will encounter trading frictions.

  2. There are different market participants trading the ADR, as opposed to the foreign stock, and often one of those is more liquid than the other. 

    • Sometimes, those participants (especially in the less liquid product) overreact to news, get caught in some kind of trading mistake, or have some other issue that makes them temporarily rush to buy or sell. 

    • This "emotional/stupid" behavior causes exploitable price distortions.

If you carefully monitor both markets and have a basic model that gives you a sense of when the price discrepancy is out of whack (hopefully due to stupid behavior), you can make money by stepping in and betting on the discrepancy closing. 

You make money because there is an economic explanation for why there is a temporary inefficiency/discrepancy, you are probably involved in a less liquid market, and you are taking some risk to try to bet on it in some way. You're getting paid for taking a logical risk.

Example of a Nonsensical Pairs Trading Scenario

Our previous, realistic example involved trading in less liquid markets and betting on an explainable discrepancy closing.

Sadly, a lot of pairs trading scenarios people discuss are super unrealistic. There are two main flavors.

Flavor 1 - Looking at Two Very Liquid Products and Looking at Super Granular Time Increments

This is a common one, and it often involves looking at minutely data because there is no obvious discrepancy when viewing longer increments (like 4 - 24 hours). Hint: this is a red flag already.

For example, a databento blog post discussed trading the discrepancy between West Texas Intermediate (WTI) crude oil futures and Brent crude oil futures. Both of these products are highly liquid. They represent nearly identical commodities (aside from their transit routes). 

There is no compelling explanation for why there would be an exploitable discrepancy between them.

In our ADR example, we mentioned that in the less liquid security, you can see price distortions from stupid behavior. But in a very liquid market, there are many participants and a ton of trading volume. A single participant's emotional/stupid behavior is not going to move the needle and cause an exploitable discrepancy. Basically, there is no logical reason to expect any exploitable discrepancies. 

The lack of exploitable opportunities will be obvious when you run a simple test using daily OHLC price data - one that assumes that you can only transact around the closing price, after adding a transaction cost and some slippage, and after considering a signal lag. One way to lag the signal is to check the price discrepancy at the open, and use it to decide if you need to make a trade around the close. 

When that test fails, the amateur mistake is to try again but with very granular data, like 1-minute time increments. These granular OHLC bars are misleading (see the RobotJames tweet here), because folks often run tests assuming that you can basically get in at the open of the next bar. The reality can be much, much worse than that.

Once you run some sensitivity tests, you'll see that the margin for error when testing a strategy on 1-minute time increments is extremely low. If you underestimate slippage by just a little bit, it could lead to wildly optimistic results when you look over many trades. It's a recipe for disaster. 

But the bigger issue is that there is no logical reason to expect any trading edge here in the first place!

Flavor 2 - Looking at Two Fairly Different Assets

Another common example involves two related stocks, such as two different regional banks. However, these assets are not nearly identical. For example, one regional bank may represent a completely different part of the US than another one. It is totally possible that one region of the US saw worse economic growth than another one, causing one regional bank to have worse financial results and causing its stock to have worse performance.

This would cause a discrepancy in the stock prices of the two regional banks, but it's not one that is inefficient or likely to revert. It occurred because these stocks represent equity in different companies, which can have different financial results and different business risks. 

Again, we have no real reason to expect we can identify price discrepancies that are inefficient and worth betting against. 

Common Mistakes in Testing

A lot of articles justify their selection of assets to pair trade based on the fact that they are cointegrated. There are two problems with this. 

Firstly, this ignores the questions that really matter: "Are these assets nearly identical? Is one of them less liquid than the other and likely to sometimes have inefficient price moves?" 

Just because something looks like it moves similarly to something else doesn't mean we should use both of them for pairs trading. 

Secondly, cointegration should not be the focus. It shows you whether two things move similarly in the long term. This can be misleading because in bull markets, a lot of stocks go up at the same time. And when the market has a correction, a lot of stocks drop at the same time. They will be cointegrated. In fact, the charts are so obviously cointegrated you don't even need to bother with the fancy calculation. 

There are plenty of pairs of assets that are cointegrated even though they aren't suitable for pairs trading.

We should first start with pairs of assets that we can logically expect to be nearly identical, but sometimes have inefficient discrepancies. Then, to test suitability for pairs trading, we shouldn't focus on cointegration.

Instead, we should focus on the important question - how strong is the relationship between our "discrepancy metric", and the returns in the following XYZ days / time increments (on a relative basis). A sensible choice for a discrepancy metric is to track the rolling Z-score of the discrepancy between the two assets. We then look at the correlation (Pearson correlation coefficient) between that Z score, and returns in the following XYZ days / time increments.

Keep in mind, when we measure the discrepancy between two assets, we may need to normalize for volatility. If for some reason, asset A is twice as volatile as asset B, we need to adjust asset A's movements before comparing them to movements of asset B, so we don't get fooled by noise. Similarly, when we compare returns in the following XYZ days, we may need to adjust for differing volatilities to ensure an apples-to-apples comparison. 

We might want to also see if that relationship is somewhat stable, so we may want to check if the correlation in the first half of the sample is similar to the second half.

There are other challenges involved in testing this, but this covers the basics. 

Summing Up

I think about 90% of pairs trading examples that I see on social media are nonsensical and unlikely to make money. Be careful about this stuff and take some time to check the logic. 

Ti piace questo post?

Offri un caffè a Entropy Chase

Altro da Entropy Chase