Aligning Multilingual News for Stock Return Prediction
Yuntao Wu, Lynn Tao, Ing-Haw Cheng, Charles Martineau, Yoshio Nozawa, John Hull, Andreas Veneris
TL;DR
The paper addresses how multilingual news can inform stock return prediction by aligning cross-language content at the sentence level using optimal transport. It embeds sentences with LaBSE, constructs a sparse, interpretable alignment map, and aggregates aligned, unaligned, and full-text embeddings to generate return signals. Return signals derived from aligned sentences show stronger correlations with realized returns, and long-short portfolios based on these signals achieve higher Sharpe ratios than those using full or unaligned text, with notable gains in the Japanese market. The approach scales to large multilingual corpora and offers a principled, interpretable mechanism to leverage cross-language information for finance, with future work aimed at broader markets and improved thresholding.
Abstract
News spreads rapidly across languages and regions, but translations may lose subtle nuances. We propose a method to align sentences in multilingual news articles using optimal transport, identifying semantically similar content across languages. We apply this method to align more than 140,000 pairs of Bloomberg English and Japanese news articles covering around 3500 stocks in Tokyo exchange over 2012-2024. Aligned sentences are sparser, more interpretable, and exhibit higher semantic similarity. Return scores constructed from aligned sentences show stronger correlations with realized stock returns, and long-short trading strategies based on these alignments achieve 10\% higher Sharpe ratios than analyzing the full text sample.
