❌

Reading view

Canonical Correlation Analysis as Reduced Rank Regression in High Dimensions

Canonical correlation analysis is a widespread technique for discovering linear relationships between two sets of variables. In high dimensions, however, standard estimates of the canonical directions cease to be consistent without assuming further structure. In this setting, a possible solution consists in leveraging the presumed sparsity of the solution: only a subset of the covariates span the canonical directions. While the last decade has seen a proliferation of sparse canonical correlation analysis methods, practical challenges regarding the scalability and adaptability of these methods still persist. To circumvent these issues, this paper suggests an alternative strategy that uses reduced rank regression to estimate the canonical directions when one of the datasets is high-dimensional while the other remains low-dimensional. By casting the problem of estimating the canonical direction as a regression problem, our estimator is able to leverage the rich statistics literature on high-dimensional regression and is easily adaptable to accommodate a wider range of structural priors. Our proposed solution maintains computational efficiency and accuracy, even in the presence of very high-dimensional data. We demonstrate the advantages of our approach through a series of simulated experiments, achieving a substantial increase in accuracy compared to existing methods. Additionally, we highlight its practical utility by applying it to three real-world datasets, where we show that our method is able to uncover scientifically meaningful associations between variables.
  •  

Sparse Topic Modeling via Spectral Decomposition and Thresholding

In probabilistic Latent Semantic Indexing (pLSI), word frequencies across document corpora are modeled through a low-rank factorization of the expected document-term matrix into topic-word and topic-document components. In this paper, we study the estimation of the topic-word matrix under a sparsity structure motivated by Zipf's law: word frequencies within each topic exhibit a rapid empirical decay, with most probability mass concentrated on a small subset of words. Motivated by this observation, we introduce a spectral estimator that adaptively thresholds rare words prior to factorization. We show that the resulting estimator achieves an $\ell_1$-error rate whose dependence on the vocabulary size $p$ is only logarithmic. Our error bounds hold across parameter regimes, including high-dimensional settings with extremely large vocabularies, a practically important scenario that has received limited theoretical attention. Unlike many existing methods, our approach does not require the separability (or anchor-word) assumption. Synthetic and real-data experiments demonstrate that the proposed procedure is computationally efficient, statistically reliable, and effective across domains with widely varying dimensions, sparsity levels, and document lengths.
  •  
❌