arXiv 2026 Preprint

Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices

Fateme Mazdarani, Carlos Toxtli-Hernández

“Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices.” arXiv preprint arXiv:2609.19243 (2026), 2026

Abstract

Spectral co-clustering is a useful tool for discovering latent structure in word-document matrices, but its reliance on singular value decomposition (SVD) can make standard formulations expensive on high-dimensional data. This paper presents two randomized approximations for normalized spectral co-clustering of bipartite text data when the numbers of document and word clusters may differ. The first method uses randomized SVD through random projection, while the second combines partial SVD with element-wise random sampling. Across real-world and synthetic datasets, both methods reduce runtime relative to the full-SVD baseline, but their behavior depends on matrix sparsity. The random projection method is the more reliable approximation across the tested settings, whereas the sampling-based method is most useful on denser matrices and provides limited benefit on already sparse text data. These results show that randomized approximations for spectral co-clustering should be selected according to the underlying structure of the data.

Cite this work

Fateme Mazdarani and Carlos Toxtli-Hernández. 2026. Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices. “Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices.” arXiv preprint arXiv:2609.19243 (2026). https://doi.org/10.48550/arXiv.2609.19243

@misc{Mazdarani2026Randomized,
  title = {Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices},
  author = {Mazdarani, Fateme and Toxtli, Carlos},
  year = {2026},
  eprint = {2609.19243},
  archiveprefix = {arXiv},
  primaryclass = {cs.LG},
  doi = {10.48550/arXiv.2609.19243},
  url = {https://arxiv.org/abs/2609.19243}
}