Abstract
Pretrained text embeddings enable content-based cold-item recommendation, but their domain-agnostic objectives may underrepresent distinctions that matter within a particular recommendation domain. We investigate lightweight autoencoder adapters that specialize these embeddings without retraining or finetuning the embedding model. The autoencoders need no interactional data to train; they merely use the set of items, whose embeddings they aim to reconstruct. We compared denoising, variational, and sparse autoencoder architectures across several embedding models and datasets. While all variants substantially improved over raw embeddings, sparse autoencoders (SAE) dominated in most settings. Qualitative inspection further reveals that some sparse dimensions capture coherent domain-specific concepts, suggesting SAE’s capability to meaningfully reorganize general-purpose embeddings. These preliminary findings motivate the inclusion of sparse semantic adaptation as a distinct component of recommender systems.
Overview
Recommender systems increasingly rely on off-the-shelf text embedding models (e.g., BGE, Qwen3) for content-based recommendation and cold-start treatment. These embeddings are intentionally optimized for broad applicability, so they must represent movies, books, recipes, and countless other concepts at once. Within a single recommendation domain, however, only a narrow subset of those distinctions is relevant — leaving domain-specific signals obscured by noise the embedding was built to carry.
This short paper investigates sparse autoencoders (SAEs) as unsupervised semantic domain adapters: lightweight adapters that specialize pretrained embeddings for a given domain without retraining or finetuning the embedding model, and without using any interaction data — the adapter is trained purely by reconstructing item-description embeddings.
Key idea: Adapt the semantic embedding space itself before it reaches the recommender, rather than learning domain-specific transformations from user–item interactions. Interactions are used only for model selection and final evaluation.
Approach in Brief
Each item's textual description is encoded with a frozen embedding model, and an ℓ1-normalized AbsTop-k SAE reconstructs those embeddings, keeping only the k largest-magnitude latent activations per item. The resulting normalized sparse code becomes the adapted item representation used for retrieval. A denoising variant (D-SAE) corrupts inputs during training to encourage a more stable semantic structure, and both are compared against dense baselines (a denoising autoencoder and a β-VAE) and the original embeddings.
Results
Across 13 datasets (MovieLens20M, GoodBooks10K, and 11 Amazon Reviews'23 categories), 3 embedding models, and 2 metrics — 78 evaluation settings in total — sparse autoencoders outperform both dense alternatives in 70 of 78 settings. The team observes a consistent progression: DAE improves over the original embeddings, β-VAE over DAE, and the sparse variants over both. Qualitative inspection of individual latent dimensions reveals coherent domain-specific concepts (e.g., brake components in Automotive, staplers and sharpeners in Office Products, baby bottles and swaddles in Baby Products), suggesting sparse adaptation reorganizes embeddings into meaningful concepts rather than merely compressing them.

Takeaway
These preliminary findings position semantic adaptation as an independent component of modern recommender systems rather than a preprocessing afterthought, and suggest that part of the gains attributed to cold-start and semantic-ID methods may stem from domain adaptation of the embedding space itself.
Full results, implementation details, and source code are available in the GitHub repository.
BibTeX
@inproceedings{vancura2026sparse,
author = {Van\v{c}ura, Vojt\v{e}ch and Medda, Giacomo and Spi\v{s}\'{a}k, Martin and Pe\v{s}ka, Ladislav},
title = {Sparse Autoencoders as Semantic Domain Adapters for Recommender Systems},
booktitle = {Proceedings of the 20th ACM Conference on Recommender Systems},
series = {RecSys '26},
year = {2026},
location = {Minneapolis, MN, USA},
publisher = {ACM},
doi = {10.1145/3773078.3841253}
}