Compresso: Espresso-Style Sparse Representation Learning for Interpretable Recommender Systems

Vojtěch Vančura, Giacomo Medda, Martin Spišák, Ladislav Peška
20th ACM Conference on Recommender Systems (RecSys) 2026
Compresso: Espresso-Style Sparse Representation Learning for Interpretable Recommender Systems - Overview

Abstract

Sparse representations can make recommender-system embeddings more compact and inspectable, but developing sparse-learning workflows typically requires substantial engineering around sparsification, training, pruning, storage, and analysis. We present Compresso, an open-source PyTorch framework that exposes this functionality through a simple and modular interface. Inspired by Italian espresso culture, where one orders a caffè while the barista handles the machinery, Compresso lets researchers focus on sparse models rather than infrastructure. The framework provides high-level sparse autoencoder training, reusable sparse tensor representations, differentiable top-k operators, sparse and masked neural parameters, pruning schedules, and composable clustering tools.

Motivation

Modern recommender systems increasingly rely on dense user and item representations for retrieval, ranking, personalization, and semantic matching. These representations capture rich semantic and collaborative information, but they are difficult to inspect, and their storage and computational costs can become substantial at scale. Sparse representations offer an alternative: each entity is described by a small set of active features, supporting compression, efficient retrieval, and analysis of recurring semantic or collaborative factors.

Despite growing interest, experimenting with sparse representations still requires substantial supporting infrastructure — sparsification operators, gradient estimators, training and pruning schedules, sparse storage, model integration, persistence, and downstream analysis. As a result, implementations are often tied to individual methods and difficult to adapt or reuse.

The Gap: Existing sparse-representation methods for recommendation target specific applications (compression, retrieval, interpretation), while general sparse-autoencoder libraries mostly serve language- and vision-model activations. Recommender-system libraries, in turn, focus on model training, evaluation, and benchmarking rather than reusable sparse representation learning.

The Compresso Framework

Compresso is an open-source PyTorch framework for learning, storing, and analyzing fixed-k sparse representations. Its name reflects Italian espresso culture: one orders a caffè, while the barista handles the grinding, brewing, and machinery. Similarly, Compresso lets researchers focus on sparse models and their applications while the framework handles the underlying infrastructure — combining a high-level fit/transform workflow, reusable sparse data structures and PyTorch components, and composable pipelines for clustering and analyzing sparse activation patterns.

Compresso workflow: dense embeddings encoded by a top-k SAE into an SRPTensor, then grouped into interpretable clusters
Figure 1: Compresso workflow. A top-k SAE converts dense user or item embeddings into fixed-k sparse representations stored in an SRPTensor. These representations can be persisted, reconstructed, and grouped into interpretable clusters based on shared activations.

Design and Architecture

Compresso's design goal is an "espresso-style" workflow in which dense embeddings become sparse representations — suitable for storage, analysis, and clustering — in only a few lines of code. The high-level workflow centers on a top-k sparse autoencoder that maps each dense vector into a latent feature space and retains the k highest-scoring activations. The scoring rule is configurable (sparsification by raw activation or absolute magnitude), and the resulting representations have a predictable storage budget while exposing semantic or collaborative signals across users, items, or contexts.

Core Objects and API

The main user-facing API is built around TopKSAETrainer, TopKSAEConfig, and SRPTensor. The trainer follows a familiar scikit-learn-style interface: fit trains the sparse autoencoder on a dense embedding matrix, while transform encodes compatible vectors into the learned sparse latent space. This separation matters in recommendation, where a model trained on an existing catalog or split can subsequently be applied to new items, users, or contexts.

The resulting SRPTensor (Sparse Row-Packed — a sparse representation with a fixed number of entries per row) stores active feature indices and values while preserving the logical dense shape. It can be converted to dense or standard sparse tensor formats, persisted, and reused for downstream retrieval, model input, or analysis.

Advanced Features

Beyond the high-level autoencoder workflow, Compresso exposes lower-level components for integrating sparsity directly into custom PyTorch models:

  • topk_ste / TopKSparsify: Apply hard top-k selection along a chosen dimension using a straight-through estimator for gradient-based optimization.
  • SRPParam: Represents a row-packed sparse parameter with fixed indices and trainable values, for known sparsity patterns.
  • MaskedParam: Pairs dense weights with scheduled binary masks for progressive row-wise pruning, when the sparse structure must be discovered during training.
  • SparsityController: Manages mask updates and optional weight rewinding.

Together, these components support both explicitly sparse parameters and gradual dense-to-sparse training, without reimplementing sparsification, pruning, or sparse storage.

Clustering and Analysis

Compresso treats sparse representations not only as compressed model inputs but also as objects for analysis. Entities that activate similar sparse features may share product categories, semantic themes, visual properties, user preferences, or collaborative signals. The ClusteringPipeline captures this structure through composable steps for clustering, linking, merging, pruning, filtering, and labeling activation patterns. Clusters can be initialized from dominant signed features, feature combinations, or representation similarity, and labeled from entity metadata or user-defined labeling functions — giving practitioners an inspectable view of which sparse factors characterize a cluster and which users or items activate them.

Online Demonstration

The interactive demo presents Compresso end-to-end: dense product metadata is embedded and encoded with a top-k sparse autoencoder, products sharing activation patterns are grouped into clusters, and the resulting groups are shown as labeled clusters of representative items. A methodology page exposes the code used for embedding preparation, sparse encoding, clustering, and labeling, while domain pages present clusters for several Amazon product categories (automotive, baby products, grocery and gourmet food, and more).

Compresso online demo showing SAE training code alongside browsable product clusters
Figure 2: Compresso online demo for the Amazon Clothing, Shoes and Jewelry domain. Left: source code for SAE training and running the clustering pipeline. Right: the interface for browsing product clusters, inspecting latent-factor details, and adjusting the number of displayed clusters and representative products. Each row is a cluster induced by shared sparse activation patterns.

Each cluster contains representative products together with its latent-factor identifier and activation direction. Users can vary the number of displayed clusters and products, inspect individual sparse factors, and trace results back to the corresponding pipeline code — connecting sparse activation patterns with recognizable product categories, semantic themes, and visual similarities.

Try it live: the online demo is available at compreapp-demo.streamlit.app.

Key Contributions

  • Espresso-style sparse learning: An open-source PyTorch framework that hides the infrastructure of sparsification, training, pruning, storage, and analysis behind a simple fit/transform interface.
  • Reusable sparse building blocks: Fixed-k SAEs, the SRPTensor representation, differentiable top-k operators, and sparse/masked parameters that drop into custom PyTorch models.
  • Interpretability by design: A composable ClusteringPipeline that organizes recurring activation patterns into inspectable, labeled clusters for studying semantic and collaborative structure.
  • End-to-end demonstration: An interactive demo over real Amazon product domains that links sparse factors to recognizable clusters and the code that produced them.

Getting Started

Compresso is available on PyPI as compresso-pytorch, with source code, installation instructions, and complete examples in the GitHub repository.

BibTeX

@inproceedings{vancura2026compresso,
  author = {Van\v{c}ura, Vojt\v{e}ch and Medda, Giacomo and Spi\v{s}\'{a}k, Martin and Pe\v{s}ka, Ladislav},
  title = {Compresso: Espresso-Style Sparse Representation Learning for Interpretable Recommender Systems},
  booktitle = {Proceedings of the 20th ACM Conference on Recommender Systems},
  series = {RecSys '26},
  year = {2026},
  location = {Minneapolis, MN, USA},
  publisher = {ACM},
  doi = {10.1145/3773078.3841254}
}