Conformal Prediction for Statistically Reliable LLM-as-a-Judge in Recommender Systems

Maddalena Amendola, Giacomo Medda, Alessandro Soccol, Ludovico Boratto, Mirko Marras, Raffaele Perego
35th ACM International Conference on Information and Knowledge Management (CIKM) 2026
Conformal Prediction for Statistically Reliable LLM-as-a-Judge in Recommender Systems - Overview

Abstract

LLM-as-a-judge has emerged as a promising strategy for evaluating recommender systems. Yet, treating simulated ratings as real feedback can make unreliable evaluation appear reliable. This risk is especially consequential in sensitive domains such as nutrition and health. In this paper, we treat the LLM-as-a-judge paradigm as a calibrated relevance simulation problem: an LLM-based simulator predicts human preferences from user and item representations, and conformal prediction converts them into uncertainty intervals. To account for the nature of recommendation, we introduce stratified conformal calibration by user rating variance. Experiments under three LLMs and seven conformal prediction methods show that point estimates of preferences hide failures on low-rating items, while conformal interval width is driven by user-level heterogeneity.

Motivation

Evaluating recommender systems remains difficult because the available options are either efficient but biased or reliable but costly. Offline metrics such as recall and hit rate are cheap and repeatable, yet they only reflect the items users happened to encounter. Online experiments and user studies provide stronger evidence, but they are slow, expensive, and impractical for every design choice. Large Language Models (LLMs) have therefore emerged as flexible judges that simulate human preferences: given user and item data, they predict the judgment a user with that profile would provide.

Unfortunately, LLM judgments may appear confident while remaining biased, miscalibrated, or misaligned. Treating simulated ratings as real feedback can make unreliable evaluation look reliable — a risk that is especially consequential in sensitive domains such as nutrition and health.

The Gap: Conformal prediction can turn LLM evaluations into prediction intervals with finite-sample coverage guarantees, and has been applied to tasks such as text summarization. But recommendation is not a routine application: the target rating is a joint property of a user–item pair, so the same item may be relevant for one user and irrelevant for another. Uncertainty therefore depends on user rating behavior, and a single global calibration may merge stable users with volatile ones.

Approach

We frame LLM-as-a-judge as a calibrated relevance simulation problem in recipe recommendation. Given a user representation and a target recipe, the LLM predicts the rating the user would assign on a 1–5 scale. The simulated rating is treated as a proxy judgment, calibrated against held-out real ratings, and converted via conformal prediction into an interval that makes simulation uncertainty explicit.

We monitor two properties of these intervals: coverage (whether the interval contains the true rating) and efficiency (how narrow it is). If the true rating is 4, the interval [3, 5] succeeds and [1, 3] fails; among successful intervals, narrow ones such as [4, 4] or [3, 4] are far more useful than [1, 5].

Human Judgment Simulation

For each user–item pair $(u, i)$, the LLM-based simulator $f_\theta$ takes a user representation $r_u$ and item metadata $m_i$ and produces a score vector over the $K$ rating levels:

\[z_{ui} = f_\theta(r_u, m_i) = \left(z_{ui}^{(1)}, \dots, z_{ui}^{(K)}\right) \in \mathbb{R}^K\]

In practice, $z_{ui}$ corresponds to the raw logits of the valid rating tokens at the rating-generation position. The point prediction is the most supported rating, $\hat{y}_{ui} = \arg\max_k z_{ui}^{(k)}$, but the full score vector is what feeds conformal calibration.

Post-hoc Conformal Calibration

  • Nonconformity scoring: On a held-out calibration set of real ratings, a nonconformity function $s(z_{ui}, y_{ui})$ measures how inconsistent the simulator's scores are with the true rating — small if the simulator supports the real rating, large if it supports a distant one.
  • Cutoff computation: For a target miscoverage $\alpha$ (we use $\alpha = 0.10$, i.e., 90% coverage), the threshold $\hat{q}_\alpha$ is the $\lceil (n+1)(1-\alpha) \rceil$-th smallest calibration score — the standard split-conformal finite-sample correction.
  • Interval construction: For an unseen pair, every candidate rating is tested label-wise, keeping $C_\alpha(z_{ui}) = \{ y \in \mathcal{Y} : s(z_{ui}, y) \le \hat{q}_\alpha \}$. Under exchangeability, this guarantees $1 - \alpha \le P(y_{ui} \in C_\alpha) \le 1 - \alpha + \frac{1}{n+1}$. Continuous intervals are mapped back to valid Likert labels without breaking the guarantee.

User-Stratified Calibration

A single global threshold mixes different uncertainty regimes: users with stable rating behavior need narrow intervals, while users with highly variable preferences need wider ones. We therefore assign each user to a stratum by their empirical rating standard deviation and compute a group-specific threshold $\hat{q}_\alpha^{(g)}$ from that stratum's calibration examples only. The same label-wise rule then applies with the group threshold, so the coverage guarantee holds separately for each group, avoiding forcing heterogeneous raters into one global uncertainty model.

Experimental Setup

  • Dataset: HUMMUS, a recipe recommendation dataset with explicit 1–5 star ratings (39,315 users, 507,335 recipes). Since ratings are heavily skewed toward high values, we sample 1,000 users from each of four rating-std groups — [0, 0.3), [0.3, 0.5), [0.5, 0.7), [0.7, ∞) — for 4,000 users total. Per user: the 2 most recent interactions for testing, the previous 2 for calibration, and the preceding 20 as prompt history.
  • Judges: Three open-source LLMs — Qwen2.5-32B-Instruct, LLaMA-3.1-8B-Instruct, and Mistral-7B-Instruct-v0.3 — chosen for reproducibility.
  • Prompts: An items-list template (the 20 previously rated recipes with their ratings) and a profile-based template (a structured user profile built from the same history). Results were similar; items-list is reported.
  • Conformal methods: CQR, Asymmetric CQR, CHR, LVD, Boosted CQR, Boosted LCP, and R2CCP, at $\alpha = 0.10$, averaged over 30 random seeds.

Results

RQ1: Point Estimates Hide Failures on Low-Rating Recipes

At the aggregate level the three judges look comparable and reasonably accurate. Per-rating MAE tells a different story: all models fail badly on low-rated recipes, with errors above 3.6 when the true rating is 1 — they often assign very high scores to recipes users actually disliked. Aggregate performance is driven by the majority high-rating class and overstates the reliability of simulated ratings.

JudgeAcc. ↑MAE ↓MAE@1 ↓MAE@2 ↓MAE@3 ↓MAE@4 ↓MAE@5 ↓
LLaMA-3.1-8B0.7540.3443.832.641.650.740.15
Mistral-7B0.7090.3823.682.391.510.670.22
Qwen2.5-32B0.7480.3403.742.501.600.710.15

RQ2: Which Conformal Methods Work Best?

Under global (non-stratified) calibration, most methods reach the 90% coverage target, with Boosted CQR the only one consistently falling short. But coverage alone is not enough: CQR variants and CHR over-cover (~95–96%), producing conservative, unnecessarily wide intervals. Among methods meeting the target, R2CCP achieves the narrowest intervals (width < 1) across all judges.

MethodLLaMA-3.1-8BMistral-7BQwen2.5-32B
Eff. ↓Cov. ↑Eff. ↓Cov. ↑Eff. ↓Cov. ↑
CQR1.64995.601.56695.101.53295.18
Asym. CQR1.63595.561.57695.391.55795.31
CHR1.72395.411.84196.591.72695.79
R2CCP0.98190.250.95990.170.93990.42
Boosted CQR1.52988.771.59389.041.49088.06
Boosted LCP0.98489.970.97790.320.94490.37
LVD1.09490.321.10290.431.13990.12

What drives the width of these intervals? Recipe-level rating variance has negligible practical association with interval width ($\rho \le 0.06$), while user-level rating variance shows a strong, consistent positive association ($\rho = 0.48$–$0.62$). Conformal uncertainty in recommendation is driven by user preference heterogeneity, not by the items.

Mean conformal interval width versus recipe rating std and user rating std
Figure 1: Drivers of uncertainty. Interval width is almost independent of recipe-level rating variance (a), but increases substantially with user-level rating variance (b).

RQ3: User-Stratified Calibration

Calibrating within each of the four user groups (1,000 users, i.e., 2,000 calibration and 2,000 test pairs per group; a Kolmogorov–Smirnov test confirms no calibration–test shift) reveals how strongly heterogeneity shapes usefulness. Interval widths grow from roughly 0–0.5 for the most consistent raters to nearly 3.0 for the most volatile ones — on a 1–5 scale, the latter cover most of the range and are weakly actionable.

Coverage versus interval width for each conformal method across user groups and LLM judges
Figure 2: Coverage–efficiency trade-off under user-stratified calibration. Higher user variance leads to wider intervals. Dashed lines mark the 90% coverage target.

R2CCP and Boosted LCP stay closest to the 90% target across groups and judges, while CQR variants and CHR over-cover, especially for low-variance users. Boosted CQR is the least reliable, often dropping below target in higher-variance groups. No method is optimal everywhere, but R2CCP offers the most stable trade-off, particularly for low- and mid-variance users.

Key Contributions

  • Stratified conformal calibration for LLM-as-a-judge: Calibrating simulated ratings within user groups defined by rating variance, so coverage guarantees hold per group.
  • Systematic evaluation: Three open-source LLMs, two user-representation strategies, and seven conformal prediction methods in recipe recommendation.
  • Practical insights: Point-prediction metrics mask failures on low-rating recipes, and interval width is driven by user — not recipe — variance.

Takeaways

Valid coverage is achievable for LLM-simulated ratings, but interval usefulness depends strongly on who is being simulated: intervals are informative for consistent users and become broad for users with highly variable tastes. This supports user-stratified calibration as a principled strategy for uncertainty-aware recommendation evaluation, and points toward pipelines that use calibrated intervals to decide when to trust an LLM judge and when to ask real users.

Source code and data are available in the GitHub repository.

BibTeX

@inproceedings{amendola2026conformal,
  author = {Amendola, Maddalena and Medda, Giacomo and Soccol, Alessandro and Boratto, Ludovico and Marras, Mirko and Perego, Raffaele},
  title = {Conformal Prediction for Statistically Reliable LLM-as-a-Judge in Recommender Systems},
  booktitle = {Proceedings of the 35th ACM International Conference on Information and Knowledge Management},
  series = {CIKM '26},
  year = {2026},
  location = {Rome, Italy},
  publisher = {ACM},
  doi = {10.1145/3799682.3839966}
}