Score-based Generative Diffusion Models for Social Recommendations
Chengyi Liu, Jiahao Zhang, Shijie Wang, Wenqi Fan, Qing Li
I Introduction
In the era of information explosion, Recommender Systems (RS) have emerged as indispensable tools for mitigating information overload and personalizing users’ content stream across e-commerce and social media platforms. Collaborative Filtering (CF), which aims to learn user and item representations from historical interactions, forms the foundation of modern recommender systems . However, CF-based methods face a persistent challenge due to their reliance on the quality of user-item interaction data: real-world interaction data used for training is often sparse, which can adversely impact the performance of these models . To address the data sparseness issue, social recommendation methods incorporate social context to enhance the modelling of user intent within user representations. This practice aligns with the idea that individuals connected through social ties often exhibit similar preferences—a phenomenon supported by social influence theory and also recognized as social homophily . In recent years, a wide range of research has focused on enhancing the performance of social recommendations by leveraging the social homophily condition, encompassing Graph Neural Networks (GNNs) methods and Self-Supervised Learning (SSL) methods .
Nevertheless, the underlying assumption of social homophily has recently been questioned. Extensive statistical and empirical analyses have suggested that leveraging the raw social graph can be unreliable and may further degrade the recommendation performance. This counter-intuition arises because of the inherent complexity and noise present in real-world social networks , introducing noisy information into learned user representations. Considering the example in Fig. 1, a male student (i.e., User A) might establish social relationships with a female classmate (i.e., User B) who has recently acquired cosmetics or feminine apparel, yet such relationships likely contribute to unreliable recommendations for the male. Consequently, these noisy social signals can propagate through social encoder layers, accumulating bias and ultimately hindering the recommender system’s effectiveness . Thus, it is necessary to develop social denoising methods to filter redundant social signals from user-user social graph, alleviating the negative impact of such signals on recommendation performance.
Efforts to address the low social homophily problem primarily rely on data-centric graph editing methods , which denoise the social graph by deleting or rewiring edges based on human-designed heuristics, resulting in improved user representations. For example, GDMSR uses a transformer layer to compute user similarity with users’ interacted item lists, thereby refining the social network according to these user similarities. However, these methods face two key limitations in achieving optimal social representation consistent with collaborative signals:
Handcrafted Similarity Metrics: These graph editing methods, guided by the heuristic rules, rely solely on handcrafted user similarity metrics (e.g., cosine similarity), which lacks generalizability and effectiveness in evaluating social networks. These metrics are typically limited to one-hop modifications of social relations and depend heavily on the representation capacity of the underlying backbone models in the social domain. As illustrated in Fig. 2 (a), handcrafted heuristics may lead to suboptimal user representations (denoted by dark circles ) that are far from the optimal representation (denoted by red star ).
Expressive Capability Limitation: These methods are unable to adaptively account for all users within a social network, which is computationally infeasible due to the vast number of users. The graph rewiring algorithms focus on a small, predefined subset with a fixed number of candidates, such as one-hop neighbours or users with high embedding similarity. This approach neglects users with similar preferences who do not conform to these predefined criteria, thereby missing opportunities to enhance the alignment between social and collaborative signals. Furthermore, extending denoising to multi-hop optimization may introduce conflicts between local and global optima in the graph editing process.
To overcome the aforementioned limitations, we propose a fundamentally novel approach to denoise social information in recommendation systems (RS) from a generative perspective. Specifically, our key idea is to model the underlying distribution of users’ social representations and directly generate denoised social representations that maximize consistency with collaborative signals. This generative paradigm is highly desirable as it allows us to transcend all possible graph structure revisions and achieve the ideal social representation, which may be unreachable through graph editing steps under computational constraints. To instantiate this novel generative paradigm, we draw inspiration from the success of diffusion models in image restoration , and propose to leverage the potential of score-based generative models (SGMs) , a strong yet unexplored class of diffusion models. Empowered by stochastic differential equations (SDEs), SGMs unify the diffusion processes, such as DDPM and SMLD , in continuous time steps. This unified continuous formulation enables a predict-and-correct generative paradigm that flexibly incorporates a wide range of numerical SDE solvers and statistical methods, thereby reducing numerical errors in the computation of diffusion models and more accurately capturing the complexities of the social denoising problem .
Despite the great potential, it is highly challenging to directly take advantage of SGMs for social denoising in social recommendations. First, the optimal social network that accurately aligns with users’ preference patterns remains unknown, leading to a lack of appropriate supervision signals (i.e., training objectives) for SGMs. Additionally, since SGMs were originally designed for image generation and denoising, applying them to social recommendations introduces multiple challenges, such as handling high-order graph structures and aligning signals from social and collaborative domains. To address the challenges above, we propose the Score-based Generative Model for Social Recommendation (SGSR) to transcend the social homophily condition in social recommendations with SGMs. Instead of manually determining the sub-optimal embeddings from edited social graphs as supervisory signals, a curriculum learning mechanism is introduced to facilitate the modelling of social representations across social graphs with incrementally sparse levels. Furthermore, a joint optimization method is developed to leverage collaborative signals to implicitly identify the optimal social representation. Our main contributions are summarized as follows:
We propose a score-based generative framework (SGSR) that effectively captures the underlying distribution of social representations and improves their collaborative consistency.
To effectively adapt the SGM in the context of social recommendation, we derive a novel classifier-free conditioning objective to achieve the personalized denoising diffusion process based on the user-item behaviour.
To address training obstacles of leveraging SGMs for social denoising, we introduce an efficient training mechanism, including a curriculum learning strategy to ease the distribution learning process and a joint training mechanism to implicitly guide the denoising process.
We conduct comprehensive comparison experiments and ablation studies on three real-world datasets. The results demonstrate that the proposed SGSR framework significantly outperforms state-of-the-art baselines, validating the effectiveness of our proposed method.
II THE PROPOSED METHOD
In this section, we first introduce the key notations and definitions used in this paper, including the basic concepts of the score-based generative diffusion models. Subsequently, we outline the proposed framework and provide a detailed description of each model component.
II-B An Overview of the Proposed Framework
To alleviate the low social homophily problem in social recommendations, we propose the SGSR framework to provide accurate recommendations with generatively denoised social representations. Specifically, the SGSR framework incorporates a score-based social denoising model, which encodes the high-order relations in the social network into latent space and denoises the social embeddings in continuous time through a diffusion process. To control the social denoising model with collaborative signals, we propose a novel classifier-free training objective and enhance the training process with a curriculum learning mechanism. Eventually, the filtered social representations are incorporated into the collaborative domain using self-supervised learning techniques to enhance the primary recommendation task. The overall framework is demonstrated in Fig. 3.
II-C Score-based Generative Social Denoising
To address the limitations of existing graph-based social denoising methods, we adopt a controllable SDE-based diffusion process to directly optimize social representations. This approach enables the continuous refinement of social representations, guided by user interaction behavior.
(1) Graph Representation Learning. To encode the high-order social relations into the denoising framework, we first transform the original social embeddings into a latent space before the diffusion process. Specifically, we employ the widely used Graph Convolutional Network (GCN) as the social encoder (denoted by ) of our denoising framework, which follows an iterative message-passing process described as follows:
Additionally, we obtain user and item representations from the collaborative domain to control the social denoising process. We adopt an arbitrary graph-based encoder (e.g., LightGCN , PinSAGE , etc.) for collaborative filtering with layers, denoted by , and obtain the collaborative representation as follows:
where is the adjacency matrix of the collaborative graph. Note that the choice of is independent of the input of our generative social denoising model, indicating that the proposed SGSR is a general framework adaptable to existing recommendation backbone models.
(2) Collaboratively Controlled Denoising Diffusion Paradigm. After capturing the long-range and non-linear relations in the social network, we apply the score-based diffusion process on the propagated social representation for further denoising. Specifically, SGMs use two Stochastic Differential Equations (SDEs) to model the noising and denoising processes of the data distribution.
Forward Process. The forward process aims to transform the original data distribution of social representations into a prior distribution (e.g., Gaussian noise) in a continuous manner. Considering the forward diffusion process with respect to the continuous time step , we first define the raw social embedding of a user as the initial state , taken from embedding matrix , and follows the prior distribution. The forward diffusion process is defined by the following SDE:
Reverse Process. The reverse-time SDE simulates the diffusion process in reverse time, reconstructing samples from a known prior that adhere to the original data distribution. We run the reverse process conditioned on the user’s collaborative information to guide the recovery of the user representation towards a collaboratively consistent direction. This can be formulated as the following conditioned reverse-time SDE :
where denotes the Wiener process evolving in reverse time, denotes the data distribution at time , and the condition vector from encodes the user’s preference pattern in the collaborative domain. The reverse SDE requires the gradient of , known as the score , which can be estimated with a time-aware scoring neural network .
Training. The core of training SGMs involves optimizing the conditional scoring network using the score-matching objective, which aligns the network’s output with the actual score function:
where is estimated by a classifier trained on the perturbed representation . Nevertheless, training an explicit classifier for each user’s social representation at every noise scale is computationally infeasible, considering that there could be thousands of users in the context of recommendations, meaning thousands of classes to train. This constraint necessitates the development of a classifier-free guidance approach for implementing score-based methods in the context of social recommendation. As an alternative, we regard the collaborative condition as disentangled pseudo-labels to achieve controllable generation, thereby circumventing the need for explicit classification gradient. Subsequently, we construct a differential loss function to facilitate the training of time-dependent conditional score network . Based on Tweedie’s formula and the properties of disentangled representations , we mathematically derive the training loss from the minimization of conditional score matching in Eq. (5):
Thus, the denoising score function depends only on the noising process and can be evaluated in closed form. Finally, the SGSR employs the following classifier-free objectives to optimize the conditional scoring network :
where is empirically set to 1 based on experimental results.
SDE Instantiation. Despite the universality of the score-based generative paradigm, it is essential to adopt specific designs for the drift coefficient and diffusion coefficient to instantiate the SGM into practice. to instantiate the SGM in practice. To this end, we implement the drift and diffusion coefficients using two widely adopted formulations: the Variance Exploding (VE) SDE, a continuous form of the Score Matching with Langevin Dynamics (SMLD) model, and the Variance Preserving (VP) SDE, a continuous form of the well-known Denoising Diffusion Probabilistic Models . Specifically, the forward processes of these two SDEs are formalized as follows:
where and represent the predefined noise scalers. The reverse processes of these SDEs can be obtained via Eq. (4).
Besides, another challenge in implementing the SGM with specific SDEs lies in the computation of the denoising score function , which is an indispensable part of the reverse process. Fortunately, the specifications of the drift and diffusion coefficients enable a closed-form solution for the Gaussian transition kernel , facilitating an efficient training process.
Thus, with the closed-form Gaussian transition kernel, it is straightforward to derive the corresponding probability density function and obtain the closed-form solution of the score functions.
Social Reperesentation Denoising. After training the score-based social denoising model, we obtain the scoring network and use it to denoise users’ social representations. Given a user’s raw social representation , the model first perturbs the raw social representation in a continuous manner using either the forward VP SDE or VE SDE in Eq. (10) to obtain the output . It then reconstructs the denoised representation by iteratively solving the collaboratively conditioned reverse-SDE (Eq. 4) with off-the-shelf SDE initial-value problem (IVP) numerical solvers . To obtain the numerical solution of a given SDE, we adopt the Predictor-Corrector (PC) method , a flexible and powerful framework for solving SDEs inherent in SGMs. The iterative steps of the Predictor-Corrector solver first use the prediction of numerical SDE solvers like the Euler–Maruyama method , and then correct the predicted trajectory with score-based MCMC methods, such as Langevin MCMC , achieving a balance of numerical accuracy and efficiency. The Algorithm 1 outlines the inference process of the denoising model, incorporating annealed Langevin dynamics as a corrector ( represents the step size for Langevin Dynamics).
II-D Curriculum-Driven Training of the Social Denoising Model
Since the optimal social representation that maximizes consistency with collaborative signals is difficult to determine, training the SGM for social graph denoising presents significant challenges due to the lack of ground truth and supervision signals. For example, in image restoration tasks, SGMs use deteriorated images as input and compute the loss function between the ground-truth image and the denoised image. This straightforward training strategy is inapplicable in the context of social denoising, as the ground-truth optimal social representation is not accessible. To overcome this limitation, we employ a curriculum learning strategy that enables the SGM to perceive social representations across social graphs with varying levels of sparsity. This approach potentially encompasses the optimal social graph, which the SGM model can identify automatically with collaborative signals.
The two-stage curriculum training mechanism implements a progressive learning paradigm, gradually exposing the SGM to increasingly sparse social distributions. The sparser social distributions, characterized by fewer relationships, encapsulate reduced learnable information, posing greater challenges for models to learn the distributions . This method forms an easy-to-hard training strategy, enhancing the robustness of the learning process. Specifically:
In the initial stage, the model retains the percentage of social connections that exhibit the most similar preference patterns for epochs. Given the collaborative representations of user and from as and , we leverage the inner product to measure the similarity .
In the second stage, a random selection mechanism keeps percentage of social relations for the subsequent epochs in the curriculum period, facilitating SGM’s access to a broader range of social representations.
This approach encourages the SGM to model the relationship between collaborative conditions and social patterns across varying levels of graph sparsity. The initial process focuses on user-level sparsification, selectively removing seemingly irrelevant social connections that are more readily learnable , while the latter method implements a holistic network sparsification, challenging the SGM to model more diverse representations .
The training objectives in Eq. (8) can be updated with the sparse social representation obtained by the same from the post-edition social graph. The updated loss, , is formulated as:
where denotes the social embedding of sampled user , represents the distribution of data at the intermediate diffusion step , representes the corresponding collaborative representation as disentangled condition. Note that the SGSR does not designate a specific sparsity level as the optimal representation; instead, it is determined by the recommendation signal through joint training strategy as described in 2.
II-E Score-based Denoising Model Optimization
Besides explicit conditioning on the observed user behaviour patterns, the SGSR leverages the recommendation signal to implicitly guide the SGM in identifying the optimal social representation. On the other hand, to integrate the produced social representation into the collaborative domain, the SGSR introduces an auxiliary SSL loss to augment the primary recommendation task.
(1) Cross-domain Contrastive Learning. The stochastic Wi- ener process and exposure bias in the interactive sampling process may introduce undesired noise to the generated social representation . To mitigate these biases and effectively leverage the learned social representation distribution, the SGSR incorporates contrastive learning techniques, known for their high efficiency in knowledge transfer . The SGSR enhances the RS with social information by maximizing the consistency of user representations across the social and collaborative domains.
As described in II-C, we construct the denoised social representation , produced by the SDE-based diffusion process as the self-supervised signal. To integrate social information into the collaborative domain, we adopt the InfoNCE loss as our contrastive learning objective to optimize the user representations by maximizing the agreement between and , which can be defined as:
where and denotes the social representation and collaborative representation of user projected to the same semantic space by multi-layer perceptions , , and is the temperature parameter.
(2) Joint Training Mechanism. The SGSR leverages the recommendation signal to guide the social representation refinement trajectory with the well-established Bayesian Personalized Ranking (BPR) loss, denoted as .
where represents the set of items interacted by user , denotes the sigmoid function, and represent the predicted scores for user over the positive item and negative sample, respectively. The joint optimization loss for the SGSR is:
where and are hyperparameters to balance the social denoising score matching and the recommendation accuracy. Details of the training process of the denoising model are shown in the Algorithm 2.
III Experiment
In this section, we evaluate the effectiveness of the proposed SGSR framework through extensive experiments.
(1) Datasets. The experiments are conducted on three real-world social recommendation datasets: Ciao Ciao and Epinions datasets: https://www.cse.msu.edu/~tangjili/trust.html, Epinions1, and Dianping https://lihui.info/data/dianping/. Each dataset records users’ ratings on items on a scale of $$ and their social connections. These three datasets have been shown to satisfy the low social homophily assumption (with graph-wise homophily ratio smaller than 0.1 ) and used to evaluate the social denoise task according to the previous literature . Following a widely used setting , we transform the explicit ratings into implicit interactions by only keeping interactions with ratings above 3, and then filter out users and items with fewer than three interactions to ensure dataset quality. For all three datasets, we split the data into training, validation, and testing sets in a ratio of 8:1:1. The statistics of the datasets are summarized in Table I.
(2) Evaluation Metrics. To evaluate top- recommendations from implicit user feedback, we use two standard ranking metrics: Recall@ (R@) and Normalized Discounted Cumulative Gain (N@). We set for evaluation and rank items among all candidates to conduct a full-rank evaluation protocol. To ensure the stability of the results, each experiment is repeated 5 times, and we report the average values.
(3) Baselines. We evaluate the SGSR’s performance against the state-of-the-art baselines across several categories: vanilla social recommendation models (SocialMF , NGCF , LightGCN , DiffNet , SEPT, MHCN ), diffusion recommendation models (DiffRec , DreamRec , DreamRec+), and social graph denoising methods (DESIGN , DSL , GBSR , SHaRe ). Descriptions of these baseline methods are provided as follows:
DiffRec : Diffrec initially introduces the diffusion algorithm to the recommendation context, which perturbs the user’s historical interaction over the whole item vector.
DreamRec : DreamRec accomplishes the sequence recommendation by generating the consequent item embedding and recommending with item retrieval techniques.
DreamRec+: DreamRec+ combines the social information with the generation guidance embedding to adopt the DreamRec to social recommendation
SocialMF : SocialMF propagates the social information into the matrix factorization model
NGCF : NGCF leverages the graph convolutional network to capture collaborative signals by propagating embeddings on the user-item interaction graph.
LightGCN : LightGCN simplifies NGCF by removing the feature transformation and nonlinear activation layers, specifically designed for CF tasks.
DiffNet : DiffNet implements layer-wise propagation mechanisms to model the diffusion of social influence in recommendation systems.
MHCN : MHCN incorporates multi-channel hypergraph convolution operations to capture complex high-order correlations among users, items, and social relations.
SEPT : SEPT introduces a socially aware self-supervised learning (SSL) framework that incorporates tri-training to capture supervisory signals from both the node itself and its neighbouring nodes.
DESIGN : DESIGN leverages knowledge distillation to integrate information from both the user-item interaction graph and the social graph.
DSL : DSL leverages SSL techniques to denoise the social representation by aligning the embedding with those who share similar preference patterns.
GBSR ): GBSR introduces a theoretically-inspired framework to maximize the mutual information between the denoised social graph and the interaction matrix
SHaRe : SHaRe designs the graph rewiring technique to discretely edit the social graphs based on the cosine similarity between user’s collaborative representations.
For all baselines, we adhere to their official implementations and recommended settings to ensure fair comparisons. Specifically, the SocialMF, NGCF, LightGCN, DiffNet, MHCN, and SEPT were implemented using the publicly available QRec https://github.com/Coder-Yu/QRec library which is built on TensorFlow. The implementations of DreamRec, Design, DSL, and GBSR were derived from their official code releases. The DreamRec+ model integrates social information obtained through the GCN encoder as the guidance embedding for DDPM, utilizing a multilayer perceptron for this integration. Moreover, we implement social denoising baselines requiring a recommendation backbone model, as well as our proposed SGSR with the most widely used LightGCN backbone. For a fair comparison, we employ a grid search strategy to tune the hyperparameters of the baseline models, exploring the vicinity of the optimal values reported in their respective original papers. All methods are optimized using the Adam optimizer. For the diffusion-based recommenders, we meticulously tune both the upper and lower bounds of the noise, as well as the number of diffusion steps, to maximize performance. Take 5 steps in DiffRec as an example, the upper bound and lower bound are searched in 0.5, 0.1, 0.05, 0.01 and 0.01, 0.001, 0.0001, 0.0001 .
(4) Parameter Settings. The proposed framework is implemented with PyTorch. The embedding dimension is tuned within 32, 64, 128, 256. We use batch sizes of 1024 for Ciao and 2048 for Epinions and Dianping, and sample one negative instance for each positive instance to compute the BPR loss. We use 3 GNN layers to encode the high-order social relations for our denoising framework. We employ the Adam optimizer with a learning rate of 0.001 to train the SGSR model. The diffusion step is evaluated across for both VE-SDE and VP-SDE formulations. We examine two predictors (reverse diffusion sampler, Euler-Maruyama) in combination with two correctors (Langevin MCMC, none). The maintain rate and in the curriculum learning mechanism are tuned within 0.7, 0.75, 0.8, 0.85, 0.9. Additionally, the temperature parameter for SSL, , which is set in . The loss weights and are tuned in .
III-B Overall Performance
We evaluate the recommendation performance of all baselines and our proposed approach SGSR. Table II summarizes the Recall and NDCG score across the Ciao, Epinions, and Dianping datasets. Our main findings include:
SGSR consistently outperforms all baseline methods, demonstrating significant improvements across all datasets. Notably, on the Ciao dataset, SGSR achieves a 5.49% improvement over the top-performing baseline. This substantial performance gain highlights the effectiveness of addressing low homophily challenges from a generative perspective. In particular, the classifier-free SGM used in SGSR effectively generates collaborative-consistent social representations, which significantly contributes to its superior performance.
In most cases, social denoising methods surpass the GNN-based social recommendation that directly utilizes the raw social graph. This observation underscores the importance of addressing the challenge of low social homophily. However, the performance of these social recommendation methods is sub-optimal compared with the proposed SGSR, likely due to their reliance on simple, handcrafted similarity metrics and the restricted scope of rewired edges.
Incorporating social information typically enhances the recommendation performance of both GNN-based models and diffusion-based approaches. However, the performance of DiffNet and DreamRec+ degrades when social networks are integrated. This suggests the presence of low social homophily and indicates the necessity to filter out redundant information from social relations.
Diffusion-based methods underperform most vanilla social recommendation models. This observation indicates that existing diffusion models for recommendation cannot be simply adapted to the social recommendation task, necessitating the need to address the unique challenges in developing a diffusion-based recommender model for social recommendations.
III-C Ablation Study
This study aims to evaluate the impact of the key components in SGSR. To comprehensively analyze the effectiveness of each module, we design three model variants:
w/o cur: This variant removes the Curriculum Learning Mechanism, allowing the SGSR to directly learn the social representation from the raw social graph . Consequently, is reduced to Eq.(8).
w/o sgm: This variant eliminates the Score-based Generative Model (SGM), instead employing traditional data augmentation techniques to generate two social views for cross-domain contrastive learning.
w/o ssl: This variant replaces the contrastive learning module with multi-layer perceptrons to fuse user representations from the social and collaborative domains.
Figure 4 presents the results of the ablation study. Firstly, the removal of the SGM resulted in significant performance degradation, underscoring its crucial role in the overall architecture. This finding highlights the importance of SGM in enhancing the model’s predictive capabilities. Secondly, we observed a reduction in performance when the curriculum training module was eliminated. This decline suggests that the inclusion of the entire raw social graph can negatively impact the model’s effectiveness, demonstrating the importance of social denoising tasks. The curriculum training approach appears to mitigate these effects by providing a more structured learning process. Lastly, the most substantial performance decrease was observed upon the removal of the Self-Supervised Learning (SSL) component. This outcome aligns with our previous experimental results and strongly indicates the effectiveness of the contrastive learning module. The SSL component plays an important role in capturing nuanced features and relationships within the data, thereby significantly contributing to the model’s overall performance.
Ablation Study on Generative Models. We further conduct comparison experiments with alternatives of generative models including VAE, DDPM , NCSN , and different variants of SGMs. For VAE, DDPM, and NCSN, we used their official code to ensure a fair comparison. For the SGM variants, the implementation details are provided in Section III-D, achieving different instantiations of SGM with various drift and diffusion coefficients. The results are summarized in the Table III.
From the table, we observe the following: (1) Diffusion-based generative models outperform VAE, likely due to diffusion algorithms breaking down denoising tasks into multiple steps, which enhances effectiveness. (2) DDPM achieves results comparable to SGM (VP-SDE without predictor), as DDPM can be seen as a discretized version of SGM . (3) Variants of SGM outperform both DDPM and NCSN, demonstrating the effectiveness of the PC sampler and justifying the advantage of SGM.
Besides, GANs are not included in this table. As mentioned earlier, generative models like GANs are not directly applicable to social denoising. Due to the challenge of dealing with unknown optimal social networks, training GANs in this context is infeasible, as it becomes difficult to generate suitable positive and negative samples for effective training.
Ablation Study on Graph Encoders. Graph encoders are employed to extract collaborative and social information from graph-structured data. We evaluate SGSR with various recommendation backbones to assess the impact of different encoders, following the designs of previous works. The results, presented in Table IV, show that the proposed method achieves comparable performance with mainstream encoders. SGSR is adaptable to different encoders, even though each encoder produces representations of varying quality. This can be attributed to SGM, which enhances the denoised social representation with the learned data distribution.
III-D Parameter Sensitivity Study
In this section, we evaluate the impact of the key parameters in SGSR.
(1) Diffusion Steps. We tune the scale of the continuous time , known as diffusion steps, on two well-established formulations: VP-SDE and VE-SDE. The results are presented in Fig. 5. For all the datasets, we observe optimal performance at 5 diffusion steps. Performance degrades as the number of time steps increases, potentially due to the introduction of additional noise in longer diffusion processes. Moreover, VP-SDE outperforms VE-SDE in most cases.
(2) PC Sampler. We implemented the Predictor-Corrector (PC) sampler using various combinations, including Euler-Maruyama and Langevin dynamics, Euler-Maruyama without correction, reverse diffusion sampler and Langevin dynamics, and reverse diffusion sampler without corrector. The reverse diffusion sampler is a discretization of DDPM . The results, presented in Fig. 6, demonstrate that the predicting and correcting paradigm consistently outperforms the non-corrector version across all three datasets. This highlights the effectiveness of the PC sampler and the flexible sampling capabilities of SDE-based diffusion models.
(3) Maintain Rate in Curriculum Learning Mechanism. We varied the maintain rate and within the set . The results are presented in Fig. 7. SGSR performance peaks when we remove and of social relations for the Dianping dataset during the two-stage curriculum learning. This underscores the necessity of social graph denoising, as including the entire raw social graph would degrade recommendation performance.
III-E Model Investigation
In this section, we provide a deeper analysis of GBSR, focusing on memory cost, time analysis, scalability limitation, and training efficiency.
(1) Memory Analysis. We conducted additional experiments to assess the memory consumption of the social denoising models on the Ciao and Epinions dataset, employing consistent settings (batch size: 1024, embedding size: 128, 3-layer GNN encoder for both social embeddings and collaborative embeddings). The Table V, reported in MiB, demonstrates that while memory efficiency is not a particular strength of SGSR among social denoising methods, memory usage remains within acceptable bounds.
(2) Time Analysis. We record the training time for various social denoising methods on three datasets, using consistent hyperparameter settings (batch size: 1024, embedding size: 128, learning rate: 0.0001, optimizer: Adam, 3-layer GNN encoder for both social embeddings and collaborative embeddings). The results are presented in table VI. Despite leveraging the 5-step diffusion process, SGSR’s time consumption remains high due to the stepwise denoising. Besides, the training of SGSR learns the distribution of social representations rather than merely estimating user similarity. Future work may focus on achieving one-step embedding generation with a consistent model to reduce time costs.
(3) Scalability Analysis. In SGSR, the SGM is applied to the social representations of users, initially learned by GNN-based encoders. Unlike diffusion methods on graphs, the computational cost of the SGM does not increase with the graph size . Therefore, the scalability of SGSR is not constrained by graph size, particularly due to the diffusion model. This also motivates the idea of addressing the low-homophily challenge from a generative perspective rather than graph editing methods.
(4) Convergence Analysis. To further investigate the training efficiency of the proposed method, we compare the convergence process of SGSR and the w/o sgm variant in terms of recall metrics and the BPR loss. As shown in Fig. 8a, the proposed SGSR exhibits significantly faster and more stable convergence, peaking around 350 epochs. In contrast, the w/o sgm variant, which directly utilizes the raw social representation, converges around 500 epochs and presents considerable fluctuations. Regarding the BPR loss trend in Fig. 8b, the w/o sgm variant shows notable fluctuations in loss reduction, whereas the SGSR framework demonstrates a smoother decrease. The stability of the training process can be attributed to the effective SGM-based denoising method and the efficiency of the curriculum training algorithm.
IV Related Work
In this section, we summarize the related works on the social recommendation task and diffusion models for recommendations.
(1) General Social Recommendation Methods. The widespread use of online social networks has alleviated the data sparsity issue in recommender systems by incorporating social context to enhance user representation modelling. Early approaches to social recommendations adopted the social homophily assumption and directly leveraged raw social networks, evolving through matrix factorization , graph neural networks , and self-supervised learning . However, these methods struggle to address the redundant or irrelevant information inherent in raw social networks, introducing noise into user representations. This problem necessitates systematic denoising methods to improve social recommendation.
(2) Data-centric Social Denoising Methods. Current social denoising methods primarily focus on modifying users’ social relations by assessing the confidence of social connections using handcrafted metrics to filter out irrelevant relations. For example, GDMSR employs a transformer layer to evaluate the confidence of social connections based on historical interactions, removing a fixed percentage of edges with the lowest similarity scores. SHaRe and MADM use cosine similarity to assess the reliability of social relations, discarding or rewiring connections with low similarity scores and keeping reliable connections. GBSR introduces a theoretically-inspired framework to maximize the mutual information between the denoised social graph and the interaction matrix, but it only drops the links in the social network while ignoring possibly beneficial edge rewiring or injecting operations. Generally speaking, existing data-centric social denoising methods can hardly reach the optimal social representation, since the best-fitted measure of social relations remains unknown and cannot be simply designed by human heuristics. Besides, graph editing methods only improve users’ social representations indirectly, and the editing path to the optimal graph structure corresponding to the optimal social representation may be computationally infeasible. This shows that it is necessary to address the low social homophily issue from an innovative generative perspective, which is the key research aim of this paper.
(3) Contrastive Social Denoising Methods. Unlike data-centric methods that modify the structure of social networks to enhance user representations, another line of research focuses on directly improving users’ representations through contrastive learning. These methods typically align user representations from different views (e.g., social or collaborative domains) to achieve a position more consistent with collaborative signals in the latent space. However, contrastive denoising approaches also rely on pre-defined user similarity measurements, which may not fully leverage the supervision signal from collaborative filtering, potentially falling short of reaching the optimal user representations. Additionally, since contrastive denoising methods adjust the training loss function to reduce noise in social networks, they are orthogonal to our proposed framework and can be adaptively integrated. We plan to explore the integration of the proposed SGSR framework with various auxiliary loss functions in future work.
IV-B Diffusion Models for Recommendations
Diffusion models, a powerful class of generative models, have significantly reshaped various recommendation tasks with their effective generation capabilities. In collaborative filtering (CF), DiffRec serves as a fundamental contribution, providing accurate recommendations from a diffusion perspective. Specifically, DiffRec encodes a user’s list of interacted items into a latent vector and employs a forward-reverse process to generate missing user-item interactions in the latent space. The diffused latent vector is then decoded into predicted user-item interactions. Following the proposal of the DiffRec framework, CF-Diff and GiffCF enhanced this approach by incorporating high-order relational capabilities through graph-level diffusion models. Another notable diffusion model, DreamRec , targets sequential recommendation tasks by using a conditional generation paradigm to generate target-item embeddings based on historical interaction sequences. Additionally, leveraging the strong generative ability of diffusion models to produce high-quality samples or representations for recommendation has shown promise in recommendations . For instance, DDRM refines noisy implicit user-item interaction data by applying a diffusion process to user and item representations, conditioned on their initial embeddings. Moreover, other tasks, such as point-of-interest (POI) recommendation and knowledge-aware recommendation , have also benefited from the success of diffusion models. However, existing diffusion recommender models are not designed to integrate denoised social information into recommendations, which differs from the primary research objective of this paper. Additionally, most existing works utilize the Denoising Diffusion Probabilistic Models (DDPM) framework, which can be seen as a specific discretized form of Score-based Generative Models (SGM) , highlighting the novelty of the proposed SGSR framework.
V Conclusion
In this paper, we address the low homophily challenge in social recommendation from a generative perspective using score-based diffusion models. Specifically, we implement an SDE-based diffusion process to produce collaboratively consistent social representations conditioned on user preference patterns. Recognizing that classifier guidance is unsuitable for recommendation contexts, we derive a novel classifier-free optimization objective to facilitate the training of SGMs. To address the lack of supervision signals in recommendation scenarios, we introduce a curriculum learning mechanism and a joint training strategy. Our evaluation of SGSR on three real-world datasets demonstrates significant improvements in recommendation performance compared to existing methods, validating the effectiveness of our social denoising approach.
For future work, a promising direction is to accelerate sampling by designing consistent functions within the recommendation context (training time analysis in Appendix VI). Additionally, exploring various SDE solvers, such as Runge-Kutta methods, could enhance the inference process by leveraging the computational flexibility of score-based diffusion models. Furthermore, while this paper primarily focuses on existing VP-SDEs and VE-SDEs, it is worthwhile to investigate more expressive and suitable SDEs tailored to the recommendation context.