资讯详情

推荐算法——矩阵分解(matrix factorization)

📅 2026/10/3 2:12:41 | 华诺云谱 👁 阅读
推荐算法——矩阵分解(matrix factorization)
Matrix factorization is a core idea in areas like recommender systems.矩阵分解是推荐系统等领域的一个核心概念。一、需解决的问题we have a user-movie matix R:我们有一个“用户-电影”的矩阵 Rusermovie1movie2movie3u151u242u335(In each cell is the users rating for the movie, ? represents an unknown rating)每个格子里是用户对电影的打分为未知打分Objective:Predict the score for the ? part based on the known scores.目标通过已知的打分预测出“?”部分的分数二、解决方法Construct two vector matrices for users and movies respectively, continuously adjust the values of the two sub-matrices so that their inner product multiplication can approximate or even equal the target matrix R.对用户、电影分别构建两个向量矩阵不断调整两个子矩阵的数值让他们同过内积相乘可以接近乃至与目标矩阵R相等。Here we need to introduce a concept:latent factors这里需要引进一个概念潜在因子(latent factors)It refers to the numerical value of a user/movie in a certain dimension. For example:它是指对用户/电影在某个维度的数值举个例子来说电影1 —— 动作片 0.8、爱情片 0.2一个电影可能含有两个潜在因子用户1 —— 动作片 0.9、爱情片 0.1一个user可能含有两个潜在因子——预测评分0.9×0.8 0.1×0.2 0.74That was just an example.上面只是一个例子。For each potential factors value, we need to train the model based on existing ratings to approximate the target matrix R, and predict the values of unknown parts, thereby providing corresponding recommendations for each user.对于每个潜在因子的值我们需要根据已有的评分进行模型训练接近目标矩阵R并且预测计算出未知部分的值从而对每个用户进行对应推荐。将其更标准的写成公式P: User latent features (users × k)P: 用户潜在特征users × kQ: Item latent features (items × k)Q: 物品潜在特征items× kk: number of latent factors k: 潜在因子的个数进而预测矩阵则为预测评分矩阵对于矩阵的每个部分则有三、加入偏置偏置bias模型里用来表示“系统性偏高/偏低”的那一项相当于在预测里加一个“基准分”让模型先把整体趋势解释掉再用隐向量去解释个性化差异。回到我们要解决的问题仅靠点积通常不够因为有些用户天生打分高/低有些电影总体评分高/低全局平均分也有意义常用的预测公式是​​​​​​​全局平均评分global mean用户偏置user bias表示用户整体打分偏高/偏低物品偏置item bias表示物品整体受欢迎程度代码import numpy as np import pandas as pd from sklearn.model_selection import train_test_split#导入用于“划分训练集/测试集”的函数 train_test_split pathrE:\\Desktop\\ml-100k\ml-100k\\u.data colums [user_id, item_id, rating, timestamp] df pd.read_csv(path, sep\t, namescolums) #df存 带行列标签的二维表格 df[user_id] df[user_id] - 1 df[item_id] df[item_id] - 1 #统计user跟item不重复的有多少 num_usersdf[user_id].nunique() num_itemsdf[item_id].nunique() print(users:,num_users,items:,num_items) #对df进行随机划分划分成 训练集测试集 train_df,test_dftrain_test_split(df,test_size0.2,random_state42)#random_state42固定随机种子让每次运行切分结果都一致便于复现和对比实验结果 #构建训练集的R矩阵 Rnp.zeros((num_users,num_items)) #构建全0的R矩阵 for i in range(len(train_df)): #遍历训练集每行 usertrain_df.iloc[i][user_id] #iloc 按位索引 itemtrain_df.iloc[i][item_id] ratingtrain_df.iloc[i][rating] R[int(user),int(item)]rating class MF: def __init__(self,R,k20,alpha0.005,beta0.02,iteration20): #类中的构造函数 self.RR self.num_usersR.shape[0] self.num_itemsR.shape[1] self.kk self.alphaalpha self.betabeta self.iterationiteration #构建训练 def train(self): #self当前这个对象本身 #初始化P、Q矩阵并随机存入0~1之间的浮点数 self.Pnp.random.rand(self.num_users,self.k) self.Qnp.random.rand(self.num_items,self.k) #初始化偏置项bias全0 self.bunp.zeros(self.num_users) self.binp.zeros(self.num_items) #计算平均值b sum_rating0 #存评分总和 count0 #存评分条数 for i in range(self.num_users): for j in range(self.num_items): if self.R[i][j]0: sum_ratingR[i][j] count1 self.bsum_rating/count #训练打印 主体部分 for it in range(self.iteration): for i in range(self.num_users): for j in range(self.num_items): if self.R[i][j]0: #只对有评分的地方训练 self.sgd(i,j) if (it1)%50: print(iteration:,it1,RMSE,self.predict_rmse()) #随机梯度下降Stochastic Gradient Descent关于 用户i对电影j def sgd(self,i,j): predself.predict(i,j) #用当前P、Q、bias的预测i对j的评分预测值:pred eself.R[i][j]-pred #误差e #更新bias self.bu[i]self.bu[i]self.alpha*(e-self.beta*self.bu[i]) self.bi[j]self.bi[j]self.alpha*(e-self.beta*self.bi[j]) #更新对应的P、Q矩阵 for k in range(self.k): p_oldself.P[i][k] self.P[i][k]self.P[i][k]self.alpha*(e*self.Q[j][k]-self.beta*self.P[i][k]) self.Q[j][k]self.Q[j][k]self.alpha*(e*p_old-self.beta*self.Q[j][k]) #用当前训练出出来的参数预测用户i对电影j的评分 def predict(self,i,j): resultself.bself.bu[i]self.bi[i] #基线预测 for k in range(self.k): #在k维下的点积 result self.P[i][k]*self.Q[j][k] return result #构建预测R矩阵 def predict_R(self): resultnp.zeros((self.num_users,self.num_items)) for i in range(self.num_users): for j in range(self.num_items): result[i][j]self.predict(i,j) return result #对于训练集计算rmse输出用直观看是否收敛是否出现数值爆炸 def predict_rmse(self): error0 count0 for i in range(self.num_users): for j in range(self.num_items): if self.R[i][j]0: predself.predict(i,j) error(self.R[i][j]-pred)**2 count1 return np.sqrt(error/count) #对训练集R进行训练 mfMF(R) mf.train() pred_Rmf.predict_R() #训练得到的预测R矩阵 #测试集rmse计算 def rmse(pred_R,test_df): error0 count0 for i in range(len(test_df)): usertest_df.iloc[i][user_id] itemtest_df.iloc[i][item_id] ture_ratingtest_df.iloc[i][rating] pred_ratingpred_R[int(user),int(item)] error(ture_rating-pred_rating)**2 count1 return np.sqrt(error/count) print(\ntest RMSE,rmse(pred_R,test_df))
📝

华诺云谱内容团队

资深建站顾问 · 行业研究员

10年+企业数字化服务经验,专注智能建站、SEO优化与品牌营销,持续输出建站技巧、行业洞察与营销干货,已帮助5000+企业实现数字化增长。

你可能需要的服务

订阅华诺云谱资讯周报

每周一封,精选建站技巧、SEO与营销干货,直达邮箱。已有 8,000+ 企业主订阅,助你少走弯路。

↑