r - 使用因子和数字预测变量执行套索正则化?

标签 r machine-learning feature-selection glmnet regularized

我有一个数据集,我想在其中执行套索以消除特征。由于我是 R 新手,我目前正在遵循 R 在线指南。数据存储在数据框中。目标已从数据帧中删除,并存储在其自己的单列数据帧中。这是一个回归问题,目标是数字。这是我尝试运行的代码:

library(glmnet)

lasso_model <- cv.glmnet(
                  x = as.matrix(train),
                  y = train_target,
                  alpha = 1)

以下是有关数据集的信息:

'data.frame':   9798 obs. of  55 variables:
$ acres: num  0.186 2.991 0.144 0.218 0.173 ...
$ above: int  1754 3030 1531 834 1022 1528 768 1184 2026 3176 ...
$ basement: int  0 1811 500 440 0 476 0 0 732 0 ...
$ baths: Factor w/ 7 levels "0","1","2","3",..: 3 4 3 3 2 3 2 2 3 3 ...
$ toilets: Factor w/ 5 levels "0","1","2","3",..: 1 3 2 1 1 2 1 1 2 2    ...
$ fireplaces: Factor w/ 6 levels "0","1","2","3",..: 2 2 2 2 1 1 1 2 2  2 ...
$ beds: Factor w/ 7 levels "1","2","3","4",..: 4 5 2 2 2 3 2 2 3 5 ...
$ rooms: Factor w/ 15 levels "0","1","2","3",..: 5 5 5 4 5 3 3 3 4 6 ...
$ age: int  103 17 13 46 116 12 93 93 42 100 ...
$ yearsfromsale: Factor w/ 3 levels "2","3","4": 2 2 2 1 2 2 3 3 1 1 ...
$ car: Factor w/ 4 levels "0","1","2","3": 1 4 3 1 1 3 1 1 4 1 ...
$ city_DES.MOINES: Factor w/ 2 levels "0","1": 2 1 1 2 2 1 2 2 2 2 ...
$ city_JOHNSTON: Factor w/ 2 levels "0","1": 1 2 2 1 1 1 1 1 1 1 ...
$ city_WEST.DES.MOINES: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ city_CLIVE: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ city_URBANDALE: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ city_ALTOONA: Factor w/ 2 levels "0","1": 1 1 1 1 1 2 1 1 1 1 ...
$ city_BONDURANT: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ city_CROCKER.TWNSHP: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ city_GRIMES: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ city_POLK.CITY: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ city_PLEASANT.HILL: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ city_WINDSOR.HEIGHTS: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50315: Factor w/ 2 levels "0","1": 1 1 1 2 1 1 1 1 1 1 ...
$ zip_50321: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 2 1 ...
$ zip_50320: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50312: Factor w/ 2 levels "0","1": 2 1 1 1 1 1 1 1 1 2 ...
$ zip_50314: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50311: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50309: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50316: Factor w/ 2 levels "0","1": 1 1 1 1 2 1 1 2 1 1 ...
$ zip_50317: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50313: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 2 1 1 1 ...
$ zip_50310: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50322: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50131: Factor w/ 2 levels "0","1": 1 1 2 1 1 1 1 1 1 1 ...
$ zip_50111: Factor w/ 2 levels "0","1": 1 2 1 1 1 1 1 1 1 1 ...
$ zip_50265: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50266: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50325: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50323: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50009: Factor w/ 2 levels "0","1": 1 1 1 1 1 2 1 1 1 1 ...
$ zip_50035: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50023: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50226: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50021: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50327: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ zip_50324: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 1 1 1 ...
$ walkout_0: Factor w/ 2 levels "0","1": 2 1 2 2 2 1 2 2 2 2 ...
$ walkout_1: Factor w/ 2 levels "0","1": 1 2 1 1 1 2 1 1 1 1 ...
$ condition_Normal: Factor w/ 2 levels "0","1": 1 2 2 1 1 2 1 1 1 1 ...
$ condition_Above.Normal: Factor w/ 2 levels "0","1": 2 1 1 2 2 1 2 1 1 2 ...
$ condition_Below.Normal: Factor w/ 2 levels "0","1": 1 1 1 1 1 1 1 2 2 1 ...
$ AC_1: Factor w/ 2 levels "0","1": 2 2 2 2 2 2 2 2 2 1 ...

当尝试运行 lasso_model 行时,这是我收到的错误:

Error in cbind2(1, newx) %*% nbeta : 
invalid class 'NA' to  dup_mMatrix_as_dgeMatrix

本质上,我希望能够确定要删除哪些变量。任何帮助都会很棒!

最佳答案

好吧,这是一个强烈的怀疑。

您的数据框中有因素。 as.matrix 将它们转换为字符串,而不是数字,并且 glmnet 不知道如何处理它们:

> df <- data.frame(a=as.factor(c('0', '1', '2')), b=as.factor(c('0', '0', '1')))
> df
  a b
1 0 0
2 1 0
3 2 1
> as.matrix(df)
     a   b  
[1,] "0" "0"
[2,] "1" "0"
[3,] "2" "1"

尝试将它们显式转换回数字(有点迂回,但应该可行):

> as.matrix(data.frame(lapply(df, function(x) as.numeric(as.character(x)))))
     a b
[1,] 0 0
[2,] 1 0
[3,] 2 1

关于r - 使用因子和数字预测变量执行套索正则化?,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/52245719/

相关文章:

r - 如何在 r 中绘制 3D 函数?

algorithm - 人工智能方法检测游戏中的作弊行为

python - 属性错误: module 'tensorflow_docs' has no attribute 'plots'

r - 将每个矩阵行与向量进行比较

r - 区分向量和 R 中的矩阵

r - 如何在图表下方添加文本到ggplot?

python - numpy 数组输入到 tensorflow/keras 神经网络的 dtype 有什么关系吗?

machine-learning - 对 TF 和 TF*IDF 向量执行 Chi-2 特征选择

python - 使用 tsfresh 仅选择一定数量的顶级功能

python - 使用 chi2 测试进行具有连续特征的特征选择 (Scikit Learn)