algorithm - 什么是分层 Bootstrap ？

标签 algorithm machine-learning data-mining

我学会了引导和分层。但什么是分层引导？它是如何工作的？

假设我们有一个包含 n 个实例(观察)的数据集，m 是类的数量。数据集应该怎么划分，训练和测试的百分比是多少？

最佳答案

您按类拆分数据集。之后，您独立地从每个子群体中抽样。您从一个子群体中抽样的实例数量应该与其比例相关。

 data
 d(i) <- { x in data | class(x) =i }
 for each class
    for j = 0..samplesize*(size(d(i))/size(data))
       sample(i) <- draw element from d(i)
 sample <- U sample(i)

如果您从类为 {'a', 'a', 'a', 'a', 'a', 'a', 'b', 'b'} 的数据集中采样四个元素，此过程确保分层样本中至少包含 b 类的一个元素。

关于algorithm - 什么是分层 Bootstrap ？，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/35328198/

上一篇：algorithm - O(N) 速度和 O(1) 内存的汉明数

下一篇：algorithm - Graph 的 Adjacency List 表示的空间复杂度

相关文章：

python - 如何在Python中迭代列表列表时标记编码

scala - Spark mllib LinearRegression 奇怪的结果

algorithm - 应该使用哪种排序算法在 O(n) 中运行？

c# - 哪种算法更适合在数据库中隐藏重要字符串？ (该字符串不是密码)

algorithm - 这里使用哪种聚类算法？

machine-learning - 聚类算法的性能分析

machine-learning - 新手: where to start given a problem to predict future success or not

twitter - 预测 Twitter 上 future 推文的情绪

algorithm - 给定一个自然数 A，我想找到所有自然数对 (B,C) 以便 B*C*(C+1) = A

algorithm - 将 Big-O 递归算法简化为线性