c# - 我可以给编译器/JIT 什么优化提示?

标签 c# .net vb.net optimization

我已经分析过了,现在我希望从我的热点中挤出所有可能的性能。

我知道 [MethodImplOptions.AggressiveInlining]ProfileOptimization class .还有其他的吗?


[编辑] 我刚刚发现[TargetedPatchingOptOut] 没关系,显然 that one is not needed .

最佳答案

是的,还有更多技巧:-)

实际上,我对优化 C# 代码进行了大量研究。到目前为止,这些是最重要的结果:

  1. 直接传递的 Func 和 Action 通常由 JIT'ter 内联。请注意,您不应将它们存储为变量,因为它们随后会被称为委托(delegate)。另见 this post了解更多详情。
  2. 注意过载。在不使用 IEquatable<T> 的情况下调用 Equals通常是一个糟糕的计划 - 所以如果你使用 f.ex。一个哈希,一定要实现正确的重载和接口(interface),因为它会保证你大量的性能。
  3. 从其他类调用的泛型从不内联。原因是概述的“魔术”here .
  4. 如果您使用数据结构,请确保尝试使用数组来代替:-) 真的,与……相比,这些东西快得要命……好吧,我想的几乎任何东西。通过使用我自己的哈希表和使用数组而不是列表,我已经优化了很多东西。
  5. 在很多情况下,表查找比计算事物或使用 vtable 查找、开关、多个 if 语句甚至计算等构造更快。如果您有分支机构,这也是一个好技巧;失败的分支预测通常会成为一个很大的痛苦。另见 this post - 这是我在 C# 中经常使用的一个技巧,在很多情况下效果很好。哦,查找表当然是数组。
  6. 尝试制作(小型)类结构。由于值类型的性质,结构的某些优化与类的不同。例如,方法调用更简单,因为编译器确切地知道将调用什么方法。此外,结构数组通常比类数组更快,因为它们每次数组操作需要的内存操作少 1 次。
  7. 不要使用多维数组。虽然我更喜欢 Foo[] , 甚至 Foo[][]通常比 Foo[,] 快.
  8. 如果您要复制数据,请在一周中的任何一天使用 Buffer.BlockCopy 而不是 Array.Copy。还要注意字符串:字符串操作可能会消耗性能。

过去也有一个名为“英特尔奔腾处理器优化”的指南,其中包含大量技巧(例如移位或乘法而不是除法)。虽然现在编译器做得很好,但有时这也会有所帮助。

当然这些只是优化;最大的性能提升通常是更改算法和/或数据结构的结果。请务必检查您可以使用哪些选项,不要过多地受 .NET 框架的限制……而且在我自己检查反编译代码之前,我有一种不信任 .NET 实现的自然倾向。 .. 有很多东西本来可以更快地实现的(大多数时候都是有充分理由的)。

HTH


Alex 向我指出 Array.Copy根据某些人的说法,实际上更快。由于我真的不知道这些年来发生了什么变化,我决定唯一正确的做法是创建一个全新的基准并对其进行测试。

如果您只对结果感兴趣,请往下看。在大多数情况下,调用 Buffer.BlockCopy明显优于 Array.Copy .在 .NET 4.5.2 上在具有 16 GB 内存(> 10 GB 空闲)的 Intel Skylake 上测试。

代码:

static void TestNonOverlapped1(int K)
{
    long total = 1000000000;
    long iter = total / K;
    byte[] tmp = new byte[K];
    byte[] tmp2 = new byte[K];
    for (long i = 0; i < iter; ++i)
    {
        Array.Copy(tmp, tmp2, K);
    }
}

static void TestNonOverlapped2(int K)
{
    long total = 1000000000;
    long iter = total / K;
    byte[] tmp = new byte[K];
    byte[] tmp2 = new byte[K];
    for (long i = 0; i < iter; ++i)
    {
        Buffer.BlockCopy(tmp, 0, tmp2, 0, K);
    }
}

static void TestOverlapped1(int K)
{
    long total = 1000000000;
    long iter = total / K;
    byte[] tmp = new byte[K + 16];
    for (long i = 0; i < iter; ++i)
    {
        Array.Copy(tmp, 0, tmp, 16, K);
    }
}

static void TestOverlapped2(int K)
{
    long total = 1000000000;
    long iter = total / K;
    byte[] tmp = new byte[K + 16];
    for (long i = 0; i < iter; ++i)
    {
        Buffer.BlockCopy(tmp, 0, tmp, 16, K);
    }
}

static void Main(string[] args)
{
    for (int i = 0; i < 10; ++i)
    {
        int N = 16 << i;

        Console.WriteLine("Block size: {0} bytes", N);

        Stopwatch sw = Stopwatch.StartNew();

        {
            sw.Restart();
            TestNonOverlapped1(N);

            Console.WriteLine("Non-overlapped Array.Copy: {0:0.00} ms", sw.Elapsed.TotalMilliseconds);
            GC.Collect(GC.MaxGeneration);
            GC.WaitForFullGCComplete();
        }

        {
            sw.Restart();
            TestNonOverlapped2(N);

            Console.WriteLine("Non-overlapped Buffer.BlockCopy: {0:0.00} ms", sw.Elapsed.TotalMilliseconds);
            GC.Collect(GC.MaxGeneration);
            GC.WaitForFullGCComplete();
        }

        {
            sw.Restart();
            TestOverlapped1(N);

            Console.WriteLine("Overlapped Array.Copy: {0:0.00} ms", sw.Elapsed.TotalMilliseconds);
            GC.Collect(GC.MaxGeneration);
            GC.WaitForFullGCComplete();
        }

        {
            sw.Restart();
            TestOverlapped2(N);

            Console.WriteLine("Overlapped Buffer.BlockCopy: {0:0.00} ms", sw.Elapsed.TotalMilliseconds);
            GC.Collect(GC.MaxGeneration);
            GC.WaitForFullGCComplete();
        }

        Console.WriteLine("-------------------------");
    }

    Console.ReadLine();
}

x86 JIT 上的结果:

Block size: 16 bytes
Non-overlapped Array.Copy: 4267.52 ms
Non-overlapped Buffer.BlockCopy: 2887.05 ms
Overlapped Array.Copy: 3305.01 ms
Overlapped Buffer.BlockCopy: 2670.18 ms
-------------------------
Block size: 32 bytes
Non-overlapped Array.Copy: 1327.55 ms
Non-overlapped Buffer.BlockCopy: 763.89 ms
Overlapped Array.Copy: 2334.91 ms
Overlapped Buffer.BlockCopy: 2158.49 ms
-------------------------
Block size: 64 bytes
Non-overlapped Array.Copy: 705.76 ms
Non-overlapped Buffer.BlockCopy: 390.63 ms
Overlapped Array.Copy: 1303.00 ms
Overlapped Buffer.BlockCopy: 1103.89 ms
-------------------------
Block size: 128 bytes
Non-overlapped Array.Copy: 361.18 ms
Non-overlapped Buffer.BlockCopy: 219.77 ms
Overlapped Array.Copy: 620.21 ms
Overlapped Buffer.BlockCopy: 577.20 ms
-------------------------
Block size: 256 bytes
Non-overlapped Array.Copy: 192.92 ms
Non-overlapped Buffer.BlockCopy: 108.71 ms
Overlapped Array.Copy: 347.63 ms
Overlapped Buffer.BlockCopy: 353.40 ms
-------------------------
Block size: 512 bytes
Non-overlapped Array.Copy: 104.69 ms
Non-overlapped Buffer.BlockCopy: 65.65 ms
Overlapped Array.Copy: 211.77 ms
Overlapped Buffer.BlockCopy: 202.94 ms
-------------------------
Block size: 1024 bytes
Non-overlapped Array.Copy: 52.93 ms
Non-overlapped Buffer.BlockCopy: 38.84 ms
Overlapped Array.Copy: 144.39 ms
Overlapped Buffer.BlockCopy: 154.09 ms
-------------------------
Block size: 2048 bytes
Non-overlapped Array.Copy: 45.64 ms
Non-overlapped Buffer.BlockCopy: 30.11 ms
Overlapped Array.Copy: 118.33 ms
Overlapped Buffer.BlockCopy: 109.16 ms
-------------------------
Block size: 4096 bytes
Non-overlapped Array.Copy: 30.93 ms
Non-overlapped Buffer.BlockCopy: 30.72 ms
Overlapped Array.Copy: 119.73 ms
Overlapped Buffer.BlockCopy: 104.66 ms
-------------------------
Block size: 8192 bytes
Non-overlapped Array.Copy: 30.37 ms
Non-overlapped Buffer.BlockCopy: 26.63 ms
Overlapped Array.Copy: 90.46 ms
Overlapped Buffer.BlockCopy: 87.40 ms
-------------------------

x64 JIT 上的结果:

Block size: 16 bytes
Non-overlapped Array.Copy: 1252.71 ms
Non-overlapped Buffer.BlockCopy: 694.34 ms
Overlapped Array.Copy: 701.27 ms
Overlapped Buffer.BlockCopy: 573.34 ms
-------------------------
Block size: 32 bytes
Non-overlapped Array.Copy: 995.47 ms
Non-overlapped Buffer.BlockCopy: 654.70 ms
Overlapped Array.Copy: 398.48 ms
Overlapped Buffer.BlockCopy: 336.86 ms
-------------------------
Block size: 64 bytes
Non-overlapped Array.Copy: 498.86 ms
Non-overlapped Buffer.BlockCopy: 329.15 ms
Overlapped Array.Copy: 218.43 ms
Overlapped Buffer.BlockCopy: 179.95 ms
-------------------------
Block size: 128 bytes
Non-overlapped Array.Copy: 263.00 ms
Non-overlapped Buffer.BlockCopy: 196.71 ms
Overlapped Array.Copy: 137.21 ms
Overlapped Buffer.BlockCopy: 107.02 ms
-------------------------
Block size: 256 bytes
Non-overlapped Array.Copy: 144.31 ms
Non-overlapped Buffer.BlockCopy: 101.23 ms
Overlapped Array.Copy: 85.49 ms
Overlapped Buffer.BlockCopy: 69.30 ms
-------------------------
Block size: 512 bytes
Non-overlapped Array.Copy: 76.76 ms
Non-overlapped Buffer.BlockCopy: 55.31 ms
Overlapped Array.Copy: 61.99 ms
Overlapped Buffer.BlockCopy: 54.06 ms
-------------------------
Block size: 1024 bytes
Non-overlapped Array.Copy: 44.01 ms
Non-overlapped Buffer.BlockCopy: 33.30 ms
Overlapped Array.Copy: 53.13 ms
Overlapped Buffer.BlockCopy: 51.36 ms
-------------------------
Block size: 2048 bytes
Non-overlapped Array.Copy: 27.05 ms
Non-overlapped Buffer.BlockCopy: 25.57 ms
Overlapped Array.Copy: 46.86 ms
Overlapped Buffer.BlockCopy: 47.83 ms
-------------------------
Block size: 4096 bytes
Non-overlapped Array.Copy: 29.11 ms
Non-overlapped Buffer.BlockCopy: 25.12 ms
Overlapped Array.Copy: 45.05 ms
Overlapped Buffer.BlockCopy: 47.84 ms
-------------------------
Block size: 8192 bytes
Non-overlapped Array.Copy: 24.95 ms
Non-overlapped Buffer.BlockCopy: 21.52 ms
Overlapped Array.Copy: 43.81 ms
Overlapped Buffer.BlockCopy: 43.22 ms
-------------------------

关于c# - 我可以给编译器/JIT 什么优化提示?,我们在Stack Overflow上找到一个类似的问题: https://stackoverflow.com/questions/16306206/

相关文章:

c# - ASP.NEt MVC 使用 Web API 返回 Razor View

c# - Windows 操作系统中的 Outlook 签名文件夹

c# - 在 .NET 中使用后将对象设置为 Null/Nothing

javascript - 正则表达式在多行上进行多个匹配

c# - ASP.Net 检查远程文件是否存在

c# - 找不到类型或命名空间名称 'BundleCollection'(是否缺少 using 指令或程序集引用?)

c# - 检查 ArrayList 中是否存在任何值的最佳方法

vb.net - 如何向选定的 DateTimePicker 添加 1 个月

javascript - 从 MVC4 中的 Controller 在我的 Razor 页面中调用 Javascript

c# - 自动属性和结构不混合?