java - 比较 Collectors.summingLong 和 Collectors.counting 的性能

coder 2024-03-07 原文

基准测试在 intel core i5, Ubuntu 下运行

java version "1.8.0_144"
Java(TM) SE Runtime Environment (build 1.8.0_144-b01)
Java HotSpot(TM) 64-Bit Server VM (build 25.144-b01, mixed mode)

我正在比较 Collectors.counting 和 Collectors.summingLong(x -> 1L) 的性能。这是基准:

public List<Integer> ints = new ArrayList<>();

Collector<Integer, ?, Long> counting = Collectors.counting();
Collector<Integer, ?, Long> summingLong = Collectors.summingLong(x -> 1L);

@Setup
public void setup() throws Exception{
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
    ints.add(new Random().nextInt(1000));
}

@Benchmark
@OutputTimeUnit(TimeUnit.NANOSECONDS)
@BenchmarkMode(Mode.AverageTime)
public Long counting() {
    return ints.stream().collect(counting);
}

@Benchmark
@OutputTimeUnit(TimeUnit.NANOSECONDS)
@BenchmarkMode(Mode.AverageTime)
public Long summingLong() {
    return ints.stream().collect(summingLong);
}

我得到的结果是 Collectors.counting Collectors.summingLong 慢了 3 倍。

所以我用 -prof perfnorm 用 25 个 fork 运行它。这是结果:

Benchmark                                        Mode  Cnt    Score     Error  Units
MyBenchmark.counting                             avgt  125   87.115 ±   3.787  ns/op
MyBenchmark.counting:CPI                         avgt   25    0.310 ±   0.011   #/op
MyBenchmark.counting:L1-dcache-load-misses       avgt   25    1.963 ±   0.171   #/op
MyBenchmark.counting:L1-dcache-loads             avgt   25  258.716 ±  41.458   #/op
MyBenchmark.counting:L1-dcache-store-misses      avgt   25    1.890 ±   0.168   #/op
MyBenchmark.counting:L1-dcache-stores            avgt   25  131.344 ±  16.433   #/op
MyBenchmark.counting:L1-icache-load-misses       avgt   25    0.035 ±   0.007   #/op
MyBenchmark.counting:LLC-loads                   avgt   25    0.411 ±   0.143   #/op
MyBenchmark.counting:LLC-stores                  avgt   25    0.007 ±   0.001   #/op
MyBenchmark.counting:branch-misses               avgt   25    0.029 ±   0.017   #/op
MyBenchmark.counting:branches                    avgt   25  139.669 ±  21.901   #/op
MyBenchmark.counting:cycles                      avgt   25  261.967 ±  29.392   #/op
MyBenchmark.counting:dTLB-load-misses            avgt   25    0.036 ±   0.003   #/op
MyBenchmark.counting:dTLB-loads                  avgt   25  258.111 ±  41.008   #/op
MyBenchmark.counting:dTLB-store-misses           avgt   25    0.001 ±   0.001   #/op
MyBenchmark.counting:dTLB-stores                 avgt   25  131.394 ±  16.166   #/op
MyBenchmark.counting:iTLB-load-misses            avgt   25    0.001 ±   0.001   #/op
MyBenchmark.counting:iTLB-loads                  avgt   25    0.001 ±   0.001   #/op
MyBenchmark.counting:instructions                avgt   25  850.262 ± 113.228   #/op
MyBenchmark.counting:stalled-cycles-frontend     avgt   25   48.493 ±   8.968   #/op
MyBenchmark.summingLong                          avgt  125   37.238 ±   0.194  ns/op
MyBenchmark.summingLong:CPI                      avgt   25    0.311 ±   0.002   #/op
MyBenchmark.summingLong:L1-dcache-load-misses    avgt   25    1.793 ±   0.013   #/op
MyBenchmark.summingLong:L1-dcache-loads          avgt   25   93.785 ±   0.640   #/op
MyBenchmark.summingLong:L1-dcache-store-misses   avgt   25    1.727 ±   0.013   #/op
MyBenchmark.summingLong:L1-dcache-stores         avgt   25   56.249 ±   0.408   #/op
MyBenchmark.summingLong:L1-icache-load-misses    avgt   25    0.020 ±   0.003   #/op
MyBenchmark.summingLong:LLC-loads                avgt   25    0.843 ±   0.117   #/op
MyBenchmark.summingLong:LLC-stores               avgt   25    0.004 ±   0.001   #/op
MyBenchmark.summingLong:branch-misses            avgt   25    0.008 ±   0.002   #/op
MyBenchmark.summingLong:branches                 avgt   25   61.472 ±   0.260   #/op
MyBenchmark.summingLong:cycles                   avgt   25  110.949 ±   0.784   #/op
MyBenchmark.summingLong:dTLB-load-misses         avgt   25    0.031 ±   0.001   #/op
MyBenchmark.summingLong:dTLB-loads               avgt   25   93.662 ±   0.616   #/op
MyBenchmark.summingLong:dTLB-store-misses        avgt   25   ≈ 10⁻³             #/op
MyBenchmark.summingLong:dTLB-stores              avgt   25   56.302 ±   0.351   #/op
MyBenchmark.summingLong:iTLB-load-misses         avgt   25    0.001 ±   0.001   #/op
MyBenchmark.summingLong:iTLB-loads               avgt   25   ≈ 10⁻³             #/op
MyBenchmark.summingLong:instructions             avgt   25  357.029 ±   1.712   #/op
MyBenchmark.summingLong:stalled-cycles-frontend  avgt   25   10.074 ±   1.096   #/op

我注意到的是:

分支、说明、周期 几乎相差 3 倍。还有缓存操作。分支似乎预测得很好，也没有太多缓存未命中(仅我的意见)。

所以问题可能出在编译代码上。使用 -prof perfasm 运行它(太长了，不能放在这里)。

在编译后的代码中，我注意到以下内容:

I. Collectors.summingLong assembly

我们有 3 个循环遍历数组和计数。首先只计算一个元素

0x00007f9abd226dfd: mov %edi,%ebp ;contains the iteration index
incl %ebp
;...
0x00007f9abd226e27: incl %edi
0x00007f9abd226e29: cmp %ebp,%edi
0x00007f9abd226e2b: jnl 0x7f9abd226e34 ;exit after the first iteration

第二次计算 4 个元素进行 1 次迭代(循环展开了吗？)并且在第一次之后退出。

0x00007f9abd226ea6: add $0x1,%rsi 
;...
0x00007f9abd226ed0: add $0x2,%rsi
;...
0x00007f9abd226efa: add $0x3,%rsi
;...
0x00007f9abd226f1c: add $0x4,%rbx
;...
0x00007f9abd226f20: mov %rbx,0x10(%r14)

第三个计算其余元素。

II。 Collectors.counting assembly

我们只有一个循环来逐个计算所有元素(未展开)。我们还在计数结果的循环内进行了内联装箱转换。此外，我们似乎没有在循环内内联装箱转换

0x00007f80dd22dfb5: mov $0x1,%esi
0x00007f80dd22dfba: nop
0x00007f80dd22dfbb: callq 0x7f80dd046420

这似乎执行 lambda e -> 1L 中提供的 1L 的装箱。但这并不清楚为什么。在执行实际添加时，我们有以下代码:

0x00007f80dd22dfec: mov $0x1,%r10d
0x00007f80dd22dff2: add 0x10(%r11),%r10

此外，我们将计数结果存储在堆栈 mov %r10d,0x10(%rsp) 中，而不是像第一种情况那样存储在堆中。

如果我对正在发生的事情的理解是正确的，我有

问题: 使用装箱转换展开的循环是否导致了 3 倍的减速？如果是这样，为什么运行时不展开 counting 情况下的循环？

请注意 Collectors.summingLong 的 GC 压力比 Collectors.counting 高 2.5。这不是很清楚(我只能猜测我们将中间值存储在 Collectors.counting 的堆栈中)。

MyBenchmark.counting                                      avgt    5    96.956 ±   4.412   ns/op
MyBenchmark.counting:·gc.alloc.rate                       avgt    5   734.538 ±  33.083  MB/sec
MyBenchmark.counting:·gc.alloc.rate.norm                  avgt    5   112.000 ±   0.001    B/op
MyBenchmark.counting:·gc.churn.PS_Eden_Space              avgt    5   731.423 ± 340.767  MB/sec
MyBenchmark.counting:·gc.churn.PS_Eden_Space.norm         avgt    5   111.451 ±  48.411    B/op
MyBenchmark.counting:·gc.churn.PS_Survivor_Space          avgt    5     0.017 ±   0.067  MB/sec
MyBenchmark.counting:·gc.churn.PS_Survivor_Space.norm     avgt    5     0.003 ±   0.010    B/op
MyBenchmark.counting:·gc.count                            avgt    5    16.000            counts
MyBenchmark.counting:·gc.time                             avgt    5    12.000                ms
MyBenchmark.summingLong                                   avgt    5    38.371 ±   1.733   ns/op
MyBenchmark.summingLong:·gc.alloc.rate                    avgt    5  1856.581 ±  81.706  MB/sec
MyBenchmark.summingLong:·gc.alloc.rate.norm               avgt    5   112.000 ±   0.001    B/op
MyBenchmark.summingLong:·gc.churn.PS_Eden_Space           avgt    5  1876.736 ± 192.503  MB/sec
MyBenchmark.summingLong:·gc.churn.PS_Eden_Space.norm      avgt    5   113.213 ±   9.916    B/op
MyBenchmark.summingLong:·gc.churn.PS_Survivor_Space       avgt    5     0.033 ±   0.072  MB/sec
MyBenchmark.summingLong:·gc.churn.PS_Survivor_Space.norm  avgt    5     0.002 ±   0.004    B/op
MyBenchmark.summingLong:·gc.count                         avgt    5    62.000            counts
MyBenchmark.summingLong:·gc.time                          avgt    5    48.000                ms

最佳答案

我没有查看程序集或对其进行分析，但查看源代码已经提供了一些信息。

summingLong() 这个收集器的结果:

new CollectorImpl<>(
            () -> new long[1],
            (a, t) -> { a[0] += mapper.applyAsLong(t); },
            (a, b) -> { a[0] += b[0]; return a; },
            a -> a[0], CH_NOID);

counting() 结果如下:

new CollectorImpl<>(
            boxSupplier(identity),
            (a, t) -> { a[0] = op.apply(a[0], mapper.apply(t)); },
            (a, b) -> { a[0] = op.apply(a[0], b[0]); return a; },
            a -> a[0], CH_NOID);

如您所见，counting() 版本还做了一些事情:

拳击
调用 op.apply(...)

由于 op 是 Long::sum，它对基元进行操作，因此涉及大量装箱和拆箱操作。

关于java - 比较 Collectors.summingLong 和 Collectors.counting 的性能，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/48521066/

有关java - 比较 Collectors.summingLong 和 Collectors.counting 的性能的更多相关文章

ruby - Ruby 的 Hash 在比较键时使用哪种相等性测试？ - 2
我有一个围绕一些对象的包装类，我想将这些对象用作散列中的键。包装对象和解包装对象应映射到相同的键。一个简单的例子是这样的:classAattr_reader:xdefinitialize(inner)@inner=innerenddefx;@inner.x;enddef==(other)@inner.x==other.xendenda=A.new(o)#oisjustanyobjectthatallowso.xb=A.new(o)h={a=>5}ph[a]#5ph[b]#nil,shouldbe5ph[o]#nil,shouldbe5我试过==、===、eq?并散列所有无济于事。
java - 等价于 Java 中的 Ruby Hash - 2
我真的很习惯使用Ruby编写以下代码:my_hash={}my_hash['test']=1Java中对应的数据结构是什么？最佳答案 HashMapmap=newHashMap();map.put("test",1);我假设？关于java-等价于Java中的RubyHash，我们在StackOverflow上找到一个类似的问题： https://stackoverflow.com/questions/22737685/
java - 从 JRuby 调用 Java 类的问题 - 2
我正在尝试使用boilerpipe来自JRuby。我看过guide从JRuby调用Java，并成功地将它与另一个Java包一起使用，但无法弄清楚为什么同样的东西不能用于boilerpipe。我正在尝试基本上从JRuby中执行与此Java等效的操作:URLurl=newURL("http://www.example.com/some-location/index.html");Stringtext=ArticleExtractor.INSTANCE.getText(url);在JRuby中试过这个:require'java'url=java.net.URL.new("http://www
java - 我的模型类或其他类中应该有逻辑吗 - 2
我只想对我一直在思考的这个问题有其他意见，例如我有classuser_controller和classuserclassUserattr_accessor:name,:usernameendclassUserController//dosomethingaboutanythingaboutusersend问题是我的User类中是否应该有逻辑user=User.newuser.do_something(user1)oritshouldbeuser_controller=UserController.newuser_controller.do_something(user1,user2)我
java - 什么相当于 ruby 的 rack 或 python 的 Java wsgi？ - 2
什么是ruby的rack或python的Java的wsgi？还有一个路由库。最佳答案来自Python标准PEP333:Bycontrast,althoughJavahasjustasmanywebapplicationframeworksavailable,Java's"servlet"APImakesitpossibleforapplicationswrittenwithanyJavawebapplicationframeworktoruninanywebserverthatsupportstheservletAPI.ht
Observability：从零开始创建 Java 微服务并监控它（二） - 2
这篇文章是继上一篇文章“Observability：从零开始创建Java微服务并监控它（一）”的续篇。在上一篇文章中，我们讲述了如何创建一个Javaweb应用，并使用Filebeat来收集应用所生成的日志。在今天的文章中，我来详述如何收集应用的指标，使用APM来监控应用并监督web服务的在线情况。源码可以在地址 https://github.com/liu-xiao-guo/java_observability 进行下载。摄入指标指标被视为可以随时更改的时间点值。当前请求的数量可以改变任何毫秒。你可能有1000个请求的峰值，然后一切都回到一个请求。这也意味着这些指标可能不准确，你还想提取最小/
【Java 面试合集】HashMap中为什么引入红黑树，而不是AVL树呢 - 2
HashMap中为什么引入红黑树，而不是AVL树呢1.概述开始学习这个知识点之前我们需要知道，在JDK1.8以及之前，针对HashMap有什么不同。JDK1.7的时候，HashMap的底层实现是数组+链表JDK1.8的时候，HashMap的底层实现是数组+链表+红黑树我们要思考一个问题，为什么要从链表转为红黑树呢。首先先让我们了解下链表有什么不好？？？2.链表上述的截图其实就是链表的结构，我们来看下链表的增删改查的时间复杂度增：因为链表不是线性结构，所以每次添加的时候，只需要移动一个节点，所以可以理解为复杂度是N(1)删：算法时间复杂度跟增保持一致查：既然是非线性结构，所以查询某一个节点的时候
【Java入门】使用Java实现文件夹的遍历 - 2
遍历文件夹我们通常是使用递归进行操作，这种方式比较简单，也比较容易理解。本文为大家介绍另一种不使用递归的方式，由于没有使用递归，只用到了循环和集合，所以效率更高一些！一、使用递归遍历文件夹整体思路1、使用File封装初始目录，2、打印这个目录3、获取这个目录下所有的子文件和子目录的数组。4、遍历这个数组，取出每个File对象4-1、如果File是否是一个文件，打印4-2、否则就是一个目录，递归调用代码实现publicclassSearchFile{publicstaticvoidmain(String[]args){//初始目录Filedir=newFile("d:/Dev");Datebeg
java - 为什么 ruby modulo 与 java/other lang 不同？ - 2
我基本上来自Java背景并且努力理解Ruby中的模运算。(5%3)(-5%3)(5%-3)(-5%-3)Java中的上述操作产生，2个-22个-2但在Ruby中，相同的表达式会产生21个-1-2.Ruby在逻辑上有多擅长这个？模块操作在Ruby中是如何实现的？如果将同一个操作定义为一个web服务，两个服务如何匹配逻辑。最佳答案在Java中，模运算的结果与被除数的符号相同。在Ruby中，它与除数的符号相同。remainder()在Ruby中与被除数的符号相同。您可能还想引用modulooperation.
java - Ruby 相当于 Java 的 Collections.unmodifiableList 和 Collections.unmodifiableMap - 2
Java的Collections.unmodifiableList和Collections.unmodifiableMap在Ruby标准API中是否有等价物？最佳答案使用freeze应用程序接口(interface):Preventsfurthermodificationstoobj.ARuntimeErrorwillberaisedifmodificationisattempted.Thereisnowaytounfreezeafrozenobject.SeealsoObject#frozen?.Thismethodretur

java - 比较 Collectors.summingLong 和 Collectors.counting 的性能

有关java - 比较 Collectors.summingLong 和 Collectors.counting 的性能的更多相关文章

随机推荐