是否超出"GC开销限制"是失败的次要原因?

And*_*niy 10 java garbage-collection jvm heap-memory out-of-memory

根据此问题的动机:错误java.lang.OutOfMemoryError:超出了GC开销限制

最近我和某人就这个错误进行了辩论.

根据我的理解,这个错误本身不能被视为JVM失败的"主要"原因.

我的意思是广泛的垃圾收集本身不是失败的原因.广泛的垃圾收集总是由很少的可用内存量引起,这导致频繁的GC调用(并且核心原因可能是内存泄漏).

如果我正确地理解了我的对手的位置,他认为很多有资格在系统中产生GC的小对象会导致他们经常收集,导致这个错误的原因.所以魔鬼不是内存泄漏或低内存限制,而是GC调用频率本身.

这是我们有不同观点的地方.

根据我的理解,有多少小符合GC过程的小对象产生(即使它不是一个好的设计,可能你应尽可能减少这个数量).如果你有足够的内存,并且没有明显的内存泄漏,那么在某些时候GC会收集这些对象的大部分,所以它应该不是问题.至少这不会导致系统崩溃.

简要回顾一下我的立场:如果你有GC overhead limit exceeded,那么要么你有一种内存泄漏,要么你只需要增加内存限制.

简要回顾一下我的对手的位置:如果你生产了许多符合GC的资格的小物件,那已经成为一个问题,因为它本身就是一个问题GC overhead limit exceeded.

我错了,错过了什么?

Ale*_*iez 13

- 部分答案 -

请注意,我使用OpenJDK(JDK 9)源作为评论此问题的基础.这个答案并不依赖于任何类型的文档或已发布的规范,并且包含了一些来自我对源代码的理解和解释的推测.

所述GC overhead limit exceeded在VM被认为是内存不足的错误的子类型和产生之后,存储器分配的尝试失败(参照(A)).

本质上,VM会跟踪完整垃圾收集的发生次数,并将其与完整GC(可在Hotspot上配置-XX:GCTimeLimit=,参见Garbage Collector Ergonomics)强制实施的限制进行比较.

跟踪完整GC计数的实现以及检测到GC开销限制时的逻辑可以在一个地方获得hotspot/src/share/vm/gc/shared/adaptiveSizePolicy.cpp.如您所见,旧的和伊甸园代中可用内存的两个附加条件是满足GC开销限制的标准:

void AdaptiveSizePolicy::check_gc_overhead_limit(
                                      size_t young_live,
                                      size_t eden_live,
                                      size_t max_old_gen_size,
                                      size_t max_eden_size,
                                      bool   is_full_gc,
                                      GCCause::Cause gc_cause,
                                      CollectorPolicy* collector_policy) {
  ...
  if (is_full_gc) {
    if (gc_cost() > gc_cost_limit &&
      free_in_old_gen < (size_t) mem_free_old_limit &&
      free_in_eden < (size_t) mem_free_eden_limit) {
      // Collections, on average, are taking too much time, and
      //      gc_cost() > gc_cost_limit
      // we have too little space available after a full gc.
      //      total_free_limit < mem_free_limit
      // where
      //   total_free_limit is the free space available in
      //     both generations
      //   total_mem is the total space available for allocation
      //     in both generations (survivor spaces are not included
      //     just as they are not included in eden_limit).
      //   mem_free_limit is a fraction of total_mem judged to be an
      //     acceptable amount that is still unused.
      // The heap can ask for the value of this variable when deciding
      // whether to thrown an OutOfMemory error.
      // Note that the gc time limit test only works for the collections
      // of the young gen + tenured gen and not for collections of the
      // permanent gen.  That is because the calculation of the space
      // freed by the collection is the free space in the young gen +
      // tenured gen.
      // At this point the GC overhead limit is being exceeded.
      inc_gc_overhead_limit_count();
      if (UseGCOverheadLimit) {
        if (gc_overhead_limit_count() >= AdaptiveSizePolicyGCTimeLimitThreshold){
          // All conditions have been met for throwing an out-of-memory
          set_gc_overhead_limit_exceeded(true);
          // Avoid consecutive OOM due to the gc time limit by resetting
          // the counter.
          reset_gc_overhead_limit_count();
      } else {
        ...
      }
Run Code Online (Sandbox Code Playgroud)

(a)何时GC overhead limit exceeded产生错误?

它实际上不会在集合本身期间发生,但是当VM尝试分配内存时 - 您可以在以下位置找到这些语句的理由hotspot/src/share/vm/gc/shared/collectedHeap.inline.hpp:

HeapWord* CollectedHeap::common_mem_allocate_noinit(KlassHandle klass, size_t size, TRAPS) {
    ...
    bool gc_overhead_limit_was_exceeded = false;
    result = Universe::heap()->mem_allocate(size, &gc_overhead_limit_was_exceeded);
    ...
    // Failure cases
   if (!gc_overhead_limit_was_exceeded) {
       report_java_out_of_memory("Java heap space");
       ...
    } else {
       report_java_out_of_memory("GC overhead limit exceeded");
       ...
    }
Run Code Online (Sandbox Code Playgroud)

(b)关于G1实施的说明

看一下mem_allocateG1实现的方法(可以在其中找到g1CollectedHeap.cpp),看起来布尔值gc_overhead_limit_was_exceeded不再使用了.如果G1 GC已启用,我不会太快得出结论GC内存开销错误不再发生 - 我需要检查一下.

结论

  • 看来你是对的,因为这个错误真的来自记忆耗尽;

  • 可以根据收集小对象的次数生成此错误的论点对我来说似乎不对,因为

    1. 我们看到VM确实需要耗尽内存才能发生此错误;
    2. 独立于第一个原因,我们还需要进一步完善声明 - 尤其是对小对象的引用.我们只谈论年轻一代的收藏吗?如果是这样,这些集合不包括在根据限制检查的GC计数中,因此永远不会有机会参与此错误,VM是否运行OOM.