Slow profile updating (pr 15524)

Jeffrey A Law law@redhat.com
Sat Nov 20 00:41:00 GMT 2004


Believe it or not we're at a point where updating of the profile in
response to a jump thread is the most expensive routine in the 
compiler for PR 15524.


If we look up update_bb_profile_for_threading we see one loop of
significance:

  else
    FOR_EACH_EDGE (c, ei, bb->succs)
      c->probability = ((c->probability * REG_BR_PROB_BASE) / (double)
prob);


So anytime we thread through some block BB, we have to walk through
all its successors to rescale their probabilities.  Needless to say
that gets rather expensive, especially if BB has a large switch
statement and several of its incoming edges are threadable.

Is it the case that the computation of c->probability actually
has to happen in the order you specified via the parenthesis?  If
not, then we could precompute REG_BR_PROB_BASE / (double) prob
outside the loop which would result in a loop like

    tmp = REG_BR_PROB_BASE / (double) prob;
    FOR_EACH_EDGE (c, ei, bb->succs)
      c->probability *= tmp;

Which would probably provide a reasonable improvement.

And if that's safe, then we'd probably want to rewrite it like

  else if (prob != REG_BR_PROB_BASE)
    {
      double tmp = REG_BR_PROB_BASE / (double) prob;
      FOR_EACH_EDGE 9c, ei, bb->succs)
        c->probability *= tmp;
    }

Which avoids the loop completely if nothing is going to change (as is
the case for pr15524).

Doing something like this would give us a net improvement of 15-20% on
PR 15524.

Alternately we might want to look at whether or not we can rescale the
successor blocks en-masse after all the redirections for a particular
block are complete.

Thoughts?

jeff



More information about the Gcc mailing list