[Bug middle-end/39326] Segmentation fault with -O1, out of memory with -O2

steven at gcc dot gnu.org gcc-bugzilla@gcc.gnu.org
Mon Mar 11 09:40:00 GMT 2013


http://gcc.gnu.org/bugzilla/show_bug.cgi?id=39326

--- Comment #45 from Steven Bosscher <steven at gcc dot gnu.org> 2013-03-11 09:40:18 UTC ---
Patches posted:

* Restrict GIMPLE loop invariant code motion of loop-invariant loads and
stores to loops with fewer memory references than a certain maximum that
is controlled with --param loops-max-datarefs-for-datadeps" from the
command line.
http://gcc.gnu.org/ml/gcc-patches/2013-03/msg00380.html

* Do not create new pseudo-registers for load-after-store transformations
in RTL dead store elimination.  This reduces the memory foot print after
DSE by ~2 percent, and avoids the compile time and memory usage explosion
in combine because it gets presented fewer single-def/single-use register
moves that are really just register copies.
http://gcc.gnu.org/ml/gcc-patches/2013-03/msg00379.html

* Make gcse.c respect -fno-gcse-lm. For the RTL PRE problem, this means
compile time is reasonable with -fno-gcse-lm.
A follow-up patch will implement some mechanism to disable load motion
automatically on extreme test cases like the one from this PR.
http://gcc.gnu.org/ml/gcc-patches/2013-03/msg00386.html


The remaining compile time bottlenecks are:

- RTL dead store limination in its analysis phase.  This is mostly time
spent in dependence tests in alias analysis for instructions in a single
basic block, so it's only a problem for test cases where there is a huge
number of loads and stores in each basic block. I don't think it is worth
speeding up DSE for such extreme cases. 

- Post-reload CSE because it is in the worst-case quadratic in the number 
of instructions in a basic block.  In most practical cases, post-reload 
CSE scales linearly with the number of instructions in a basic block, but
with a large constant bound. It looks up and down through the instruction
chain to see if a reg is not  clobbered between a use and a def.  Because
it only has to do so with  hard registers  the typical bound is closer to
"number of insns in basic block" * "number of hard registers".  This is
fine, I am not going to try and improve this.



More information about the Gcc-bugs mailing list