Notes from the version control BOF at the summit

Tom Lord lord@emf.net
Mon Jun 14 21:09:00 GMT 2004


    > From: Florian Weimer <fw@deneb.enyo.de>

    > > CVS makes easy a centralized model and makes _barely_ possible a
    > > decentralized model.  Arch makes both about equally easy (archive
    > > boundaries are largely transparent in arch).

    > This is a strong claim, but I think it isn't quite true (only if you
    > are accustomed to all the little annoyances and are no longer feeling
    > the pain, so to speak).

I think you contradict yourself.  If we assume that everything you say
in your parentheses is true, then that means that Arch makes the
decentralized model easy _modulo_ having to get over some or many
little annoyances.   A finite sum of little annoyances is a decent
approximation of 0 so I think I'm right when I say that arch makes
both cases "about as easy".

Of course, two things:  (1) We've been and will continue to be killing off
"little annoyances", one by one, until they are gone.  (2) I think
we're pretty far along that path and some of what is starting to annoy
you is just the underlying, fundamental, inescapable stuff -- you're
annoyed at water for being wet and at the wind for being made of air
in coherent motion.


    >> The current first generation of Arch users, having an unbiased
    >> choice, have collectively decided that decentralized is usually the
    >> best thing.

    > Centralized vs decentralized is not just a technical distinction
    > (given enough effort, you can turn one tool that favors one choice
    > into something for the other).  It can also ave great impact on
    > development.

Which was my point.   Arch made decentralized easier and, perhaps
surprisingly, people embraced that.   Very, very, very early on it was
far more common for people contributing to arch to send plain 'ol
regular diffs.   Nowadays, for the most part, _nobody_bothers_ because
it's easier just to use arch.


    > I've already seen two projects which declined considerably after
    > version control software was introduced.  Previously, distributed
    > development was implemented by ad-hoc patch submission, and frequent
    > releases synchronized everyone with the current state of development
    > and provided clear reference points.  With version control, regular
    > releases have almost stopped (and security fixes suddenly take a very
    > long time for one of the projects).  This is not tied to the type of
    > version control, by the way; one project is centralized, one is
    > extremely distributed (de facto, it even lacks an official release
    > tree).

This is superstitious nonsense.   

Perhaps the maintainers' of your projects are exercising a new-found
degree of freedom that revision control provided them.   Perhaps they
don't think that the security fixes you have in mind are actually a
priority and, because they have revision control, there is no rush to
make a release after integrating it because co-developers have access
to the fix.

Perhaps these projects haven't "declined" at all but have instead
become more efficient and effective in ways you haven't yet
recognized.



    > The GCC Release Manager will continue to do an excellent job, 

Here, here.  (From my perspective, GCC RM has become just that sort of
job:   to do it, you have to play the role of "boy/girl scout" -- just
do everything you do Right and always be eagerly proactive about
making positive progress.   In return, the rest of us will gratefully
cheer you on and remark on your Virtues.)

    > but I fear that a drastic change in the underlying revision
    > control system could make the developers' and even his much more
    > complicated.

Oh, yeah.  GCC is the hardest (and most worthwhile) project-win in
town.   I like it that way.

On the one hand, GCC could surely benefit, medium and longer term,
from project infrastructure changes.  On the other hand, it's one of
the tensest projects on the planet when it comes to weighing the
short-term trade-offs of any such changes.  So it's a _hard_ choice.

Y'all are fascinating.  Superlatively great usage of CVS -- extreme
demonstration of professionalism in all of its boringly tedious
detail.  Very critical project.  Very large project with many
contributors.  High commit rate.  Lots of established project
infrastructure.  GCC is one of the Everests of the current revision
control competition.

I hope you don't blow it by whoring yourself out to this or that
revctl project out of some process of political gladhanding.   Your
track record as a demonstration of successful engineering process is a
book of world records and it would be a shame to tarnish that.

I've tried to be very clear about my opinion: that arch is _nearly_
there for you and that it is so close that the best way to close the
gap may be to start trying to use it in isolation from the GCC
mainline, and learn from experience what more needs to be fixed.  And,
at the same time, arch is light years ahead of all of the alternatives
you might consider.  Some of those alternatives represent design
choices that arch rejected for very good reasons.  Others of those
alternatives represent clever ideas that can be (in some cases have
been) added to arch --- the other system being built around those
ideas but neglecting many other important aspects that arch covers.
Arch basically rocks, especially when regarded as a foundation for
moving forward.


    > > I don't understand what you mean by "versioned branch creation".
    > > Do you think it's important enough to elaborate on?

    > The archive does not reflect which branches are in active development
    > and which aren't.  I think this might be a usability issue and leads
    > to hacks such as periodic archive deprecation (e.g. a switch from
    > lord@emf.net--2003b to lord@emf.net--2004).

You're talking about a presentation issue.   If I ask to see which
branches there are, arch will currently, unconditionally, show me all
of them.   You wish it would show just a subset of "active branches".

Yes, indeed, those kinds of fine-points of user interface are one of
two active focusses of current arch development.  It's a sign of
arch's winningness that the remaining issues have been boiled down to
such minor matters.


    > > Regarding "push-mirror" --- at the cost of a tiny amount of work, plus
    > > the cost that the admin of the mirrored arch agrees to run an rsync
    > > server -- you can avoid "push-mirror" entirely and mirror just using
    > > rsync.   It'll be much faster.

    > I know.  A strength of arch is the availability of tolerable
    > workarounds.  But this also means that there is less incentive to fix
    > these issues.

So, you have a problem.   The problem has not one but two solutions
(rsync and push-mirror) which, usefully, make different tradeoffs.
One of the two solutions (rsync) is exactly right for the problem you
have.   And your response to this is to call the exactly right
solution a "tolerable workaround"?

I honestly don't get it.   Why is it taboo that you should be asked to
use rsync?   Geeze, it _works_, don't _fix_ it.


    > >     > arch-pqm seems to require that each GCC developer sets up and
    > >     > maintains his or her own repository.

    > > Eewww, scary.   It's not just pqm that's that way.   Arch itself makes
    > > a similar encouragement.  

    > > Try it.  You'll like it.   The first one's free. 

    > Well, it's not, I currently pay extra for the web server that servers
    > some of my arch changesets. 8-)

Yeah, but none of that money goes to me.  Seems like it should, no?  A
penny a month for every megabyte of arch archives?  :-) Soon, I hope
to be able to afford a beer (plus tip) for that.  Eventually, a
dinner!




    > >     > I'm not sure how many developers actually have to set up branches and
    > >     > merge between them.  

    > > Hehe.  Do you mean in arch world or GCC?

    > In GCC, right now.

    > > In arch world, that's pretty much _all_ that we do.

    > I know.  I'm not sure if this makes that much sense for all
    > developers.  The model which is enforced (well, strongly encouraged)
    > by arch certainly makes easy branching and merging mandatory.  But
    > this does not mean that its the most important aspect of a version
    > control system that should dominate the choice of a CVS successor
    > because other systems might get away with less branching.

There's a fundamental, revctl independent limit on what you can do
without branching:

	C := the number of committers
        M := the average number of minutes between commits from an
             individual committers

      then:

	P := M / C
	     the number of minutes for which an average checked out
             tree is up-to-date

When P is too low, all committers are forced to work blindly --
committing changes whose effect they can not accurately anticipate.
Isn't that bad news?

Last I checked, P for GCC was measured in fractions of an hour.

To be sure, and to explain why GCC hasn't collapsed in a cloud of
chaos, in a large and largely modular program like GCC, checking in
changes from a tree that is not fully up-to-date is something that can
often be gotten away with ("you change your part without breaking
other stuff and I'll do the same").  But that only buys you so much.
Don't try to triple the number of committers, for example.

Aside from fundamental limits, branches are so useful just as a user
interface to change that they should be more common than they are
among CVS users, the problem being that CVS makes them awkward.

    > >From another message of yours:

    > > A few learn svn, a few learn arch, darcs, monotone, .... whichever
    > > there's interest in.  Experiment a bit.  Study how other projects
    > > get along with these systems.  Try to imagine how it fits into
    > > GCC-world.  If you can imagine it making GCC-world much better, then
    > > start evaluating what's involved with switching and advocating for a
    > > switch.

    > There are some things that could encourage such experiments with
    > actual data from the GCC project.  I will ask around a bit.

Nifty.

-t



More information about the Gcc mailing list