Maybe we should just accumulate the corpus instead of replacing it? That should be easier on git right?