Keep the raw data, always
Whatever processing you do, keep the original. Not because you'll need it. Because you will eventually discover a bug in the processing, and without the raw version every conclusion you drew is unverifiable.
I learned this after finding an error in how I was aggregating. Everything downstream of it was suspect, and the only reason I could rebuild rather than start over was that the untouched source was still there.
Storage is cheap. Redoing six months of research is not.