A computer readable collection of text or speech
For example the Brown corpus is a million-word collection of samples from 500 written English texts from different genres (newspaper, fiction, non-fiction, academic, etc.).
You'll need to make sure you consider all those properties of a text when you are using it for processing purposes. And there are even more properties of corpora that it's important to consider. Who collected this corpus?
Whenever you build a corpus, you should be documenting these decisions in a datasheet for the corpus.