Changes: - Add optional `tok` tokenizer (a faster but different option than nltk) (#10) - Fix naive vocab checking for tokens (#14) - Allow lowercasing (#13)