Dismiss Notice
Join Physics Forums Today!
The friendliest, high quality science and math community on the planet! Everyone who loves science is here!

Text file with all English words and their part of speech

  1. Feb 16, 2015 #1
    Hey all, been wanting to get into NLP (natural language processing) but I require a text file with a list of all English words (not the definitions) and a tag indicating their part of speech, I know it exists because I had it on my old laptop but I can't seem to refind it. Any help apreciated.
  2. jcsd
  3. Feb 16, 2015 #2


    User Avatar
    Gold Member

    ALL the words in English? That's going to be one hell of a file. And mostly useless. Of the 1,000,000+ words in English (depending on who you believe), an average speaker has a vocab of about 6,000 to 8,000 words and a highly educated one has under 20,000 so even highly educated English speakers use less than 2% of the words in the language (and may have "receptive" knowledge of another 1% or less). I suspect that your list problably had 20,000 to 30,000 words, not "all" the words in English.
  4. Feb 16, 2015 #3


    User Avatar
    Gold Member

    I won't be able to help you find your file, but if you want a dictionary with words in it https://github.com/TheBerkin/Rantionary/blob/master/Prepositions.dic [Broken]is one. It has pronunciation as well.
    Last edited by a moderator: May 7, 2017
  5. Feb 16, 2015 #4
    These guys are often used as corpora for natural language, and their database is downloadable (free). Python NLTK uses this, as do a lot of other NLP libraries.
    Last edited: Feb 16, 2015
  6. Feb 17, 2015 #5
    You might want to search for the 'Brown Corpus', one of the earliest best known corpus with parts of speech. I don't think any two groups of computational linguists agree on the parts of speech; you may not even need parts of speech data depending on what you're doing.
  7. Feb 18, 2015 #6

    That's the complete list of sources used by the Python natural language toolkit. Wordnet and Brown Corpus are in there, as are others. That's quite a good library.
Share this great discussion with others via Reddit, Google+, Twitter, or Facebook