Description
The explain_document_dl pipeline is a pretrained English NLP pipeline that handles basic text processing and named entity recognition. It can processes text and performs tasks like sentence detection, tokenization, spelling correction, lemmatization, stemming, part-of-speech tagging, word embeddings, and NER.
Included Models:
- DocumentAssembler
- SentenceDetector
- TokenizerModel
- NorvigSweetingModel
- LemmatizerModel
- Stemmer
- PerceptronModel
- WordEmbeddingsModel
- NerDLModel
- NerConverter
How to use
from sparknlp.pretrained import PretrainedPipeline
pipeline = PretrainedPipeline('explain_document_dl', lang = 'en')
annotations = pipeline.fullAnnotate("The Mona Lisa is an oil painting from the 16th century.")[0]
import com.johnsnowlabs.nlp.pretrained.PretrainedPipeline
val pipeline = new PretrainedPipeline("explain_document_dl", lang = "en")
val result = pipeline.fullAnnotate("The Mona Lisa is an oil painting from the 16th century.")(0)
Results
+-------+-----+---+----+--------+------+
| token |begin|end| pos| lemma | ner |
+-------+-----+---+----+--------+------+
| The | 0 | 2| DT | The | O |
| Mona | 4 | 7| NNP| Mona | B-PER|
| Lisa | 9 | 12| NNP| Lisa | I-PER|
| is | 14 | 15| VBZ| be | O |
| an | 17 | 18| DT | an | O |
| oil | 20 | 22| NN | oil | O |
| painting|24 | 31| NN | painting| O |
| from | 33 | 36| IN | from | O |
| the | 38 | 40| DT | the | O |
| 16th | 42 | 45| JJ | 16th | O |
| century| 47 | 54| NN | century| O |
+-------+-----+---+----+--------+------+
Model Information
| Model Name: | explain_document_dl |
| Type: | pipeline |
| Compatibility: | Spark NLP 6.3.0+ |
| License: | Open Source |
| Edition: | Official |
| Language: | en |
| Size: | 176.1 MB |
Included Models
- DocumentAssembler
- SentenceDetector
- RegexTokenizer
- NorvigSweetingModel
- LemmatizerModel
- Stemmer
- PerceptronModel
- WordEmbeddingsModel
- NerDLModel
- NerConverter