Description
A BERT-base-uncased token classification model that scores how much each token in a piece of text actually matters to its meaning, rather than treating every token equally. Given a question and an answer, it assigns an importance weight to each token in the answer - so a content word like “Paris” gets weighted far more heavily than a filler word like “the” or “is”.
This is MARS (Bakman et al., via TruthTorchLM): the intended use is token-importance-weighted confidence scoring, where a model’s uncertainty about the parts of an answer that carry meaning should count more than its uncertainty about grammatical filler. In Spark NLP, load it with MarsTokenImportance, which attaches per-token importance scores as metadata for downstream use.
How to use
mars = MarsTokenImportance.pretrained("mars_token_importance", "en").setInputCols(["question", "completions"]).setOutputCol("completions_with_mars")
val mars = MarsTokenImportance.pretrained("mars_token_importance", "en").setInputCols(Array("question", "completions")).setOutputCol("completions_with_mars")
Model Information
| Model Name: | mars_token_importance |
| Compatibility: | Spark NLP 7.0.0+ |
| License: | Open Source |
| Edition: | Official |
| Input Labels: | [question, completions] |
| Output Labels: | [completions_with_mars] |
| Language: | en |
| Size: | 407.2 MB |