MARS Token Importance

Description

A BERT-base-uncased token classification model that scores how much each token in a piece of text actually matters to its meaning, rather than treating every token equally. Given a question and an answer, it assigns an importance weight to each token in the answer - so a content word like “Paris” gets weighted far more heavily than a filler word like “the” or “is”.

This is MARS (Bakman et al., via TruthTorchLM): the intended use is token-importance-weighted confidence scoring, where a model’s uncertainty about the parts of an answer that carry meaning should count more than its uncertainty about grammatical filler. In Spark NLP, load it with MarsTokenImportance, which attaches per-token importance scores as metadata for downstream use.

Download Copy S3 URI

How to use

mars = MarsTokenImportance.pretrained("mars_token_importance", "en").setInputCols(["question", "completions"]).setOutputCol("completions_with_mars")
val mars = MarsTokenImportance.pretrained("mars_token_importance", "en").setInputCols(Array("question", "completions")).setOutputCol("completions_with_mars")

Model Information

Model Name: mars_token_importance
Compatibility: Spark NLP 7.0.0+
License: Open Source
Edition: Official
Input Labels: [question, completions]
Output Labels: [completions_with_mars]
Language: en
Size: 407.2 MB