Packages

package similarity

Ordering
  1. Alphabetic
Visibility
  1. Public
  2. All

Type Members

  1. class BM25Approach extends AnnotatorApproach[BM25Model]

    Trains a BM25 (Okapi BM25) lexical ranker over a corpus of tokenized documents.

    Trains a BM25 (Okapi BM25) lexical ranker over a corpus of tokenized documents.

    BM25 is a bag-of-words retrieval function that ranks documents against a query based on the query terms appearing in each document. Because the score of a document depends on corpus-level statistics (how many documents contain a term, and the average document length), BM25 has to be implemented as a two-phase Estimator/Model pair:

    • BM25Approach (this class) scans the full corpus once during fit() and learns:
      • the total document count N
      • the document frequency df(t) of every vocabulary term
      • the average document length avgdl
      • the inverse document frequency idf(t) of every term
    • BM25Model reuses those statistics at query time to score every document against a user-provided query, emitting a BM25_RANKINGS annotation with the relevance score.

    The IDF uses the non-negative (Lucene / Elasticsearch) variant:

    idf(t) = ln(1 + (N - df(t) + 0.5) / (df(t) + 0.5))

    and each document is scored as:

    score(D, Q) = sum over t in Q of
      idf(t) * (tf(t, D) * (k1 + 1)) / (tf(t, D) + k1 * (1 - b + b * |D| / avgdl))

    The input is a column of TOKEN annotations, so BM25 is normally placed after a Tokenizer (optionally followed by a Normalizer and/or StopWordsCleaner). For the produced model and usage examples see BM25Model.

    The learned vocabulary (document frequencies and IDF) is collected to the driver during fit(). For corpora with a very large number of distinct terms, raise minDocFreq to prune rare terms and keep the driver-side vocabulary bounded.

    Example

    import com.johnsnowlabs.nlp.base.DocumentAssembler
    import com.johnsnowlabs.nlp.annotators.{StopWordsCleaner, Tokenizer}
    import com.johnsnowlabs.nlp.annotators.similarity.BM25Approach
    import org.apache.spark.ml.Pipeline
    
    val documentAssembler = new DocumentAssembler()
      .setInputCol("text")
      .setOutputCol("document")
    
    val tokenizer = new Tokenizer()
      .setInputCols("document")
      .setOutputCol("token")
    
    val stopWords = new StopWordsCleaner()
      .setInputCols("token")
      .setOutputCol("clean_token")
      .setCaseSensitive(false)
    
    val bm25 = new BM25Approach()
      .setInputCols("clean_token")
      .setOutputCol("bm25_rankings")
      .setK1(1.2)
      .setB(0.75)
      .setMinDocFreq(1)
      .setCaseSensitive(false)
    
    val pipeline = new Pipeline().setStages(
      Array(documentAssembler, tokenizer, stopWords, bm25))
    
    val model = pipeline.fit(corpus)
    model.stages.last.asInstanceOf[BM25Model].setQuery("vitamin C health benefits fruits")
    model.transform(corpus).selectExpr("explode(bm25_rankings) as ranking").show(false)
  2. class BM25Model extends AnnotatorModel[BM25Model] with HasSimpleAnnotate[BM25Model] with ParamsAndFeaturesWritable

    Fitted model produced by BM25Approach.

    Fitted model produced by BM25Approach. It holds the corpus-level statistics (IDF map, average document length and document count) and scores every document in a dataset against a user-provided query using the Okapi BM25 ranking function.

    The query is provided at transform time, so the same fitted model can be reused for many different queries ("fit once, query many times"). Provide it either as a raw string with setQuery(...) (the model splits it on non-word characters) or — recommended when the corpus was analyzed by a non-trivial pipeline — as already-analyzed tokens with setQueryTokens(...), so the query and the documents are tokenized/normalized identically (see the analyzer-symmetry note on the query parameter). For every input document the model emits a single BM25_RANKINGS annotation whose result is the BM25 score and whose metadata contains:

    • bm25_score — the BM25 relevance score of the document for the current query
    • num_query_terms_matched — how many distinct query terms occur in the document
    • query — the query the document was scored against
    • doc_len — the number of tokens in the document

    Example

    import com.johnsnowlabs.nlp.base.DocumentAssembler
    import com.johnsnowlabs.nlp.annotators.{StopWordsCleaner, Tokenizer}
    import com.johnsnowlabs.nlp.annotators.similarity.{BM25Approach, BM25Model}
    import org.apache.spark.ml.Pipeline
    import org.apache.spark.sql.functions.{col, explode}
    
    val documentAssembler = new DocumentAssembler().setInputCol("text").setOutputCol("document")
    val tokenizer = new Tokenizer().setInputCols("document").setOutputCol("token")
    val stopWords = new StopWordsCleaner().setInputCols("token").setOutputCol("clean_token")
    val bm25 = new BM25Approach().setInputCols("clean_token").setOutputCol("bm25_rankings")
    
    val model = new Pipeline()
      .setStages(Array(documentAssembler, tokenizer, stopWords, bm25))
      .fit(corpus)
    
    val bm25Model = model.stages.last.asInstanceOf[BM25Model]
    bm25Model.setQuery("vitamin C health benefits fruits")
    
    model.transform(corpus)
      .select(explode(col("bm25_rankings")).alias("ranking"))
      .select(col("ranking.metadata")("bm25_score").alias("bm25_score"))
      .show(false)
  3. class DocumentSimilarityRankerApproach extends AnnotatorApproach[DocumentSimilarityRankerModel] with HasEnableCachingProperties

    Annotator that uses LSH techniques present in Spark ML lib to execute approximate nearest neighbors search on top of sentence embeddings.

    Annotator that uses LSH techniques present in Spark ML lib to execute approximate nearest neighbors search on top of sentence embeddings.

    It aims to capture the semantic meaning of a document in a dense, continuous vector space and return it to the ranker search.

    For instantiated/pretrained models, see DocumentSimilarityRankerModel.

    For extended examples of usage, see the jupyter notebook Document Similarity Ranker for Spark NLP.

    Example

    import com.johnsnowlabs.nlp.base._
    import com.johnsnowlabs.nlp.annotator._
    import com.johnsnowlabs.nlp.annotators.similarity.DocumentSimilarityRankerApproach
    import com.johnsnowlabs.nlp.finisher.DocumentSimilarityRankerFinisher
    import org.apache.spark.ml.Pipeline
    
    import spark.implicits._
    
    val documentAssembler = new DocumentAssembler()
      .setInputCol("text")
      .setOutputCol("document")
    
    val sentenceEmbeddings = RoBertaSentenceEmbeddings
      .pretrained()
      .setInputCols("document")
      .setOutputCol("sentence_embeddings")
    
    val documentSimilarityRanker = new DocumentSimilarityRankerApproach()
      .setInputCols("sentence_embeddings")
      .setOutputCol("doc_similarity_rankings")
      .setSimilarityMethod("brp")
      .setNumberOfNeighbours(1)
      .setBucketLength(2.0)
      .setNumHashTables(3)
      .setVisibleDistances(true)
      .setIdentityRanking(false)
    
    val documentSimilarityRankerFinisher = new DocumentSimilarityRankerFinisher()
      .setInputCols("doc_similarity_rankings")
      .setOutputCols(
        "finished_doc_similarity_rankings_id",
        "finished_doc_similarity_rankings_neighbors")
      .setExtractNearestNeighbor(true)
    
    // Let's use a dataset where we can visually control similarity
    // Documents are coupled, as 1-2, 3-4, 5-6, 7-8 and they were create to be similar on purpose
    val data = Seq(
      "First document, this is my first sentence. This is my second sentence.",
      "Second document, this is my second sentence. This is my second sentence.",
      "Third document, climate change is arguably one of the most pressing problems of our time.",
      "Fourth document, climate change is definitely one of the most pressing problems of our time.",
      "Fifth document, Florence in Italy, is among the most beautiful cities in Europe.",
      "Sixth document, Florence in Italy, is a very beautiful city in Europe like Lyon in France.",
      "Seventh document, the French Riviera is the Mediterranean coastline of the southeast corner of France.",
      "Eighth document, the warmest place in France is the French Riviera coast in Southern France.")
      .toDF("text")
    
    val pipeline = new Pipeline().setStages(
      Array(
        documentAssembler,
        sentenceEmbeddings,
        documentSimilarityRanker,
        documentSimilarityRankerFinisher))
    
    val result = pipeline.fit(data).transform(data)
    
    result
      .select("finished_doc_similarity_rankings_id", "finished_doc_similarity_rankings_neighbors")
      .show(10, truncate = false)
    +-----------------------------------+------------------------------------------+
    |finished_doc_similarity_rankings_id|finished_doc_similarity_rankings_neighbors|
    +-----------------------------------+------------------------------------------+
    |1510101612                         |[(1634839239,0.12448559591306324)]        |
    |1634839239                         |[(1510101612,0.12448559591306324)]        |
    |-612640902                         |[(1274183715,0.1220122862046063)]         |
    |1274183715                         |[(-612640902,0.1220122862046063)]         |
    |-1320876223                        |[(1293373212,0.17848855164122393)]        |
    |1293373212                         |[(-1320876223,0.17848855164122393)]       |
    |-1548374770                        |[(-1719102856,0.23297156732534166)]       |
    |-1719102856                        |[(-1548374770,0.23297156732534166)]       |
    +-----------------------------------+------------------------------------------+
  4. class DocumentSimilarityRankerModel extends AnnotatorModel[DocumentSimilarityRankerModel] with HasSimpleAnnotate[DocumentSimilarityRankerModel] with HasEmbeddingsProperties with ParamsAndFeaturesWritable

    Instantiated model of the DocumentSimilarityRankerApproach.

    Instantiated model of the DocumentSimilarityRankerApproach. For usage and examples see the documentation of the main class.

  5. case class IndexedNeighbors(neighbors: Array[Int]) extends NeighborAnnotation with Product with Serializable
  6. case class IndexedNeighborsWithDistance(neighbors: Array[(Int, Double)]) extends NeighborAnnotation with Product with Serializable
  7. sealed trait NeighborAnnotation extends AnyRef
  8. case class NeighborsResultSet(result: (Int, NeighborAnnotation)) extends Product with Serializable
  9. class PairwiseVectorSimilarity extends AnnotatorModel[PairwiseVectorSimilarity] with HasSimpleAnnotate[PairwiseVectorSimilarity] with ParamsAndFeaturesWritable

    Computes pairwise vector similarity between two sets of sentence embeddings.

    Computes pairwise vector similarity between two sets of sentence embeddings.

    The annotator takes **two** SENTENCE_EMBEDDINGS input columns (e.g. query embeddings and document embeddings already joined on the same row) and, for every row, scores all N×M pairs between the embeddings in column A and the embeddings in column B. Each pair produces one VECTOR_SIMILARITY output annotation whose result holds the score as a String, and whose metadata holds typed fields for easy extraction.

    Sign conventions

    method        range         higher means
    ----------    ----------    ------------
    cosine        [-1.0, 1.0]   more similar
    dotProduct    (−∞, +∞)      more similar
    euclidean     (−∞, 0.0]     more similar (0.0 = identical vectors)

    The euclidean method returns the **negative** L2 distance so that "higher is better" holds uniformly across all three methods. A score of 0.0 means the two vectors are identical; more negative values indicate less similarity.

    Important: input data shape

    Each input column should contain **exactly one** embedding per row for standard document retrieval. If a column contains N > 1 embeddings (e.g. produced by SentenceDetector + embedder), all N×M cross-pairs are scored and returned as separate annotations. To compare queries against a corpus, join them first:

    Example

    import com.johnsnowlabs.nlp.annotators.similarity.PairwiseVectorSimilarity
    import org.apache.spark.sql.functions.{col, desc, explode}
    
    // Assume embeddingPipeline produces a "embeddings" column of SENTENCE_EMBEDDINGS.
    val queryDf  = embeddingPipeline.transform(queries).select(col("embeddings").as("query_emb"))
    val corpusDf = embeddingPipeline.transform(corpus).select(col("embeddings").as("doc_emb"), col("id"))
    
    // CrossJoin to get one (query, document) pair per row, then score.
    val paired = queryDf.crossJoin(corpusDf)
    
    val pvs = new PairwiseVectorSimilarity()
      .setInputCols("query_emb", "doc_emb")
      .setOutputCol("similarity")
      .setSimilarityMethod("cosine")
    
    pvs.transform(paired)
      .select(explode(col("similarity")).as("s"))
      .select(
        col("s.metadata")("sentence_a_text").as("query"),
        col("s.metadata")("sentence_b_text").as("document"),
        col("s.result").cast("double").as("score"))
      .orderBy(desc("score"))
      .show(false)
  10. trait ReadableBM25Model extends ParamsAndFeaturesReadable[BM25Model]
  11. trait ReadableDocumentSimilarityRanker extends ParamsAndFeaturesReadable[DocumentSimilarityRankerModel]
  12. trait ReadablePairwiseVectorSimilarity extends ParamsAndFeaturesReadable[PairwiseVectorSimilarity]

Value Members

  1. object BM25Approach extends DefaultParamsReadable[BM25Approach] with Serializable

    This is the companion object of BM25Approach.

    This is the companion object of BM25Approach. Please refer to that class for the documentation.

  2. object BM25Model extends ReadableBM25Model with Serializable

    This is the companion object of BM25Model.

    This is the companion object of BM25Model. Please refer to that class for the documentation.

  3. object DocumentSimilarityRankerApproach extends DefaultParamsReadable[DocumentSimilarityRankerApproach] with Serializable

    This is the companion object of DocumentSimilarityRankerApproach.

    This is the companion object of DocumentSimilarityRankerApproach. Please refer to that class for the documentation.

  4. object DocumentSimilarityRankerModel extends ReadableDocumentSimilarityRanker with Serializable
  5. object DocumentSimilarityUtil
  6. object PairwiseVectorSimilarity extends ReadablePairwiseVectorSimilarity with Serializable

Ungrouped