package similarity
- Alphabetic
- Public
- All
Type Members
-
class
BM25Approach extends AnnotatorApproach[BM25Model]
Trains a BM25 (Okapi BM25) lexical ranker over a corpus of tokenized documents.
Trains a BM25 (Okapi BM25) lexical ranker over a corpus of tokenized documents.
BM25 is a bag-of-words retrieval function that ranks documents against a query based on the query terms appearing in each document. Because the score of a document depends on corpus-level statistics (how many documents contain a term, and the average document length), BM25 has to be implemented as a two-phase Estimator/Model pair:
- BM25Approach (this class) scans the full corpus once during
fit()and learns:- the total document count
N - the document frequency
df(t)of every vocabulary term - the average document length
avgdl - the inverse document frequency
idf(t)of every term
- the total document count
- BM25Model reuses those statistics at query time to score every document against a
user-provided query, emitting a
BM25_RANKINGSannotation with the relevance score.
The IDF uses the non-negative (Lucene / Elasticsearch) variant:
idf(t) = ln(1 + (N - df(t) + 0.5) / (df(t) + 0.5))
and each document is scored as:
score(D, Q) = sum over t in Q of idf(t) * (tf(t, D) * (k1 + 1)) / (tf(t, D) + k1 * (1 - b + b * |D| / avgdl))
The input is a column of
TOKENannotations, so BM25 is normally placed after a Tokenizer (optionally followed by a Normalizer and/or StopWordsCleaner). For the produced model and usage examples see BM25Model.The learned vocabulary (document frequencies and IDF) is collected to the driver during
fit(). For corpora with a very large number of distinct terms, raise minDocFreq to prune rare terms and keep the driver-side vocabulary bounded.Example
import com.johnsnowlabs.nlp.base.DocumentAssembler import com.johnsnowlabs.nlp.annotators.{StopWordsCleaner, Tokenizer} import com.johnsnowlabs.nlp.annotators.similarity.BM25Approach import org.apache.spark.ml.Pipeline val documentAssembler = new DocumentAssembler() .setInputCol("text") .setOutputCol("document") val tokenizer = new Tokenizer() .setInputCols("document") .setOutputCol("token") val stopWords = new StopWordsCleaner() .setInputCols("token") .setOutputCol("clean_token") .setCaseSensitive(false) val bm25 = new BM25Approach() .setInputCols("clean_token") .setOutputCol("bm25_rankings") .setK1(1.2) .setB(0.75) .setMinDocFreq(1) .setCaseSensitive(false) val pipeline = new Pipeline().setStages( Array(documentAssembler, tokenizer, stopWords, bm25)) val model = pipeline.fit(corpus) model.stages.last.asInstanceOf[BM25Model].setQuery("vitamin C health benefits fruits") model.transform(corpus).selectExpr("explode(bm25_rankings) as ranking").show(false)
- BM25Approach (this class) scans the full corpus once during
-
class
BM25Model extends AnnotatorModel[BM25Model] with HasSimpleAnnotate[BM25Model] with ParamsAndFeaturesWritable
Fitted model produced by BM25Approach.
Fitted model produced by BM25Approach. It holds the corpus-level statistics (IDF map, average document length and document count) and scores every document in a dataset against a user-provided query using the Okapi BM25 ranking function.
The query is provided at transform time, so the same fitted model can be reused for many different queries ("fit once, query many times"). Provide it either as a raw string with
setQuery(...)(the model splits it on non-word characters) or — recommended when the corpus was analyzed by a non-trivial pipeline — as already-analyzed tokens withsetQueryTokens(...), so the query and the documents are tokenized/normalized identically (see the analyzer-symmetry note on thequeryparameter). For every input document the model emits a singleBM25_RANKINGSannotation whoseresultis the BM25 score and whose metadata contains:bm25_score— the BM25 relevance score of the document for the current querynum_query_terms_matched— how many distinct query terms occur in the documentquery— the query the document was scored againstdoc_len— the number of tokens in the document
Example
import com.johnsnowlabs.nlp.base.DocumentAssembler import com.johnsnowlabs.nlp.annotators.{StopWordsCleaner, Tokenizer} import com.johnsnowlabs.nlp.annotators.similarity.{BM25Approach, BM25Model} import org.apache.spark.ml.Pipeline import org.apache.spark.sql.functions.{col, explode} val documentAssembler = new DocumentAssembler().setInputCol("text").setOutputCol("document") val tokenizer = new Tokenizer().setInputCols("document").setOutputCol("token") val stopWords = new StopWordsCleaner().setInputCols("token").setOutputCol("clean_token") val bm25 = new BM25Approach().setInputCols("clean_token").setOutputCol("bm25_rankings") val model = new Pipeline() .setStages(Array(documentAssembler, tokenizer, stopWords, bm25)) .fit(corpus) val bm25Model = model.stages.last.asInstanceOf[BM25Model] bm25Model.setQuery("vitamin C health benefits fruits") model.transform(corpus) .select(explode(col("bm25_rankings")).alias("ranking")) .select(col("ranking.metadata")("bm25_score").alias("bm25_score")) .show(false)
-
class
DocumentSimilarityRankerApproach extends AnnotatorApproach[DocumentSimilarityRankerModel] with HasEnableCachingProperties
Annotator that uses LSH techniques present in Spark ML lib to execute approximate nearest neighbors search on top of sentence embeddings.
Annotator that uses LSH techniques present in Spark ML lib to execute approximate nearest neighbors search on top of sentence embeddings.
It aims to capture the semantic meaning of a document in a dense, continuous vector space and return it to the ranker search.
For instantiated/pretrained models, see DocumentSimilarityRankerModel.
For extended examples of usage, see the jupyter notebook Document Similarity Ranker for Spark NLP.
Example
import com.johnsnowlabs.nlp.base._ import com.johnsnowlabs.nlp.annotator._ import com.johnsnowlabs.nlp.annotators.similarity.DocumentSimilarityRankerApproach import com.johnsnowlabs.nlp.finisher.DocumentSimilarityRankerFinisher import org.apache.spark.ml.Pipeline import spark.implicits._ val documentAssembler = new DocumentAssembler() .setInputCol("text") .setOutputCol("document") val sentenceEmbeddings = RoBertaSentenceEmbeddings .pretrained() .setInputCols("document") .setOutputCol("sentence_embeddings") val documentSimilarityRanker = new DocumentSimilarityRankerApproach() .setInputCols("sentence_embeddings") .setOutputCol("doc_similarity_rankings") .setSimilarityMethod("brp") .setNumberOfNeighbours(1) .setBucketLength(2.0) .setNumHashTables(3) .setVisibleDistances(true) .setIdentityRanking(false) val documentSimilarityRankerFinisher = new DocumentSimilarityRankerFinisher() .setInputCols("doc_similarity_rankings") .setOutputCols( "finished_doc_similarity_rankings_id", "finished_doc_similarity_rankings_neighbors") .setExtractNearestNeighbor(true) // Let's use a dataset where we can visually control similarity // Documents are coupled, as 1-2, 3-4, 5-6, 7-8 and they were create to be similar on purpose val data = Seq( "First document, this is my first sentence. This is my second sentence.", "Second document, this is my second sentence. This is my second sentence.", "Third document, climate change is arguably one of the most pressing problems of our time.", "Fourth document, climate change is definitely one of the most pressing problems of our time.", "Fifth document, Florence in Italy, is among the most beautiful cities in Europe.", "Sixth document, Florence in Italy, is a very beautiful city in Europe like Lyon in France.", "Seventh document, the French Riviera is the Mediterranean coastline of the southeast corner of France.", "Eighth document, the warmest place in France is the French Riviera coast in Southern France.") .toDF("text") val pipeline = new Pipeline().setStages( Array( documentAssembler, sentenceEmbeddings, documentSimilarityRanker, documentSimilarityRankerFinisher)) val result = pipeline.fit(data).transform(data) result .select("finished_doc_similarity_rankings_id", "finished_doc_similarity_rankings_neighbors") .show(10, truncate = false) +-----------------------------------+------------------------------------------+ |finished_doc_similarity_rankings_id|finished_doc_similarity_rankings_neighbors| +-----------------------------------+------------------------------------------+ |1510101612 |[(1634839239,0.12448559591306324)] | |1634839239 |[(1510101612,0.12448559591306324)] | |-612640902 |[(1274183715,0.1220122862046063)] | |1274183715 |[(-612640902,0.1220122862046063)] | |-1320876223 |[(1293373212,0.17848855164122393)] | |1293373212 |[(-1320876223,0.17848855164122393)] | |-1548374770 |[(-1719102856,0.23297156732534166)] | |-1719102856 |[(-1548374770,0.23297156732534166)] | +-----------------------------------+------------------------------------------+
-
class
DocumentSimilarityRankerModel extends AnnotatorModel[DocumentSimilarityRankerModel] with HasSimpleAnnotate[DocumentSimilarityRankerModel] with HasEmbeddingsProperties with ParamsAndFeaturesWritable
Instantiated model of the DocumentSimilarityRankerApproach.
Instantiated model of the DocumentSimilarityRankerApproach. For usage and examples see the documentation of the main class.
- case class IndexedNeighbors(neighbors: Array[Int]) extends NeighborAnnotation with Product with Serializable
- case class IndexedNeighborsWithDistance(neighbors: Array[(Int, Double)]) extends NeighborAnnotation with Product with Serializable
- sealed trait NeighborAnnotation extends AnyRef
- case class NeighborsResultSet(result: (Int, NeighborAnnotation)) extends Product with Serializable
-
class
PairwiseVectorSimilarity extends AnnotatorModel[PairwiseVectorSimilarity] with HasSimpleAnnotate[PairwiseVectorSimilarity] with ParamsAndFeaturesWritable
Computes pairwise vector similarity between two sets of sentence embeddings.
Computes pairwise vector similarity between two sets of sentence embeddings.
The annotator takes **two**
SENTENCE_EMBEDDINGSinput columns (e.g. query embeddings and document embeddings already joined on the same row) and, for every row, scores all N×M pairs between the embeddings in column A and the embeddings in column B. Each pair produces oneVECTOR_SIMILARITYoutput annotation whoseresultholds the score as a String, and whosemetadataholds typed fields for easy extraction.Sign conventions
method range higher means ---------- ---------- ------------ cosine [-1.0, 1.0] more similar dotProduct (−∞, +∞) more similar euclidean (−∞, 0.0] more similar (0.0 = identical vectors)
The
euclideanmethod returns the **negative** L2 distance so that "higher is better" holds uniformly across all three methods. A score of0.0means the two vectors are identical; more negative values indicate less similarity.Important: input data shape
Each input column should contain **exactly one** embedding per row for standard document retrieval. If a column contains N > 1 embeddings (e.g. produced by
SentenceDetector+ embedder), all N×M cross-pairs are scored and returned as separate annotations. To compare queries against a corpus, join them first:Example
import com.johnsnowlabs.nlp.annotators.similarity.PairwiseVectorSimilarity import org.apache.spark.sql.functions.{col, desc, explode} // Assume embeddingPipeline produces a "embeddings" column of SENTENCE_EMBEDDINGS. val queryDf = embeddingPipeline.transform(queries).select(col("embeddings").as("query_emb")) val corpusDf = embeddingPipeline.transform(corpus).select(col("embeddings").as("doc_emb"), col("id")) // CrossJoin to get one (query, document) pair per row, then score. val paired = queryDf.crossJoin(corpusDf) val pvs = new PairwiseVectorSimilarity() .setInputCols("query_emb", "doc_emb") .setOutputCol("similarity") .setSimilarityMethod("cosine") pvs.transform(paired) .select(explode(col("similarity")).as("s")) .select( col("s.metadata")("sentence_a_text").as("query"), col("s.metadata")("sentence_b_text").as("document"), col("s.result").cast("double").as("score")) .orderBy(desc("score")) .show(false)
- trait ReadableBM25Model extends ParamsAndFeaturesReadable[BM25Model]
- trait ReadableDocumentSimilarityRanker extends ParamsAndFeaturesReadable[DocumentSimilarityRankerModel]
- trait ReadablePairwiseVectorSimilarity extends ParamsAndFeaturesReadable[PairwiseVectorSimilarity]
Value Members
-
object
BM25Approach extends DefaultParamsReadable[BM25Approach] with Serializable
This is the companion object of BM25Approach.
This is the companion object of BM25Approach. Please refer to that class for the documentation.
-
object
BM25Model extends ReadableBM25Model with Serializable
This is the companion object of BM25Model.
This is the companion object of BM25Model. Please refer to that class for the documentation.
-
object
DocumentSimilarityRankerApproach extends DefaultParamsReadable[DocumentSimilarityRankerApproach] with Serializable
This is the companion object of DocumentSimilarityRankerApproach.
This is the companion object of DocumentSimilarityRankerApproach. Please refer to that class for the documentation.
- object DocumentSimilarityRankerModel extends ReadableDocumentSimilarityRanker with Serializable
- object DocumentSimilarityUtil
- object PairwiseVectorSimilarity extends ReadablePairwiseVectorSimilarity with Serializable