Description
This model generates aligned text and image embeddings for multimodal retrieval workflows. It is used with BiEncoderMultimodalEmbeddings, which consumes paired DOCUMENT and IMAGE annotations and produces two embedding columns: one for the text side and one for the image side. The embeddings can be indexed in a vector database and used for text-to-image, image-to-text, text-to-text, or image-to-image retrieval and RAG pipelines.
Predicted Entities
Open in Colab Download Copy S3 URI
How to use
Use BiEncoderMultimodalEmbeddings.pretrained(“ops_mm_embedding_v1_2b”, “en”) with a DOCUMENT input column and an IMAGE input column. The model outputs SENTENCE_EMBEDDINGS in derived columns named <outputCol>_doc_embeddings and <outputCol>_image_embeddings.
from pyspark.ml import Pipeline
from sparknlp.reader import ReaderAssembler, LayoutAlignerForVision
from sparknlp.annotator import BiEncoderMultimodalEmbeddings
reader = (
ReaderAssembler()
.setContentPath("file:///path/to/document.html")
.setContentType("text/html")
.setOutputCol("reader")
.setOutputAsDocument(False)
)
vision_aligner = (
LayoutAlignerForVision()
.setInputCols(["reader_text", "reader_image"])
.setOutputCol("vision_pair")
.setExplodeDocs(True)
.setAddNeighborText(True)
)
opsmm = (
BiEncoderMultimodalEmbeddings.pretrained("ops_mm_embedding_v1_2b", "en")
.setInputCols(["vision_pair_doc", "vision_pair_image"])
.setOutputCol("opsmm")
.setBatchSize(1)
)
pipeline = Pipeline(stages=[reader, vision_aligner, opsmm])
result = pipeline.fit(spark.emptyDataFrame).transform(spark.emptyDataFrame)
result.select("opsmm_doc_embeddings", "opsmm_image_embeddings").show(truncate=False)
from pyspark.ml import Pipeline
from sparknlp.reader import ReaderAssembler, LayoutAlignerForVision
from sparknlp.annotator import BiEncoderMultimodalEmbeddings
reader = (
ReaderAssembler()
.setContentPath("file:///path/to/document.html")
.setContentType("text/html")
.setOutputCol("reader")
.setOutputAsDocument(False)
)
vision_aligner = (
LayoutAlignerForVision()
.setInputCols(["reader_text", "reader_image"])
.setOutputCol("vision_pair")
.setExplodeDocs(True)
.setAddNeighborText(True)
)
opsmm = (
BiEncoderMultimodalEmbeddings.pretrained("ops_mm_embedding_v1_2b", "en")
.setInputCols(["vision_pair_doc", "vision_pair_image"])
.setOutputCol("opsmm")
.setBatchSize(1)
)
pipeline = Pipeline(stages=[reader, vision_aligner, opsmm])
result = pipeline.fit(spark.emptyDataFrame).transform(spark.emptyDataFrame)
result.select("opsmm_doc_embeddings", "opsmm_image_embeddings").show(truncate=False)
Results
The model produces 1536-dimensional embeddings for both text and image inputs. It does not produce labels, entities, or generated text
Model Information
| Model Name: | ops_mm_embedding_v1_2b |
| Compatibility: | Spark NLP 6.4.1+ |
| License: | Open Source |
| Edition: | Official |
| Input Labels: | [vision_pair_doc, vision_pair_image] |
| Output Labels: | [mm] |
| Language: | en |
| Size: | 3.0 GB |
PREVIOUSModernBERT Base ONNX
NEXTSegment Any Text