Browse by type
[
][#github-license]
[
][#pypi-package]
[
][#pypi-package]
[
][#docs-package]
This framework provides an easy method to compute embeddings for accessing, using, and training state-of-the-art embedding and reranker models. It can be used to compute embeddings using Sentence Transformer models (quickstart), to calculate similarity scores using Cross-Encoder (a.k.a. reranker) models (quickstart), to generate sparse embeddings using Sparse Encoder models (quickstart) or to compute token-level embeddings for ColBERT-style late-interaction retrieval using Multi-Vector Encoder models (quickstart). This unlocks a wide range of applications, including semantic search, semantic textual similarity, and paraphrase mining.
A wide selection of over 15,000 pre-trained Sentence Transformers models are available for immediate use on 🤗 Hugging Face, including many of the state-of-the-art models from the Massive Text Embeddings Benchmark (MTEB) leaderboard. Additionally, it is easy to train or finetune your own embedding models, reranker models, sparse encoder models or multi-vector encoder models using Sentence Transformers, enabling you to create custom models for your specific use cases.
For the full documentation, see www.SBERT.net.
We recommend Python 3.10+, PyTorch 2.2+, and transformers v5.0+.
pip install -U sentence-transformers
See Installation in the docs for uv, conda, source, and editable installs, CUDA setup, and extras ([image], [audio], [video], [train], [onnx], [openvino], [dev]).
See Quickstart in our documentation.
First download a pretrained embedding a.k.a. Sentence Transformer model.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
Then provide some texts to the model.
sentences = [
"The weather is lovely today.",
"It's so sunny outside!",
"He drove to the stadium.",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# => (3, 384)
And that's already it. We now have numpy arrays with the embeddings, one for each text. We can use these to compute similarities.
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[1.0000, 0.6660, 0.1046],
# [0.6660, 1.0000, 0.1411],
# [0.1046, 0.1411, 1.0000]])
First download a pretrained reranker a.k.a. Cross Encoder model.
from sentence_transformers import CrossEncoder
# 1. Load a pretrained CrossEncoder model
model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2")
Then provide some texts to the model.
# The texts for which to predict similarity scores
query = "How many people live in Berlin?"
passages = [
"Berlin had a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers.",
"Berlin has a yearly total of about 135 million day visitors, making it one of the most-visited cities in the European Union.",
"In 2013 around 600,000 Berliners were registered in one of the more than 2,300 sport and fitness clubs.",
]
# 2a. predict scores for pairs of texts
scores = model.predict([(query, passage) for passage in passages])
print(scores)
# => [8.607139 5.506266 6.352977]
And we're good to go. You can also use model.rank to avoid having to perform the reranking manually:
# 2b. Rank a list of passages for a query
ranks = model.rank(query, passages, return_documents=True)
print("Query:", query)
for rank in ranks:
print(f"- #{rank['corpus_id']} ({rank['score']:.2f}): {rank['text']}")
"""
Query: How many people live in Berlin?
- #0 (8.61): Berlin had a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers.
- #2 (6.35): In 2013 around 600,000 Berliners were registered in one of the more than 2,300 sport and fitness clubs.
- #1 (5.51): Berlin has a yearly total of about 135 million day visitors, making it one of the most-visited cities in the European Union.
"""
First download a pretrained sparse embedding a.k.a. Sparse Encoder model.
from sentence_transformers import SparseEncoder
# 1. Load a pretrained SparseEncoder model
model = SparseEncoder("naver/splade-cocondenser-ensembledistil")
# The sentences to encode
sentences = [
"The weather is lovely today.",
"It's so sunny outside!",
"He drove to the stadium.",
]
# 2. Calculate sparse embeddings by calling model.encode()
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 30522] - sparse representation with vocabulary size dimensions
# 3. Calculate the embedding similarities
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 35.629, 9.154, 0.098],
# [ 9.154, 27.478, 0.019],
# [ 0.098, 0.019, 29.553]])
# 4. Check sparsity stats
stats = SparseEncoder.sparsity(embeddings)
print(f"Sparsity: {stats['sparsity_ratio']:.2%}")
# Sparsity: 99.84%
First download a pretrained multi-vector a.k.a. late-interaction (ColBERT-style) model.
from sentence_transformers import MultiVectorEncoder
# 1. Load a pretrained MultiVectorEncoder model
model = MultiVectorEncoder("lightonai/GTE-ModernColBERT-v1")
queries = ["What is the capital of France?"]
documents = [
"Paris is the capital of France.",
"Berlin is the capital of Germany.",
]
# 2. Encode queries and documents into sequences of token-level embeddings
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings[0].shape, document_embeddings[0].shape)
# (10, 128) (9, 128) # one 128-dimensional vector per token
# 3. Score them with late interaction (MaxSim)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[9.6037, 9.4055]])
We provide a large list of pretrained models for more than 100 languages. Some models are general purpose models, while others produce embeddings for specific use cases.
Tip: Using an AI coding agent (Claude Code, Codex, Cursor, Gemini CLI, ...)? Install the
train-sentence-transformersHugging Face Agent Skill viahf skills add train-sentence-transformers [--claude] [--global]and ask your agent to fine-tune a model on your data.
This framework allows you to fine-tune your own sentence embedding methods, so that you get task-specific sentence embeddings. You have various options to choose from in order to get perfect sentence embeddings for your specific task.
Some highlights across the different types of training are:
The following Hugging Face blog posts complement this documentation with narrative walkthroughs and full training examples:
Training guides:
Multimodal:
Efficiency techniques:
You can use this framework for:
Computing Sentence Embeddings
Semantic Textual Similarity
Semantic Search
Retrieve & Re-Rank
[Dense only Retrieval](https://www.sbert.net/examples/sentence_transformer/app
$ claude mcp add sentence-transformers \
-- python -m otcore.mcp_server <graph>