MCPcopy Create free account

hub / github.com/huggingface/sentence-transformers / functions

Functions3,759 in github.com/huggingface/sentence-transformers

Functiontest_check_version_requirements_unmet
()
tests/util/test_environment.py:151
Functiontest_check_version_requirements_unmet_python_excluded_but_others_listed
()
tests/util/test_environment.py:180
Functiontest_check_version_requirements_unmet_python_omits_pip_install
`pip install "python>=99.0"` isn't a command anyone can run.
tests/util/test_environment.py:170
Functiontest_chunk_ranges_budget_packing
Greedy packing keeps each chunk's per_item_token * width * count under the budget, isolates a long outlier in its own chunk, and never emits an em
tests/util/test_similarity.py:617
Methodtest_class_pool_matches_module_helper_directly
(self)
tests/multi_vector_encoder/test_token_pooling.py:94
Methodtest_classifier_dropout
(self, caplog)
tests/util/test_decorators.py:232
Functiontest_classifier_dropout_default_value
(reranker_bert_tiny_model: CrossEncoder)
tests/cross_encoder/test_model.py:36
Functiontest_classifier_dropout_is_set
()
tests/cross_encoder/test_model.py:30
Methodtest_classifier_dropout_merges_into_existing_config_kwargs
(self, caplog)
tests/util/test_decorators.py:243
Functiontest_clip
()
tests/sentence_transformer/test_model.py:971
Functiontest_cmnrl_attributes_and_config_back_compat
CachedMultipleNegativesRankingLoss composes an inner MultipleNegativesRankingLoss, but the documented attributes and ``get_config_dict`` keys must
tests/sentence_transformer/losses/test_cmnrl.py:349
Functiontest_cmnrl_gather_across_devices_offset
The ``gather_across_devices`` path has no other test. Simulate world=2/rank=1 by having ``all_gather_with_grad`` return ``[local; local]``: the ga
tests/sentence_transformer/losses/test_cmnrl.py:496
Functiontest_cmnrl_hyperparameters_stay_assignable
The delegating properties must keep supporting assignment, as the plain attributes did before the rebase. That includes assigning an ``nn.Module``
tests/sentence_transformer/losses/test_cmnrl.py:435
Functiontest_cmnrl_matryoshka
``MatryoshkaLoss(CachedMultipleNegativesRankingLoss(...))``, the combination the visual document retrieval example ships with, must match ``Matryo
tests/sentence_transformer/losses/test_cmnrl.py:397
Functiontest_cmnrl_same_grad
( train_samples_mnrl: list[tuple[str, str, str]], train_samples_cmnrl: list[tuple[str, str, str]],
tests/sentence_transformer/losses/test_cmnrl.py:90
Functiontest_cmnrl_token_budget_matches_mnrl
With ``mini_batch_num_tokens``, the embedding passes pack by token count, but the loss and gradient must still exactly match MultipleNegativesRank
tests/sentence_transformer/losses/test_cmnrl.py:468
Functiontest_colbert_kd_scores_are_the_block_diagonal_of_the_matrix
KD scores are each query's own document group, the query-major block diagonal of the full matrix. Split out of the bfloat16 test below so it keeps
tests/multi_vector_encoder/test_model.py:1012
Functiontest_colbert_scorers_match_the_maxsim_keyword_surface
Every colbert scorer takes and forwards chunk_elements and length_normalize, so one kwargs dict works across the family and the default training p
tests/multi_vector_encoder/test_model.py:1737
Functiontest_colbert_scores_keep_float32_through_delegation
The losses' default similarity_fct surface must keep maxsim's float32 scores, also under a future fused reimplementation.
tests/multi_vector_encoder/test_model.py:1030
Functiontest_colbert_scoring_callable_query_major
()
tests/multi_vector_encoder/test_model.py:1539
Functiontest_collator_assigns_task_by_position_regardless_of_name
The collator assigns ``task`` by column POSITION (column 0 = query, the rest = document) to match the losses, which score positionally. Column nam
tests/multi_vector_encoder/test_trainer.py:175
Functiontest_collator_pads_to_batch_longest_not_model_max_length
A fresh MVE has ``document_length=None``. Before the fix the collator forced ``padding="max_length"`` which silently padded every batch to ``token
tests/multi_vector_encoder/test_trainer.py:121
Functiontest_collator_router_mapping_overrides_positional_default
``router_mapping`` overrides the positional default per column.
tests/multi_vector_encoder/test_trainer.py:198
Functiontest_collator_stamps_resolved_task_into_batch
The collator stamps the task each column was tokenized with as ``{column}_task`` so the losses re-run the model under the same task (instead of re
tests/multi_vector_encoder/test_trainer.py:217
Functiontest_collator_within_column_rectangular_across_ragged_columns
Within each column every row shares ``T`` (the loss reshapes per-column on ``dim=0``), but columns are independently padded to their own batch-lon
tests/multi_vector_encoder/test_trainer.py:148
Functiontest_column_names
Test specifying custom column names.
tests/util/test_hard_negatives.py:127
Methodtest_combined_modality_builds_dicts
A model with tuple modality ("image", "text") builds multimodal dicts from text+image columns.
tests/base/test_model_card.py:564
Methodtest_combined_only_model
A Kosmos-like model with only ("image", "text") still builds combined dicts.
tests/base/test_model_card.py:630
Functiontest_community_detection_all_points_in_one_community
Test case where all points form a single community due to a low threshold.
tests/util/test_retrieval.py:89
Functiontest_community_detection_gpu_support
Test case for GPU support (if available).
tests/util/test_retrieval.py:186
Functiontest_community_detection_large_batch_size
Test case with a large dataset and batching.
tests/util/test_retrieval.py:158
Functiontest_community_detection_min_community_size_filtering
Test case where communities are filtered based on minimum size.
tests/util/test_retrieval.py:105
Functiontest_community_detection_no_communities_high_threshold
Test case where no communities are found due to a high threshold.
tests/util/test_retrieval.py:75
Functiontest_community_detection_numpy_input
Test case where input is a numpy array instead of a torch tensor.
tests/util/test_retrieval.py:142
Functiontest_community_detection_overlapping_communities
Test case with overlapping communities (resolved by the function).
tests/util/test_retrieval.py:122
Functiontest_community_detection_similarities_equal_to_threshold
Members whose similarity equals the threshold exactly must not be truncated. The candidate window grew only while the smallest candidate was stri
tests/util/test_retrieval.py:167
Functiontest_community_detection_two_clear_communities
Test case with two clear communities.
tests/util/test_retrieval.py:55
Methodtest_compound_modality_expanded
A dict with multiple modalities (e.g. text+image) should be expanded into separate content items, not wrapped as a single compound-typed item.
tests/base/test_modality.py:1401
Methodtest_compound_modality_preserves_user_order_in_both_roles
(self, modalities)
tests/base/test_modality.py:1335
Functiontest_config_and_model_none_without_transformer
``.config`` and ``.model`` must degrade to ``None`` when the module stack has no underlying HuggingFace Transformer rather than raising AttributeE
tests/cross_encoder/test_model.py:419
Methodtest_config_args_deprecated
(self, caplog)
tests/base/modules/test_transformer.py:381
Functiontest_config_delegates_to_underlying_transformers_model
`model.config` should return the underlying transformers PretrainedConfig so integrations like Deepspeed and transformers.Trainer can read `hidden
tests/base/test_model.py:252
Functiontest_config_is_settable
Assigning ``model.config`` should not raise and should round-trip through the underlying transformers model (fixes optimum ONNX export regression
tests/base/test_model.py:270
Functiontest_config_returns_none_when_no_underlying_transformers_model
If the model has no underlying transformers PreTrainedModel (e.g. StaticEmbedding-only), `model.config` returns `None` instead of raising.
tests/base/test_model.py:261
Functiontest_config_setter_on_model_without_underlying_transformers_model
Assigning ``model.config`` on a StaticEmbedding-only model (no underlying transformers model) should be silently ignored, i.e. ``model.config`` ke
tests/base/test_model.py:282
Methodtest_conflicting_sampling_rates_raise
https://github.com/huggingface/sentence-transformers/issues/3874
tests/base/test_modality.py:714
Methodtest_consistent_hash
The same encoded source produces the same hash.
tests/base/test_model_card.py:1063
Methodtest_contains_expected_modalities
(self)
tests/base/test_modality.py:392
Methodtest_content_not_str_or_list
Non-string, non-list content is treated as non-text.
tests/base/test_modality.py:1553
Functiontest_contrastive_loss_does_not_warn_for_binary_labels
()
tests/sentence_transformer/losses/test_contrastive.py:44
Functiontest_contrastive_loss_only_checks_the_first_batch
()
tests/sentence_transformer/losses/test_contrastive.py:56
Functiontest_contrastive_loss_warns_once_for_non_binary_labels
()
tests/sentence_transformer/losses/test_contrastive.py:27
Functiontest_conversion_from_sentence_transformer_keeps_prompts_and_similarity
The SentenceTransformer branch reuses the saved modules and appends a SparseAutoEncoder (CSR), so the source's prompts and similarity stay meaning
tests/sparse_encoder/test_model.py:936
Functiontest_conversion_ignores_prompts_from_sparse_save
Converting a SparseEncoder (or CrossEncoder) save rebuilds the default MVE modules, so the source's prompts and default_prompt_name are not inheri
tests/multi_vector_encoder/test_model.py:901
Functiontest_conversion_ignores_source_prompts_and_default_prompt_name
Converting a CrossEncoder save rebuilds this family's default modules, so the source's prompts and default_prompt_name no longer describe the load
tests/sentence_transformer/test_model.py:1411
Functiontest_conversion_ignores_source_prompts_and_default_prompt_name
Converting a SentenceTransformer save appends a fresh classification head, so the source's embedding prompts and default_prompt_name are not inher
tests/cross_encoder/test_model.py:1114
Functiontest_conversion_ignores_source_prompts_and_similarity
Converting a MultiVectorEncoder save rebuilds this family's default modules, so the source's prompts, default_prompt_name and similarity_fn_name a
tests/sparse_encoder/test_model.py:918
Functiontest_conversion_ignores_source_similarity_fn_name
A SparseEncoder save carries similarity_fn_name="dot", chosen for sparse embeddings. The converted dense model scores with the dense default inste
tests/sentence_transformer/test_model.py:1438
Functiontest_convert_dense_sentence_transformer_resets_similarity_to_maxsim
A dense SentenceTransformer is converted to a MultiVectorEncoder on load. Its saved ``similarity_fn_name`` ("cosine" / "dot" can't score ragged pe
tests/multi_vector_encoder/test_model.py:772
Functiontest_convert_dense_st_with_dense_head_redirects_to_token_level
Converting a dense SentenceTransformer WITH a Dense head (LaBSE-shape) redirects the head to token level: the conversion drops the Pooling, so sen
tests/multi_vector_encoder/test_model.py:797
Methodtest_convert_kwargs_dropped_for_rank
rank no longer declares them, so they are dropped as no-ops and the warning names rank.
tests/util/test_decorators.py:347
Methodtest_convert_kwargs_kept_for_predict
predict still declares both, so they must pass through untouched and unwarned.
tests/util/test_decorators.py:337
Functiontest_convert_to_tensor_is_rejected_by_name
Variable-length embeddings cannot stack, so unlike every other model type there is no `convert_to_tensor` here. Copied-over calls are common enoug
tests/multi_vector_encoder/test_model.py:736
Functiontest_convert_to_tensor_negative_stride_still_rejected
Negative strides were never convertible and still are not. `torch.from_numpy` rejects them with the same error `torch.tensor` raised before the vi
tests/util/test_tensor.py:96
Functiontest_convert_to_tensor_read_only_array_does_not_warn
Read-only buffers (memmaps, broadcast views) take the copying path instead of torch's non-writable-tensor warning.
tests/util/test_tensor.py:104
Functiontest_convert_to_tensor_stacks_a_list_of_arrays_without_reading_elementwise
A list of arrays (what encoding with `convert_to_numpy=True` returns) went through `torch.tensor`, which reads it one element at a time. Stacking
tests/util/test_tensor.py:84
Functiontest_convert_to_tensor_views_numpy_without_copying
A numpy corpus is viewed, not copied: the copy costs as much as the scoring it feeds.
tests/util/test_tensor.py:77
Functiontest_cos_sim_sparse
Test cosine similarity between sparse and dense representations.
tests/util/test_similarity.py:151
Functiontest_cosent_penalises_the_lower_labelled_pair
The exponent is s(k,l) - s(i,j): the score of the pair with the *lower* expected similarity minus the score of the pair with the higher one, so ra
tests/sentence_transformer/losses/test_cosent.py:29
Methodtest_covers_non_text_pairs
Non-text pairs reach pair_to_messages by a different route, but take the same roles.
tests/base/test_modality.py:1717
Functiontest_create_minibatch_keeps_left_padding
Leading padding is left in place: trimming it would shift the positions a model derives from the sequence length.
tests/sentence_transformer/losses/test_gradcache.py:464
Functiontest_create_minibatch_keeps_query_expansion_positions
ColBERT expansion positions carry attention_mask 0 but are still scored, so the trailing trim must not drop them: doing so silently shrinks a 32-t
tests/sentence_transformer/losses/test_gradcache.py:475
Functiontest_create_minibatch_trims_trailing_padding
A padded mini-batch of short sequences must be embedded at its own width, not the whole batch's padded width, or mini_batch_num_tokens would not b
tests/sentence_transformer/losses/test_gradcache.py:444
Functiontest_cross_encoder
(reranker_bert_tiny_model: CrossEncoder)
tests/base/test_model.py:308
Functiontest_cross_encoder
Test using a cross-encoder for rescoring.
tests/util/test_hard_negatives.py:210
Methodtest_cross_encoder_custom_examples
(self)
tests/cross_encoder/test_model_card.py:204
Methodtest_cross_encoder_default_examples
(self)
tests/cross_encoder/test_model_card.py:194
Functiontest_cross_encoder_detailed
Test using a cross-encoder with different parameters.
tests/util/test_hard_negatives.py:230
Methodtest_cross_encoder_multi_label
Multi-label: shape includes num_labels, no rank section.
tests/cross_encoder/test_model_card.py:215
Methodtest_cross_encoder_multimodal_pairs
Multimodal pairs use _format_snippet_value for non-text elements and skip rank section.
tests/cross_encoder/test_model_card.py:245
Methodtest_cross_encoder_multimodal_run_usage_snippet
run_usage_snippet with multimodal pairs gracefully falls back when predict fails.
tests/cross_encoder/test_model_card.py:271
Methodtest_cross_encoder_single_label_has_rank
Single-label: shape is (n,), rank section present.
tests/cross_encoder/test_model_card.py:230
Functiontest_cross_encoder_with_multiple_positives_non_faiss
( multiple_positive_dataset: Dataset, static_retrieval_mrl_en_v1_model: SentenceTransformer, reran
tests/util/test_hard_negatives.py:256
Functiontest_csr_max_active_dims_passed_to_forward
(csr_bert_tiny_model: SparseEncoder, monkeypatch: pytest.MonkeyPatch)
tests/sparse_encoder/test_model.py:291
Functiontest_csr_outputs
(csr_bert_tiny_model: SparseEncoder, is_inference: bool, expected_keys: set)
tests/sparse_encoder/modules/test_csr.py:58
Methodtest_custom_keep_columns
(self, queries: Dataset, documents: Dataset)
tests/util/test_dataset.py:197
Methodtest_custom_model_id
(self)
tests/multi_vector_encoder/test_model_card.py:73
Methodtest_custom_role
(self)
tests/base/test_modality.py:509
Functiontest_data
()
tests/sentence_transformer/evaluation/test_information_retrieval_evaluator.py:62
Functiontest_data_collator
( stsb_bert_tiny_model: SentenceTransformer, stsb_dataset_dict: DatasetDict, has_bos_token: bool,
tests/sentence_transformer/test_trainer.py:702
Methodtest_data_uri
(self)
tests/base/test_modality.py:36
Methodtest_data_uri_not_affected_by_space_check
data: URIs bypass the URL space check and should still be detected as images.
tests/base/test_modality.py:337
Methodtest_dataset_dict_rejected
(self, queries: Dataset)
tests/util/test_dataset.py:260
Functiontest_decode_batch_with_empty_sample
(splade_bert_tiny_model: SparseEncoder)
tests/sparse_encoder/test_model.py:168
Functiontest_decode_empty_tensor
(splade_bert_tiny_model: SparseEncoder)
tests/sparse_encoder/test_model.py:134
Functiontest_decode_handles_sparse_dense_inputs
( splade_bert_tiny_model: SparseEncoder, texts: list[str] | str, format_type: str )
tests/sparse_encoder/test_model.py:99
Functiontest_decode_invalid_input_type
(splade_bert_tiny_model: SparseEncoder)
tests/sparse_encoder/test_model.py:155
Functiontest_decode_invalid_ndim
(splade_bert_tiny_model: SparseEncoder)
tests/sparse_encoder/test_model.py:161
Functiontest_decode_invalid_top_k
(splade_bert_tiny_model: SparseEncoder, top_k: int)
tests/sparse_encoder/test_model.py:148
Functiontest_decode_returns_sorted_weights
( splade_bert_tiny_model: SparseEncoder, texts: list[str] | str, top_k: int | None )
tests/sparse_encoder/test_model.py:192
← previousnext →2,101–2,200 of 3,759, ranked by callers