Skip to content
API ReferenceZilliz CloudMilvusAttu

Full-Text Search

Full-text search enables keyword-based text retrieval using BM25 scoring. Milvus supports this through collection functions that automatically tokenize text and convert it to sparse vectors at insert and query time.

  1. You define a VarChar field for text and a SparseFloatVector field for the generated vectors
  2. You add a BM25 function to the collection that maps the text field to the sparse vector field
  3. When you insert text, Milvus automatically tokenizes it and generates sparse vectors
  4. When you search, you pass text queries and Milvus automatically converts them to sparse vectors for matching

Before creating a collection, you can test how an analyzer tokenizes text using runAnalyzer():

const result = await client.runAnalyzer({
text: ['machine learning is great', 'deep learning fundamentals'],
analyzer_params: {
type: 'standard',
},
});
console.log('Results:', result.results);
// Each text input produces an AnalyzerResult with tokens

You can also test analyzers that are configured on a collection field, request detailed token data, and include token hashes:

const result = await client.runAnalyzer({
db_name: 'default',
collection_name: 'articles',
field_name: 'body',
analyzer_names: ['english_analyzer'],
text: ['Running analyzers with offsets and hashes'],
with_detail: true,
with_hash: true,
});
result.results[0].tokens.forEach((token) => {
console.log(token.token, token.start_offset, token.end_offset, token.hash);
});

runAnalyzer() accepts either analyzer_params for an ad-hoc analyzer configuration or collection context fields (db_name, collection_name, field_name, analyzer_names) to use analyzers defined in a schema.

Type Description
standard Standard tokenizer with lowercase filter
english English-specific tokenizer with stemming
chinese Chinese text tokenizer
Milvus 3.0

For analyzers that load a custom dictionary or synonym file, register the server-visible object before referencing it from analyzer parameters:

await client.addFileResource({
name: 'product_dictionary',
path: 'analyzer-resources/product-dictionary.txt',
});
const resources = await client.listFileResources();
console.log(resources.resources);

The file must already be available through the storage configured for Milvus. addFileResource() registers metadata; it does not upload a local Node.js file. Remove the registration when it is no longer used:

await client.removeFileResource({ name: 'product_dictionary' });
import {
MilvusClient,
DataType,
FunctionType,
IndexType,
MetricType,
ConsistencyLevelEnum,
} from '@zilliz/milvus2-sdk-node';
const client = new MilvusClient({ address: 'localhost:19530' });
// 1. Create collection with text and sparse vector fields
await client.createCollection({
collection_name: 'articles',
fields: [
{
name: 'id',
data_type: DataType.Int64,
is_primary_key: true,
autoID: true,
},
{
name: 'title',
data_type: DataType.VarChar,
max_length: 256,
enable_analyzer: true,
},
{
name: 'body',
data_type: DataType.VarChar,
max_length: 10000,
enable_analyzer: true,
},
{ name: 'sparse_vector', data_type: DataType.SparseFloatVector },
],
functions: [
{
name: 'bm25',
type: FunctionType.BM25,
input_field_names: ['body'],
output_field_names: ['sparse_vector'],
},
],
});

You can add the BM25 output field, function, and bound index to an existing collection in one schema alteration:

await client.addFunctionField({
collection_name: 'articles',
field: {
name: 'title_sparse',
data_type: DataType.SparseFloatVector,
is_function_output: true,
},
function: {
name: 'title_bm25',
type: FunctionType.BM25,
input_field_names: ['title'],
output_field_names: ['title_sparse'],
params: {},
},
index_name: 'title_bm25_index',
extra_params: {
index_type: IndexType.SPARSE_INVERTED_INDEX,
metric_type: MetricType.BM25,
},
});

Pass a text string as the search data. Milvus uses the BM25 function to convert it to a sparse vector automatically. The example uses Strong consistency and flushes the inserted documents before searching so that the small test dataset is ready for BM25 retrieval. Avoid flushing after every insert in a production ingestion pipeline.

// Create index and load
await client.createIndex({
collection_name: 'articles',
field_name: 'sparse_vector',
index_type: 'SPARSE_INVERTED_INDEX',
metric_type: 'BM25',
});
await client.loadCollectionSync({ collection_name: 'articles' });
// Insert documents
await client.insert({
collection_name: 'articles',
data: [
{
title: 'Introduction to Machine Learning',
body: 'Machine learning is a branch of AI...',
},
{
title: 'Deep Learning Basics',
body: 'Deep learning uses neural networks...',
},
],
});
// Wait for the example documents to be flushed before searching
await client.flushSync({ collection_names: ['articles'] });
// Search by text
const results = await client.search({
collection_name: 'articles',
data: ['machine learning algorithms'],
anns_field: 'sparse_vector',
metric_type: 'BM25',
consistency_level: ConsistencyLevelEnum.Strong,
limit: 10,
output_fields: ['title', 'body'],
});
console.log('Results:', results.results);

Use the highlighter parameter to highlight matched text fragments in search results.

Highlights keyword matches using the text field’s analyzer. Set highlight_search_text: true to highlight the BM25 search terms, and explicitly set metric_type: 'BM25' on the search request. Creating a BM25 index alone does not supply this request parameter to the highlighter.

In this collection, the BM25 function takes body as its input, so search-term highlighting applies to body. Including title in output_fields returns its source text but does not automatically highlight it.

const results = await client.search({
collection_name: 'articles',
data: ['machine learning'],
anns_field: 'sparse_vector',
metric_type: 'BM25',
consistency_level: ConsistencyLevelEnum.Strong,
limit: 10,
output_fields: ['title', 'body'],
highlighter: {
type: 0, // HighlightType.Lexical
highlight_search_text: true,
pre_tags: ['<em>'],
post_tags: ['</em>'],
fragment_size: 100,
num_of_fragments: 3,
},
});

For the first example document, results.results[0].highlight.body contains:

{
"fragments": ["<em>Machine</em> <em>learning</em> is a branch of AI..."],
"scores": []
}

Highlights semantically similar text through a model service. Milvus must have the Zilliz model-service provider endpoint configured on the server, with access to an appropriate highlighting model deployment. The highlighter parameter does not configure or start that service.

A standalone Milvus instance without this provider configuration returns an error such as Create SemanticHighlight failed: zilliz client config error, lost endpoint config. Use the following example only after configuring the service, and replace YOUR_HIGHLIGHT_MODEL_DEPLOYMENT_ID with your deployment ID:

const results = await client.search({
collection_name: 'articles',
data: ['machine learning'],
anns_field: 'sparse_vector',
metric_type: 'BM25',
consistency_level: ConsistencyLevelEnum.Strong,
limit: 10,
output_fields: ['title', 'body'],
highlighter: {
type: 1, // HighlightType.Semantic
model_deployment_id: 'YOUR_HIGHLIGHT_MODEL_DEPLOYMENT_ID',
queries: ['machine learning'],
input_fields: ['body'],
pre_tags: ['<mark>'],
post_tags: ['</mark>'],
threshold: 0.5,
},
});

queries must contain one string for each search query. Set highlight_only: true to request only highlighted fragments from the model service. The source fields requested through output_fields are still returned in the search hits.

Search hits include highlight data for the fields processed by the highlighter. Each highlighted field contains fragments and scores arrays. Lexical highlighting returns an empty scores array; semantic highlighting can return scores from the model service.

results.results.forEach((hit) => {
console.log('Title:', hit.title);
console.log('Highlight:', hit.highlight);
const bodyHighlight = hit.highlight?.body;
bodyHighlight?.fragments.forEach((fragment, index) => {
console.log('Fragment:', fragment);
const score = bodyHighlight.scores[index];
if (score !== undefined) {
console.log('Score:', score);
}
});
});
// List functions on a collection
const desc = await client.describeCollection({ collection_name: 'articles' });
console.log('Functions:', desc.schema.functions);
// Alter a function
await client.alterCollectionFunction({
collection_name: 'articles',
function_name: 'bm25',
function: {
name: 'bm25',
type: FunctionType.BM25,
input_field_names: ['body'],
output_field_names: ['sparse_vector'],
params: { key: 'new_value' },
},
});
// Drop the function, output field, and bound index
await client.dropFunctionField({
collection_name: 'articles',
function_name: 'bm25',
});