The opensearch backend uses OpenSearch as the vector store. Knowledge text is embedded by an embedding model, written to OpenSearch via LlamaIndex, and retrieved by similarity.
When to use
- You already run an OpenSearch cluster and want self-hosted vector search;
- You need persistence, sharing across processes and instances, and full control over indexing and retrieval.
Dependencies
Usage
Before running, configure MODEL_EMBEDDING_NAME, MODEL_EMBEDDING_DIM, MODEL_EMBEDDING_API_BASE, and MODEL_EMBEDDING_API_KEY. The extensions extra includes llama-index, embedding adapters, and vector-store connectors. Text is sent to the configured embedding service. For BytePlus or another provider, explicitly set the matching endpoint, model, and credentials; the default Ark endpoint does not automatically switch
Also configure the OpenSearch variables below and a valid CA certificate; grant index creation, write, and search permissions
You can also pass connection and embedding config explicitly via backend_config:
Parameters
KnowledgeBase parameters
Constructor parameters
backend_config supports the following settings:
OpenSearch connection config
opensearch_config is an OpensearchConfig with env prefix DATABASE_OPENSEARCH_:
Embedding config
embedding_config is an EmbeddingModelConfig with env prefix MODEL_EMBEDDING_:
Environment variables
index must follow OpenSearch naming rules: all lowercase, only a-z0-9_-., and not starting with _ or -; otherwise initialization fails. In production, set cert_path to enable certificate verification and avoid security risks.
The example search should return the annual-leave policy. For empty results, check successful ingestion, matching embedding dimensions, completed server processing, network access, and permissions. Managed ingestion may not be immediately searchable. Running the configured Agent also requires model credentials