Metis can talk directly to llama.cpp's HTTP server, letting you run chat and embedding models locally without any cloud dependency.
The llama.cpp server can be used as a drop-in replacement for OpenAI-compatible providers and exposes the following endpoints:
/v1/chat/completions/v1/completions/v1/embeddings/v1/models/v1/responses
-
Install the llama.cpp server.
-
Pull at least one chat model and one embedding model (GGUF format).
- Chat model examples
- e.g. for 8GB system,
Llama3.1-8B - e.g. for 16GB system,
Qwen3.5-9B - e.g. for 24GB system,
Qwen3.6-35B-A3B
- e.g. for 8GB system,
- Embedding model example
- e.g.
nomic-embed-text-v1.5
- e.g.
- Chat model examples
-
Start the llama.cpp server with both the chat and embedding models. The server can serve multiple models simultaneously.
llama-server \ --model /path/to/model.gguf \ --ctx-size 128000 \ --embedding
The
--embeddingflag enables the embeddings endpoint. Without it, Metis cannot build the vector store. The--ctx-sizeshould match or exceedmetis_engine.max_token_length.
Add or adjust the llm_provider block in your metis.yaml:
llm_provider:
name: "llamacpp"
base_url: "http://localhost:8080/v1"
model: "Llama3.1-8B"
embedding_provider:
name: "llamacpp"
base_url: "http://localhost:8080/v1"
code_embedding_model: "nomic-embed-text-v1.5"
docs_embedding_model: "nomic-embed-text-v1.5"base_urldefaults tohttp://localhost:8080/v1if not configured.namemust be"llamacpp"(case-insensitive).modelis required underllm_provider.code_embedding_model/docs_embedding_modelare required underembedding_provideronly when the Index capability cannot reuse embedding models supplied by the vector backend.- An API key is not required by the llama.cpp server; Metis uses a placeholder by default.
Once the server responds, run uv run metis --interactive --codebase-path <path> (or metis --interactive inside your virtual environment) and use the usual index, review_code, review_dir or review_file commands. Metis will route model requests through the OpenAI Responses API and embedding requests through the OpenAI-compatible embeddings API.