Use the Correct Local API Root
| Runtime | Default OpenAI-compatible Base URL | API key |
|---|---|---|
| Ollama | http://localhost:11434/v1 | Placeholder value for SDKs |
| LM Studio | http://localhost:1234/v1 | Not required unless enabled |
| vLLM | http://localhost:8000/v1 | Matches --api-key |
| llama.cpp | http://localhost:8080/v1 | Matches --api-key if set |
Start the Server with an Explicit Port
| Runtime | Example start command | Official source |
|---|---|---|
| Ollama | ollama serve | OpenAI compatibility |
| LM Studio | lms server start --port 1234 --cors | Server CLI |
| vLLM | vllm serve MODEL --api-key token-abc123 | OpenAI server |
| llama.cpp | llama-server -m model.gguf --port 8080 | Server README |
Commands and defaults were checked against the linked primary sources on 12 Aug 2026.
Read the Model ID Before You Send Chat
A display name is not always the API model ID. Read /v1/models. Copy the returned id exactly, including case, organization, and tags.
curl http://localhost:11434/v1/models
# Replace the address for another runtime:
# LM Studio http://localhost:1234/v1/models
# vLLM http://localhost:8000/v1/models
# llama.cpp http://localhost:8080/v1/modelscurl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "MODEL_ID_FROM_V1_MODELS",
"messages": [{"role": "user", "content": "Reply with OK"}]
}'LM Studio can list downloaded or loaded models based on its JIT setting. vLLM exposes the served model name. Ollama model IDs can include tags.
Allow the Browser Origin without Exposing the Server
Command-line clients do not use CORS. Browser-direct tests do. Allow the exact page origin and keep the server bound to localhost unless you need network access.
| Runtime | Browser access setting | Security note |
|---|---|---|
| Ollama | OLLAMA_ORIGINS="http://localhost:3000" ollama serve | Use the exact site origin |
| LM Studio | lms server start --port 1234 --cors | Enable authentication if you expose it |
| vLLM | --allowed-origins '["http://localhost:3000"]' | A restricted list is safer than a wildcard |
| llama.cpp | --cors-origins http://localhost:3000 | Restrict the default wildcard |
See the Ollama origins setting, LM Studio server settings, vLLM server arguments, and llama.cpp server arguments.
Verify the Runtime and Model Separately
LLMCompat checks the API root and Model List before it sends billable or compute-heavy probes. It then tests chat, streaming, tools, and JSON Mode as separate contracts.