Local OpenAI-Compatible APIs: Ollama, vLLM, LM Studio, and llama.cpp

Use the runtime's OpenAI-compatible /v1 root. Check /v1/models before chat. Local endpoints stay in browser-direct mode, so browser tests also need a correct CORS policy.

Published
Updated
Reading time
9 minutes
OLLAMAlocalhost:11434/v1
LM STUDIOlocalhost:1234/v1
VLLMlocalhost:8000/v1
LLAMA.CPPlocalhost:8080/v1

Use the Correct Local API Root

RuntimeDefault OpenAI-compatible Base URLAPI key
Ollamahttp://localhost:11434/v1Placeholder value for SDKs
LM Studiohttp://localhost:1234/v1Not required unless enabled
vLLMhttp://localhost:8000/v1Matches --api-key
llama.cpphttp://localhost:8080/v1Matches --api-key if set
Run the model-list preflight

Start the Server with an Explicit Port

RuntimeExample start commandOfficial source
Ollamaollama serveOpenAI compatibility
LM Studiolms server start --port 1234 --corsServer CLI
vLLMvllm serve MODEL --api-key token-abc123OpenAI server
llama.cppllama-server -m model.gguf --port 8080Server README

Commands and defaults were checked against the linked primary sources on 12 Aug 2026.

Read the Model ID Before You Send Chat

A display name is not always the API model ID. Read /v1/models. Copy the returned id exactly, including case, organization, and tags.

Model List requests
curl http://localhost:11434/v1/models

# Replace the address for another runtime:
# LM Studio  http://localhost:1234/v1/models
# vLLM       http://localhost:8000/v1/models
# llama.cpp  http://localhost:8080/v1/models
Minimal chat request
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MODEL_ID_FROM_V1_MODELS",
    "messages": [{"role": "user", "content": "Reply with OK"}]
  }'
Check model loading rules

LM Studio can list downloaded or loaded models based on its JIT setting. vLLM exposes the served model name. Ollama model IDs can include tags.

Allow the Browser Origin without Exposing the Server

Command-line clients do not use CORS. Browser-direct tests do. Allow the exact page origin and keep the server bound to localhost unless you need network access.

RuntimeBrowser access settingSecurity note
OllamaOLLAMA_ORIGINS="http://localhost:3000" ollama serveUse the exact site origin
LM Studiolms server start --port 1234 --corsEnable authentication if you expose it
vLLM--allowed-origins '["http://localhost:3000"]'A restricted list is safer than a wildcard
llama.cpp--cors-origins http://localhost:3000Restrict the default wildcard

See the Ollama origins setting, LM Studio server settings, vLLM server arguments, and llama.cpp server arguments.

Verify the Runtime and Model Separately

LLMCompat checks the API root and Model List before it sends billable or compute-heavy probes. It then tests chat, streaming, tools, and JSON Mode as separate contracts.