Skip to content

Cloud Models ​

Ollama's cloud models are a new kind of model in Ollama that can run without a powerful GPU. Instead, cloud models are automatically offloaded to Ollama's cloud service while offering the same capabilities as local models, making it possible to keep using your local tools while running larger models that wouldn't fit on a personal computer.

Supported models ​

For a list of supported models, see Ollama's model library.

Running Cloud models ​

Ollama's cloud models require an account on ollama.com. To sign in or create an account, run:

ollama signin

CLI ​

To run a cloud model, open the terminal and run:

ollama run gpt-oss:120b-cloud

Python ​

First, pull a cloud model so it can be accessed:

ollama pull gpt-oss:120b-cloud

Next, install Ollama's Python library:

pip install ollama

Next, create and run a simple Python script:

python
from ollama import Client

client = Client()

messages = [
  {
    'role': 'user',
    'content': 'Why is the sky blue?',
  },
]

for part in client.chat('gpt-oss:120b-cloud', messages=messages, stream=True):
  print(part['message']['content'], end='', flush=True)

JavaScript ​

First, pull a cloud model so it can be accessed:

ollama pull gpt-oss:120b-cloud

Next, install Ollama's JavaScript library:

npm i ollama

Then use the library to run a cloud model:

typescript


const ollama = new Ollama();

const response = await ollama.chat({
  model: "gpt-oss:120b-cloud",
  messages: [{ role: "user", content: "Explain quantum computing" }],
  stream: true,
});

for await (const part of response) {
  process.stdout.write(part.message.content);
}

cURL ​

First, pull a cloud model so it can be accessed:

ollama pull gpt-oss:120b-cloud

Run the following cURL command to run the command via Ollama's API:

curl http://localhost:11434/api/chat -d '{
  "model": "gpt-oss:120b-cloud",
  "messages": [{
    "role": "user",
    "content": "Why is the sky blue?"
  }],
  "stream": false
}'

Cloud API access ​

Cloud models can also be accessed directly on ollama.com's API. In this mode, ollama.com acts as a remote Ollama host.

Authentication ​

For direct access to ollama.com's API, first create an API key.

Then, set the OLLAMA_API_KEY environment variable to your API key.

export OLLAMA_API_KEY=your_api_key

Listing models ​

For models available directly via Ollama's API, models can be listed via:

curl https://ollama.com/api/tags

Generating a response ​

Python ​

First, install Ollama's Python library

pip install ollama

Then make a request

python

from ollama import Client

client = Client(
    host="https://ollama.com",
    headers={'Authorization': 'Bearer ' + os.environ.get('OLLAMA_API_KEY')}
)

messages = [
  {
    'role': 'user',
    'content': 'Why is the sky blue?',
  },
]

for part in client.chat('gpt-oss:120b', messages=messages, stream=True):
  print(part['message']['content'], end='', flush=True)

JavaScript ​

First, install Ollama's JavaScript library:

npm i ollama

Next, make a request to the model:

typescript


const ollama = new Ollama({
  host: "https://ollama.com",
  headers: {
    Authorization: "Bearer " + process.env.OLLAMA_API_KEY,
  },
});

const response = await ollama.chat({
  model: "gpt-oss:120b",
  messages: [{ role: "user", content: "Explain quantum computing" }],
  stream: true,
});

for await (const part of response) {
  process.stdout.write(part.message.content);
}

cURL ​

Generate a response via Ollama's chat API:

curl https://ollama.com/api/chat \
  -H "Authorization: Bearer $OLLAMA_API_KEY" \
  -d '{
    "model": "gpt-oss:120b",
    "messages": [{
      "role": "user",
      "content": "Why is the sky blue?"
    }],
    "stream": false
  }'

Local only ​

Ollama can run in local-only mode by disabling Ollama's cloud features.

Retirements ​

Ollama will occasionally deprecate and retire older cloud models as newer and better open-source models are released. Tools and applications relying on Ollama Cloud models may need to be updated to keep working. Impacted users will be notified in advance of model deprecation and retirement. Deprecations will be communicated through email and on the Ollama website.

Ollama Cloud model retirement does not affect local models.

Upcoming retirements ​

Retirement dateModelRecommended alternative
July 31, 2026minimax-m2.5minimax-m2.7
July 31, 2026kimi-k2.5kimi-k2.6

Past retirements ​

July 15, 2026 | Model | Recommended alternative | | --- | --- | | `deepseek-v3.1:671b` | `deepseek-v4-flash` | | `deepseek-v3.2` | `deepseek-v4-flash` | | `devstral-2:123b` | `mistral-large-3:675b` | | `devstral-small-2:24b` | | | `ministral-3:14b` | | | `ministral-3:3b` | | | `ministral-3:8b` | | | `gemini-3-flash-preview` | `minimax-m3` | | `gemma3:12b` | `gemma4:31b` | | `gemma3:27b` | `gemma4:31b` | | `gemma3:4b` | `gemma4:31b` | | `glm-4.7` | `glm-5.2` | | `glm-5` | `glm-5.2` | | `minimax-m2.1` | `minimax-m3` | | `qwen3-coder-next` | `qwen3.5:397b` | | `qwen3-coder:480b` | `qwen3.5:397b` |
June 30, 2026 | Model | Recommended alternative | | --- | --- | | `rnj-1:8b` | |
June 16, 2026 | Model | Recommended alternative | | --- | --- | | `kimi-k2-thinking` | `kimi-k2.6` | | `kimi-k2:1t` | `kimi-k2.6` | | `minimax-m2` | `minimax-m3` | | `glm-4.6` | `glm-5.1` | | `qwen3-next:80b` | `qwen3.5` | | `qwen3-vl:235b` | `qwen3.5` | | `qwen3-vl:235b-instruct` | `qwen3.5` | | `cogito-2.1:671b` | `deepseek-v4-flash` |