Skip to content

API ​

Nota: la documentazione API di Ollama è in fase di migrazione su https://docs.ollama.com/api

Endpoints ​

Convenzioni ​

Nomi dei modelli ​

I nomi dei modelli seguono il formato model:tag, dove model può avere un namespace opzionale come example/model. Alcuni esempi sono orca-mini:3b-q8_0 e llama3:70b. Il tag è opzionale e, se non fornito, avrà come valore predefinito latest. Il tag viene utilizzato per identificare una versione specifica.

Durate ​

Tutte le durate sono restituite in nanosecondi.

Risposte in streaming ​

Alcuni endpoint restituiscono risposte in streaming come oggetti JSON. Lo streaming può essere disabilitato fornendo {"stream": false} per questi endpoint.

Genera un completamento ​

POST /api/generate

Genera una risposta per un prompt fornito con un modello specificato. Si tratta di un endpoint in streaming, quindi verrà restituita una serie di risposte. L'oggetto di risposta finale includerà statistiche e dati aggiuntivi relativi alla richiesta.

Parametri ​

  • model: (obbligatorio) il nome del modello
  • prompt: il prompt per cui generare una risposta
  • suffix: il testo dopo la risposta del modello
  • images: (opzionale) un elenco di immagini codificate in base64 (per modelli multimodali come llava)
  • think: (per modelli di pensiero) il modello deve pensare prima di rispondere? Può essere un valore booleano o un livello di pensiero ("low", "medium", "high" o "max")

Parametri avanzati (opzionali):

  • format: il formato in cui restituire la risposta. Il formato può essere json o uno schema JSON
  • options: parametri aggiuntivi del modello elencati nella documentazione del Modelfile come temperature
  • system: messaggio di sistema (sovrascrive quello definito nel Modelfile)
  • template: il template di prompt da utilizzare (sovrascrive quello definito nel Modelfile)
  • stream: se false la risposta verrà restituita come un singolo oggetto di risposta, invece di un flusso di oggetti
  • raw: se true non verrà applicata alcuna formattazione al prompt. Puoi scegliere di utilizzare il parametro raw se stai specificando un prompt completo con template nella richiesta all'API
  • keep_alive: controlla per quanto tempo il modello rimarrà caricato in memoria dopo la richiesta (predefinito: 5m)
  • context (deprecato): il parametro context restituito da una richiesta precedente a /generate, può essere utilizzato per mantenere una memoria conversazionale breve

Parametri sperimentali per la generazione di immagini (solo per modelli di generazione di immagini):

WARNING

Questi parametri sono sperimentali e potrebbero cambiare nelle versioni future.

  • width: larghezza dell'immagine generata in pixel
  • height: altezza dell'immagine generata in pixel
  • steps: numero di passaggi di diffusione

Output strutturati ​

Gli output strutturati sono supportati fornendo uno schema JSON nel parametro format. Il modello genererà una risposta che corrisponde allo schema. Vedi l'esempio di output strutturati di seguito.

Modalità JSON ​

Abilita la modalità JSON impostando il parametro format su json. In questo modo la risposta sarà strutturata come un oggetto JSON valido. Vedi l'esempio di modalità JSON di seguito.

IMPORTANT

È importante istruire il modello a utilizzare JSON nel prompt. In caso contrario, il modello potrebbe generare grandi quantità di spazi bianchi.

Esempi ​

Richiesta di generazione (in streaming) ​

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Why is the sky blue?"
}'
Risposta ​

Viene restituito un flusso di oggetti JSON:

json
{
  "model": "llama3.2",
  "created_at": "2023-08-04T08:52:19.385406455-07:00",
  "response": "The",
  "done": false
}

La risposta finale nel flusso include anche dati aggiuntivi relativi alla generazione:

  • total_duration: tempo impiegato per generare la risposta
  • load_duration: tempo impiegato in nanosecondi per caricare il modello
  • prompt_eval_count: numero di token nel prompt
  • prompt_eval_duration: tempo impiegato in nanosecondi per valutare il prompt
  • eval_count: numero di token nella risposta
  • eval_duration: tempo in nanosecondi impiegato per generare la risposta
  • context: una codifica della conversazione utilizzata in questa risposta, può essere inviata nella richiesta successiva per mantenere una memoria conversazionale
  • response: vuota se la risposta è stata trasmessa in streaming, altrimenti contiene la risposta completa

Per calcolare la velocità di generazione della risposta in token al secondo (token/s), dividi eval_count / eval_duration * 10^9.

json
{
  "model": "llama3.2",
  "created_at": "2023-08-04T19:22:45.499127Z",
  "response": "",
  "done": true,
  "context": [1, 2, 3],
  "total_duration": 10706818083,
  "load_duration": 6338219291,
  "prompt_eval_count": 26,
  "prompt_eval_duration": 130079000,
  "eval_count": 259,
  "eval_duration": 4232710000
}

Richiesta (nessuno streaming) ​

Richiesta ​

È possibile ricevere una risposta in una singola risposta quando lo streaming è disattivato.

shell
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Why is the sky blue?",
  "stream": false
}'
Risposta ​

Se stream è impostato su false, la risposta sarà un singolo oggetto JSON:

json
{
  "model": "llama3.2",
  "created_at": "2023-08-04T19:22:45.499127Z",
  "response": "The sky is blue because it is the color of the sky.",
  "done": true,
  "context": [1, 2, 3],
  "total_duration": 5043500667,
  "load_duration": 5025959,
  "prompt_eval_count": 26,
  "prompt_eval_duration": 325953000,
  "eval_count": 290,
  "eval_duration": 4709213000
}

Richiesta (con suffisso) ​

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "codellama:code",
  "prompt": "def compute_gcd(a, b):",
  "suffix": "    return result",
  "options": {
    "temperature": 0
  },
  "stream": false
}'
Risposta ​
json5
{
  "model": "codellama:code",
  "created_at": "2024-07-22T20:47:51.147561Z",
  "response": "\n  if a == 0:\n    return b\n  else:\n    return compute_gcd(b % a, a)\n\ndef compute_lcm(a, b):\n  result = (a * b) / compute_gcd(a, b)\n",
  "done": true,
  "done_reason": "stop",
  "context": [...],
  "total_duration": 1162761250,
  "load_duration": 6683708,
  "prompt_eval_count": 17,
  "prompt_eval_duration": 201222000,
  "eval_count": 63,
  "eval_duration": 953997000
}

Richiesta (output strutturati) ​

Richiesta ​
shell
curl -X POST http://localhost:11434/api/generate -H "Content-Type: application/json" -d '{
  "model": "llama3.1:8b",
  "prompt": "Ollama is 22 years old and is busy saving the world. Respond using JSON",
  "stream": false,
  "format": {
    "type": "object",
    "properties": {
      "age": {
        "type": "integer"
      },
      "available": {
        "type": "boolean"
      }
    },
    "required": [
      "age",
      "available"
    ]
  }
}'
Risposta ​
json
{
  "model": "llama3.1:8b",
  "created_at": "2024-12-06T00:48:09.983619Z",
  "response": "{\n  \"age\": 22,\n  \"available\": true\n}",
  "done": true,
  "done_reason": "stop",
  "context": [1, 2, 3],
  "total_duration": 1075509083,
  "load_duration": 567678166,
  "prompt_eval_count": 28,
  "prompt_eval_duration": 236000000,
  "eval_count": 16,
  "eval_duration": 269000000
}

Richiesta (modalità JSON) ​

IMPORTANT

Quando format è impostato su json, l'output sarà sempre un oggetto JSON ben formato. È importante anche istruire il modello a rispondere in JSON.

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "What color is the sky at different times of the day? Respond using JSON",
  "format": "json",
  "stream": false
}'
Risposta ​
json
{
  "model": "llama3.2",
  "created_at": "2023-11-09T21:07:55.186497Z",
  "response": "{\n\"morning\": {\n\"color\": \"blue\"\n},\n\"noon\": {\n\"color\": \"blue-gray\"\n},\n\"afternoon\": {\n\"color\": \"warm gray\"\n},\n\"evening\": {\n\"color\": \"orange\"\n}\n}\n",
  "done": true,
  "context": [1, 2, 3],
  "total_duration": 4648158584,
  "load_duration": 4071084,
  "prompt_eval_count": 36,
  "prompt_eval_duration": 439038000,
  "eval_count": 180,
  "eval_duration": 4196918000
}

Il valore di response sarà una stringa contenente JSON simile a:

json
{
  "morning": {
    "color": "blue"
  },
  "noon": {
    "color": "blue-gray"
  },
  "afternoon": {
    "color": "warm gray"
  },
  "evening": {
    "color": "orange"
  }
}

Richiesta (con immagini) ​

Per inviare immagini a modelli multimodali come llava o bakllava, fornisci un elenco di images codificate in base64:

Richiesta ​

shell
curl http://localhost:11434/api/generate -d '{
  "model": "llava",
  "prompt":"What is in this picture?",
  "stream": false,
  "images": ["iVBORw0KGgoAAAANSUhEUgAAAG0AAABmCAYAAADBPx+VAAAACXBIWXMAAAsTAAALEwEAmpwYAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAA3VSURBVHgB7Z27r0zdG8fX743i1bi1ikMoFMQloXRpKFFIqI7LH4BEQ+NWIkjQuSWCRIEoULk0gsK1kCBI0IhrQVT7tz/7zZo888yz1r7MnDl7z5xvsjkzs2fP3uu71nNfa7lkAsm7d++Sffv2JbNmzUqcc8m0adOSzZs3Z+/XES4ZckAWJEGWPiCxjsQNLWmQsWjRIpMseaxcuTKpG/7HP27I8P79e7dq1ars/yL4/v27S0ejqwv+cUOGEGGpKHR37tzJCEpHV9tnT58+dXXCJDdECBE2Ojrqjh071hpNECjx4cMHVycM1Uhbv359B2F79+51586daxN/+pyRkRFXKyRDAqxEp4yMlDDzXG1NPnnyJKkThoK0VFd1ELZu3TrzXKxKfW7dMBQ6bcuWLW2v0VlHjx41z717927ba22U9APcw7Nnz1oGEPeL3m3p2mTAYYnFmMOMXybPPXv2bNIPpFZr1NHn4HMw0KRBjg9NuRw95s8PEcz/6DZELQd/09C9QGq5RsmSRybqkwHGjh07OsJSsYYm3ijPpyHzoiacg35MLdDSIS/O1yM778jOTwYUkKNHWUzUWaOsylE00MyI0fcnOwIdjvtNdW/HZwNLGg+sR1kMepSNJXmIwxBZiG8tDTpEZzKg0GItNsosY8USkxDhD0Rinuiko2gfL/RbiD2LZAjU9zKQJj8RDR0vJBR1/Phx9+PHj9Z7REF4nTZkxzX4LCXHrV271qXkBAPGfP/atWvu/PnzHe4C97F48eIsRLZ9+3a3f/9+87dwP1JxaF7/3r17ba+5l4EcaVo0lj3SBq5kGTJSQmLWMjgYNei2GPT1MuMqGTDEFHzeQSP2wi/jGnkmPJ/nhccs44jvDAxpVcxnq0F6eT8h4ni/iIWpR5lPyA6ETkNXoSukvpJAD3AsXLiwpZs49+fPn5ke4j10TqYvegSfn0OnafC+Tv9ooA/JPkgQysqQNBzagXY55nO/oa1F7qvIPWkRL12WRpMWUvpVDYmxAPehxWSe8ZEXL20sadYIozfmNch4QJPAfeJgW3rNsnzphBKNJM2KKODo1rVOMRYik5ETy3ix4qWNI81qAAirizgMIc+yhTytx0JWZuNI03qsrgWlGtwjoS9XwgUhWGyhUaRZZQNNIEwCiXD16tXcAHUs79co0vSD8rrJCIW98pzvxpAWyyo3HYwqS0+H0BjStClcZJT5coMm6D2LOF8TolGJtK9fvyZpyiC5ePFi9nc/oJU4eiEP0jVoAnHa9wyJycITMP78+eMeP37sXrx44d6+fdt6f82aNdkx1pg9e3Zb5W+RSRE+n+VjksQWifvVaTKFhn5O8my63K8Qabdv33b379/PiAP//vuvW7BggZszZ072/+TJk91YgkafPn166zXB1rQHFvouAWHq9z3SEevSUerqCn2/dDCeta2jxYbr69evk4MHDyY7d+7MjhMnTiTPnz9Pfv/+nfQT2ggpO2dMF8cghuoM7Ygj5iWCqRlGFml0QC/ftGmTmzt3rmsaKDsgBSPh0/8yPeLLBihLkOKJc0jp8H8vUzcxIA1k6QJ/c78tWEyj5P3o4u9+jywNPdJi5rAH9x0KHcl4Hg570eQp3+vHXGyrmEeigzQsQsjavXt38ujRo44LQuDDhw+TW7duRS1HGgMxhNXHgflaNTOsHyKvHK5Ijo2jbFjJBQK9YwFd6RVMzfgRBmEfP37suBBm/p49e1qjEP2mwTViNRo0VJWH1deMXcNK08uUjVUu7s/zRaL+oLNxz1bpANco4npUgX4G2eFbpDFyQoQxojBCpEGSytmOH8qrH5Q9vuzD6ofQylkCUmh8DBAr+q8JCyVNtWQIidKQE9wNtLSQnS4jDSsxNHogzFuQBw4cyM61UKVsjfr3ooBkPSqqQHesUPWVtzi9/vQi1T+rJj7WiTz4Pt/l3LxUkr5P2VYZaZ4URpsE+st/dujQoaBBYokbrz/8TJNQYLSonrPS9kUaSkPeZyj1AWSj+d+VBoy1pIWVNed8P0Ll/ee5HdGRhrHhR5GGN0r4LGZBaj8oFDJitBTJzIZgFcmU0Y8ytWMZMzJOaXUSrUs5RxKnrxmbb5YXO9VGUhtpXldhEUogFr3IzIsvlpmdosVcGVGXFWp2oU9kLFL3dEkSz6NHEY1sjSRdIuDFWEhd8KxFqsRi1uM/nz9/zpxnwlESONdg6dKlbsaMGS4EHFHtjFIDHwKOo46l4TxSuxgDzi+rE2jg+BaFruOX4HXa0Nnf1lwAPufZeF8/r6zD97WK2qFnGjBxTw5qNGPxT+5T/r7/7RawFC3j4vTp09koCxkeHjqbHJqArmH5UrFKKksnxrK7FuRIs8STfBZv+luugXZ2pR/pP9Ois4z+TiMzUUkUjD0iEi1fzX8GmXyuxUBRcaUfykV0YZnlJGKQpOiGB76x5GeWkWWJc3mOrK6S7xdND+W5N6XyaRgtWJFe13GkaZnKOsYqGdOVVVbGupsyA/l7emTLHi7vwTdirNEt0qxnzAvBFcnQF16xh/TMpUuXHDowhlA9vQVraQhkudRdzOnK+04ZSP3DUhVSP61YsaLtd/ks7ZgtPcXqPqEafHkdqa84X6aCeL7YWlv6edGFHb+ZFICPlljHhg0bKuk0CSvVznWsotRu433alNdFrqG45ejoaPCaUkWERpLXjzFL2Rpllp7PJU2a/v7Ab8N05/9t27Z16KUqoFGsxnI9EosS2niSYg9SpU6B4JgTrvVW1flt1sT+0ADIJU2maXzcUTraGCRaL1Wp9rUMk16PMom8QhruxzvZIegJjFU7LLCePfS8uaQdPny4jTTL0dbee5mYokQsXTIWNY46kuMbnt8Kmec+LGWtOVIl9cT1rCB0V8WqkjAsRwta93TbwNYoGKsUSChN44lgBNCoHLHzquYKrU6qZ8lolCIN0Rh6cP0Q3U6I6IXILYOQI513hJaSKAorFpuHXJNfVlpRtmYBk1Su1obZr5dnKAO+L10Hrj3WZW+E3qh6IszE37F6EB+68mGpvKm4eb9bFrlzrok7fvr0Kfv727dvWRmdVTJHw0qiiCUSZ6wCK+7XL/AcsgNyL74DQQ730sv78Su7+t/A36MdY0sW5o40ahslXr58aZ5HtZB8GH64m9EmMZ7FpYw4T6QnrZfgenrhFxaSiSGXtPnz57e9TkNZLvTjeqhr734CNtrK41L40sUQckmj1lGKQ0rC37x544r8eNXRpnVE3ZZY7zXo8NomiO0ZUCj2uHz58rbXoZ6gc0uA+F6ZeKS/jhRDUq8MKrTho9fEkihMmhxtBI1DxKFY9XLpVcSkfoi8JGnToZO5sU5aiDQIW716ddt7ZLYtMQlhECdBGXZZMWldY5BHm5xgAroWj4C0hbYkSc/jBmggIrXJWlZM6pSETsEPGqZOndr2uuuR5rF169a2HoHPdurUKZM4CO1WTPqaDaAd+GFGKdIQkxAn9RuEWcTRyN2KSUgiSgF5aWzPTeA/lN5rZubMmR2bE4SIC4nJoltgAV/dVefZm72AtctUCJU2CMJ327hxY9t7EHbkyJFseq+EJSY16RPo3Dkq1kkr7+q0bNmyDuLQcZBEPYmHVdOBiJyIlrRDq41YPWfXOxUysi5fvtyaj+2BpcnsUV/oSoEMOk2CQGlr4ckhBwaetBhjCwH0ZHtJROPJkyc7UjcYLDjmrH7ADTEBXFfOYmB0k9oYBOjJ8b4aOYSe7QkKcYhFlq3QYLQhSidNmtS2RATwy8YOM3EQJsUjKiaWZ+vZToUQgzhkHXudb/PW5YMHD9yZM2faPsMwoc7RciYJXbGuBqJ1UIGKKLv915jsvgtJxCZDubdXr165mzdvtr1Hz5LONA8jrUwKPqsmVesKa49S3Q4WxmRPUEYdTjgiUcfUwLx589ySJUva3oMkP6IYddq6HMS4o55xBJBUeRjzfa4Zdeg56QZ43LhxoyPo7Lf1kNt7oO8wWAbNwaYjIv5lhyS7kRf96dvm5Jah8vfvX3flyhX35cuX6HfzFHOToS1H4BenCaHvO8pr8iDuwoUL7tevX+b5ZdbBair0xkFIlFDlW4ZknEClsp/TzXyAKVOmmHWFVSbDNw1l1+4f90U6IY/q4V27dpnE9bJ+v87QEydjqx/UamVVPRG+mwkNTYN+9tjkwzEx+atCm/X9WvWtDtAb68Wy9LXa1UmvCDDIpPkyOQ5ZwSzJ4jMrvFcr0rSjOUh+GcT4LSg5ugkW1Io0/SCDQBojh0hPlaJdah+tkVYrnTZowP8iq1F1TgMBBauufyB33x1v+NWFYmT5KmppgHC+NkAgbmRkpD3yn9QIseXymoTQFGQmIOKTxiZIWpvAatenVqRVXf2nTrAWMsPnKrMZHz6bJq5jvce6QK8J1cQNgKxlJapMPdZSR64/UivS9NztpkVEdKcrs5alhhWP9NeqlfWopzhZScI6QxseegZRGeg5a8C3Re1Mfl1ScP36ddcUaMuv24iOJtz7sbUjTS4qBvKmstYJoUauiuD3k5qhyr7QdUHMeCgLa1Ear9NquemdXgmum4fvJ6w1lqsuDhNrg1qSpleJK7K3TF0Q2jSd94uSZ60kK1e3qyVpQK6PVWXp2/FC3mp6jBhKKOiY2h3gtUV64TWM6wDETRPLDfSakXmH3w8g9Jlug8ZtTt4kVF0kLUYYmCCtD/DrQ5YhMGbA9L3ucdjh0y8kOHW5gU/VEEmJTcL4Pz/f7mgoAbYkAAAAAElFTkSuQmCC"]
}'

Risposta ​

json
{
  "model": "llava",
  "created_at": "2023-11-03T15:36:02.583064Z",
  "response": "A happy cartoon character, which is cute and cheerful.",
  "done": true,
  "context": [1, 2, 3],
  "total_duration": 2938432250,
  "load_duration": 2559292,
  "prompt_eval_count": 1,
  "prompt_eval_duration": 2195557000,
  "eval_count": 44,
  "eval_duration": 736432000
}

Richiesta (modalità raw) ​

In alcuni casi, potresti voler bypassare il sistema di templating e fornire un prompt completo. In questo caso, puoi utilizzare il parametro raw per disabilitare il templating. Tieni anche presente che la modalità raw non restituirà un context.

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "mistral",
  "prompt": "[INST] why is the sky blue? [/INST]",
  "raw": true,
  "stream": false
}'

Richiesta (output riproducibili) ​

Per output riproducibili, imposta seed su un numero:

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "mistral",
  "prompt": "Why is the sky blue?",
  "options": {
    "seed": 123
  }
}'
Risposta ​
json
{
  "model": "mistral",
  "created_at": "2023-11-03T15:36:02.583064Z",
  "response": " The sky appears blue because of a phenomenon called Rayleigh scattering.",
  "done": true,
  "total_duration": 8493852375,
  "load_duration": 6589624375,
  "prompt_eval_count": 14,
  "prompt_eval_duration": 119039000,
  "eval_count": 110,
  "eval_duration": 1779061000
}

Richiesta di generazione (con opzioni) ​

Se desideri impostare opzioni personalizzate per il modello in fase di esecuzione invece che nel Modelfile, puoi farlo utilizzando il parametro options. Questo esempio imposta tutte le opzioni disponibili, ma puoi impostarle singolarmente e omettere quelle che non desideri sovrascrivere.

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Why is the sky blue?",
  "stream": false,
  "options": {
    "num_keep": 5,
    "seed": 42,
    "num_predict": 100,
    "draft_num_predict": 4,
    "top_k": 20,
    "top_p": 0.9,
    "min_p": 0.0,
    "typical_p": 0.7,
    "repeat_last_n": 33,
    "temperature": 0.8,
    "repeat_penalty": 1.2,
    "presence_penalty": 1.5,
    "frequency_penalty": 1.0,
    "penalize_newline": true,
    "stop": ["\n", "user:"],
    "numa": false,
    "num_ctx": 1024,
    "num_batch": 2,
    "num_gpu": 1,
    "main_gpu": 0,
    "use_mmap": true,
    "num_thread": 8
  }
}'
Risposta ​
json
{
  "model": "llama3.2",
  "created_at": "2023-08-04T19:22:45.499127Z",
  "response": "The sky is blue because it is the color of the sky.",
  "done": true,
  "context": [1, 2, 3],
  "total_duration": 4935886791,
  "load_duration": 534986708,
  "prompt_eval_count": 26,
  "prompt_eval_duration": 107345000,
  "eval_count": 237,
  "eval_duration": 4289432000
}

Carica un modello ​

Se viene fornito un prompt vuoto, il modello verrà caricato in memoria.

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2"
}'
Risposta ​

Viene restituito un singolo oggetto JSON:

json
{
  "model": "llama3.2",
  "created_at": "2023-12-18T19:52:07.071755Z",
  "response": "",
  "done": true
}

Scarica un modello ​

Se viene fornito un prompt vuoto e il parametro keep_alive è impostato su 0, un modello verrà scaricato dalla memoria.

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2",
  "keep_alive": 0
}'
Risposta ​

Viene restituito un singolo oggetto JSON:

json
{
  "model": "llama3.2",
  "created_at": "2024-09-12T03:54:03.516566Z",
  "response": "",
  "done": true,
  "done_reason": "unload"
}

Genera un completamento di chat ​

POST /api/chat

Genera il messaggio successivo in una chat con un modello fornito. Questo è un endpoint di streaming, quindi verrà restituita una serie di risposte. Lo streaming può essere disabilitato utilizzando "stream": false. L'oggetto di risposta finale includerà statistiche e dati aggiuntivi dalla richiesta.

Parametri ​

  • model: (obbligatorio) il nome del modello
  • messages: i messaggi della chat, può essere utilizzato per mantenere la memoria della conversazione
  • tools: elenco di strumenti in JSON che il modello può utilizzare se supportato
  • think: (per i modelli con capacità di pensiero) il modello deve pensare prima di rispondere? Può essere un valore booleano o un livello di pensiero ("low", "medium", "high" o "max").

L'oggetto message ha i seguenti campi:

  • role: il ruolo del messaggio, ovvero system, user, assistant o tool
  • content: il contenuto del messaggio
  • thinking: (per i modelli con capacità di pensiero) il processo di pensiero del modello
  • images (opzionale): un elenco di immagini da includere nel messaggio (per i modelli multimodali come llava)
  • tool_calls (opzionale): un elenco di strumenti in JSON che il modello desidera utilizzare
  • tool_name (opzionale): aggiungi il nome dello strumento che è stato eseguito per informare il modello del risultato

Parametri avanzati (opzionali):

  • format: il formato in cui restituire la risposta. Il formato può essere json o uno schema JSON.
  • options: parametri aggiuntivi del modello elencati nella documentazione del [Modelfile](. /modelfile. mdx#valid-parameters-and-values) come temperature
  • stream: se false la risposta verrà restituita come un singolo oggetto di risposta, invece che come un flusso di oggetti
  • keep_alive: controlla per quanto tempo il modello rimarrà caricato in memoria dopo la richiesta (predefinito: 5m)

Chiamata di strumenti ​

La chiamata di strumenti è supportata fornendo un elenco di strumenti nel parametro tools. Il modello genererà una risposta che include un elenco di chiamate a strumenti. Vedi l'esempio Richiesta di chat (Streaming con strumenti) di seguito. I modelli possono anche spiegare il risultato della chiamata allo strumento nella risposta. Vedi l'esempio Richiesta di chat (Con cronologia, con strumenti) di seguito. [Vedi i modelli con capacità di chiamata di strumenti](https://ollama. com/search? c=tool).

Output strutturati ​

Gli output strutturati sono supportati fornendo uno schema JSON nel parametro format. Il modello genererà una risposta che corrisponde allo schema. Vedi l'esempio Richiesta di chat (Output strutturati) di seguito.

Esempi ​

Richiesta di chat (Streaming) ​

Richiesta ​

Invia un messaggio di chat con una risposta in streaming.

shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [
    {
      "role": "user",
      "content": "why is the sky blue? "
    }
  ]
}'
Risposta ​

Viene restituito un flusso di oggetti JSON:

json
{
  "model": "llama3. 2",
  "created_at": "2023-08-04T08:52:19. 385406455-07:00",
  "message": {
    "role": "assistant",
    "content": "The",
    "images": null
  },
  "done": false
}

Risposta finale:

json
{
  "model": "llama3. 2",
  "created_at": "2023-08-04T19:22:45. 499127Z",
  "message": {
    "role": "assistant",
    "content": ""
  },
  "done": true,
  "total_duration": 4883583458,
  "load_duration": 1334875,
  "prompt_eval_count": 26,
  "prompt_eval_duration": 342546000,
  "eval_count": 282,
  "eval_duration": 4535599000
}

Richiesta di chat (Streaming con strumenti) ​

Richiesta ​
shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [
    {
      "role": "user",
      "content": "what is the weather in tokyo? "
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the weather in a given city",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string",
              "description": "The city to get the weather for"
            }
          },
          "required": ["city"]
        }
      }
    }
  ],
  "stream": true
}'
Risposta ​

Viene restituito un flusso di oggetti JSON:

json
{
  "model": "llama3. 2",
  "created_at": "2025-07-07T20:22:19. 184789Z",
  "message": {
    "role": "assistant",
    "content": "",
    "tool_calls": [
      {
        "function": {
          "name": "get_weather",
          "arguments": {
            "city": "Tokyo"
          }
        }
      }
    ]
  },
  "done": false
}

Risposta finale:

json
{
  "model": "llama3. 2",
  "created_at": "2025-07-07T20:22:19. 19314Z",
  "message": {
    "role": "assistant",
    "content": ""
  },
  "done_reason": "stop",
  "done": true,
  "total_duration": 182242375,
  "load_duration": 41295167,
  "prompt_eval_count": 169,
  "prompt_eval_duration": 24573166,
  "eval_count": 15,
  "eval_duration": 115959084
}

Richiesta di chat (Nessuno streaming) ​

Richiesta ​
shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [
    {
      "role": "user",
      "content": "why is the sky blue? "
    }
  ],
  "stream": false
}'
Risposta ​
json
{
  "model": "llama3. 2",
  "created_at": "2023-12-12T14:13:43. 416799Z",
  "message": {
    "role": "assistant",
    "content": "Hello! How are you today? "
  },
  "done": true,
  "total_duration": 5191566416,
  "load_duration": 2154458,
  "prompt_eval_count": 26,
  "prompt_eval_duration": 383809000,
  "eval_count": 298,
  "eval_duration": 4799921000
}

Richiesta di chat (Nessuno streaming, con strumenti) ​

Richiesta ​
shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [
    {
      "role": "user",
      "content": "what is the weather in tokyo? "
    }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the weather in a given city",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string",
              "description": "The city to get the weather for"
            }
          },
          "required": ["city"]
        }
      }
    }
  ],
  "stream": false
}'
Risposta ​
json
{
  "model": "llama3. 2",
  "created_at": "2025-07-07T20:32:53. 844124Z",
  "message": {
    "role": "assistant",
    "content": "",
    "tool_calls": [
      {
        "function": {
          "name": "get_weather",
          "arguments": {
            "city": "Tokyo"
          }
        }
      }
    ]
  },
  "done_reason": "stop",
  "done": true,
  "total_duration": 3244883583,
  "load_duration": 2969184542,
  "prompt_eval_count": 169,
  "prompt_eval_duration": 141656333,
  "eval_count": 18,
  "eval_duration": 133293625
}

Richiesta di chat (Output strutturati) ​

Richiesta ​
shell
curl -X POST http://localhost:11434/api/chat -H "Content-Type: application/json" -d '{
  "model": "llama3. 1",
  "messages": [{"role": "user", "content": "Ollama is 22 years old and busy saving the world. Return a JSON object with the age and availability. "}],
  "stream": false,
  "format": {
    "type": "object",
    "properties": {
      "age": {
        "type": "integer"
      },
      "available": {
        "type": "boolean"
      }
    },
    "required": [
      "age",
      "available"
    ]
  },
  "options": {
    "temperature": 0
  }
}'
Risposta ​
json
{
  "model": "llama3. 1",
  "created_at": "2024-12-06T00:46:58. 265747Z",
  "message": {
    "role": "assistant",
    "content": "{\"age\": 22, \"available\": false}"
  },
  "done_reason": "stop",
  "done": true,
  "total_duration": 2254970291,
  "load_duration": 574751416,
  "prompt_eval_count": 34,
  "prompt_eval_duration": 1502000000,
  "eval_count": 12,
  "eval_duration": 175000000
}

Richiesta di chat (Con cronologia) ​

Invia un messaggio di chat con una cronologia della conversazione. Puoi utilizzare lo stesso approccio per avviare la conversazione utilizzando prompting multi-shot o chain-of-thought.

Richiesta ​
shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [
    {
      "role": "user",
      "content": "why is the sky blue? "
    },
    {
      "role": "assistant",
      "content": "due to rayleigh scattering. "
    },
    {
      "role": "user",
      "content": "how is that different than mie scattering? "
    }
  ]
}'
Risposta ​

Viene restituito un flusso di oggetti JSON:

json
{
  "model": "llama3. 2",
  "created_at": "2023-08-04T08:52:19. 385406455-07:00",
  "message": {
    "role": "assistant",
    "content": "The"
  },
  "done": false
}

Risposta finale:

json
{
  "model": "llama3. 2",
  "created_at": "2023-08-04T19:22:45. 499127Z",
  "done": true,
  "total_duration": 8113331500,
  "load_duration": 6396458,
  "prompt_eval_count": 61,
  "prompt_eval_duration": 398801000,
  "eval_count": 468,
  "eval_duration": 7701267000
}

Richiesta di chat (Con cronologia, con strumenti) ​

Richiesta ​
shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [
    {
      "role": "user",
      "content": "what is the weather in Toronto? "
    },
    // il messaggio del modello aggiunto alla cronologia
    {
      "role": "assistant",
      "content": "",
      "tool_calls": [
        {
          "function": {
            "name": "get_weather",
            "arguments": {
              "city": "Toronto"
            }
          }
        }
      ]
    },
    // il risultato della chiamata allo strumento aggiunto alla cronologia
    {
      "role": "tool",
      "content": "11 degrees celsius",
      "tool_name": "get_weather"
    }
  ],
  "stream": false,
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get the weather in a given city",
        "parameters": {
          "type": "object",
          "properties": {
            "city": {
              "type": "string",
              "description": "The city to get the weather for"
            }
          },
          "required": ["city"]
        }
      }
    }
  ]
}'
Risposta ​
json
{
  "model": "llama3. 2",
  "created_at": "2025-07-07T20:43:37. 688511Z",
  "message": {
    "role": "assistant",
    "content": "The current temperature in Toronto is 11°C. "
  },
  "done_reason": "stop",
  "done": true,
  "total_duration": 890771750,
  "load_duration": 707634750,
  "prompt_eval_count": 94,
  "prompt_eval_duration": 91703208,
  "eval_count": 11,
  "eval_duration": 90282125
}

Richiesta di chat (con immagini) ​

Richiesta ​

Invia un messaggio di chat con immagini. Le immagini devono essere fornite come array, con le singole immagini codificate in Base64.

shell
curl http://localhost:11434/api/chat -d '{
  "model": "llava",
  "messages": [
    {
      "role": "user",
      "content": "what is in this image?"

",
      "images": ["iVBORw0KGgoAAAANSUhEUgAAAG0AAABmCAYAAADBPx+VAAAACXBIWXMAAAsTAAALEwEAmpwYAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAA3VSURBVHgB7Z27r0zdG8fX743i1bi1ikMoFMQloXRpKFFIqI7LH4BEQ+NWIkjQuSWCRIEoULk0gsK1kCBI0IhrQVT7tz/7zZo888yz1r7MnDl7z5xvsjkzs2fP3uu71nNfa7lkAsm7d++Sffv2JbNmzUqcc8m0adOSzZs3Z+/XES4ZckAWJEGWPiCxjsQNLWmQsWjRIpMseaxcuTKpG/7HP27I8P79e7dq1ars/yL4/v27S0ejqwv+cUOGEGGpKHR37tzJCEpHV9tnT58+dXXCJDdECBE2Ojrqjh071hpNECjx4cMHVycM1Uhbv359B2F79+51506daxN/+pyRkRFXKyRDAqxEp4yMlDDzXG1NPnnyJKkThoK0VFd1ELZu3TrzXKxKfW7dMBQ6bcuWLW2v0VlHjx41z717927ba22U9APcw7Nnz1oGEPeL3m3p2mTAYYnFmMOMXybPPXv2bNIPpFZr1NHn4HMw0KRBjg9NuRw95s8PEcz/6DZELQd/09C9QGq5RsmSRybqkwHGjh07OsJSsYYm3ijPpyHzoiacg35MLdDSIS/O1yM778jOTwYUkKNHWUzUWaOsylE00MyI0fcnOwIdjvtNdW/HZwNLGg+sR1kMepSNJXmIwxBZiG8tDTpEZzKg0GItNsosY8USkxDhD0Rinuiko2gfL/RbiD2LZAjU9zKQJj8RDR0vJBR1/Phx9+PHj9Z7REF4nTZkxzX4LCXHrV271qXkBAPGfP/atWvu/PnzHe4C97F48eIsRLZ9+3a3f/9+87dwP1JxaF7/3r17ba+5l4EcaVo0lj3SBq5kGTJSQmLWMjgYNei2GPT1MuMqGTDEFHzeQSP2wi/jGnkmPJ/nhccs44jvDAxpVcxnq0F6eT8h4ni/iIWpR5lPyA6ETkNXoSukvpJAD3AsXLiwpZs49+fPn5ke4j10TqYvegSfn0OnafC+Tv9ooA/JPkgQysqQNBzagXY55nO/oa1F7qvIPWkRL12WRpMWUvpVDYmxAPehxWSe8ZEXL20sadYIozfmNch4QJPAfeJgW3rNsnzphBKNJM2KKODo1rVOMRYik5ETy3ix4qWNI81qAAirizgMIc+yhTytx0JWZuNI03qsrgWlGtwjoS9XwgUhWGyhUaRZZQNNIEwCiXD16tXcAHUs79co0vSD8rrJCIW98pzvxpAWyyo3HYwqS0+H0BjStClcZJT5coMm6D2LOF8TolGJtK9fvyZpyiC5ePFi9nc/oJU4eiEP0jVoAnHa9wyJycITMP78+eMeP37sXrx44d6+fdt6f82aNdkx1pg9e3Zb5W+RSRE+n+VjksQWifvVaTKFhn5O8my63K8Qabdv33b379/PiAP//vuvW7BggZszZ072/+TJk91YgkafPn166zXB1rQHFvouAWHq9z3SEevSUerqCn2/dDCeta2jxYbr69evk4MHDyY7d+7MjhMnTiTPnz9Pfv/+nfQT2ggpO2dMF8cghuoM7Ygj5iWCqRlGFml0QC/ftGmTmzt3rmsaKDsgBSPh0/8yPeLLBihLkOKJc0jp8H8vUzcxIA1k6QJ/c78tWEyj5P3o4u9+jywNPdJi5rAH9x0KHcl4Hg570eQp3+vHXGyrmEeigzQsQsjavXt38ujRo44LQuDDhw+TW7duRS1HGgMxhNXHgflaNTOsHyKvHK5Ijo2jbFjJBQK9YwFd6RVMzfgRBmEfP37suBBm/p49e1qjEP2mwTViNRo0VJWH1deMXcNK08uUjVUu7s/zRaL+oLNxz1bpANco4npUgX4G2eFbpDFyQoQxojBCpEGSytmOH8qrH5Q9vuzD6ofQylkCUmh8DBAr+q8JCyVNtWQIidKQE9wNtLSQnS4jDSsxNHogzFuQBw4cyM61UKVsjfr3ooBkPSqqQHesUPWVtzi9/vQi1T+rJj7WiTz4Pt/l3LxUkr5P2VYZaZ4URpsE+st/dujQoaBBYokbrz/8TJNQYLSonrPS9kUaSkPeZyj1AWSj+d+VBoy1pIWVNed8P0Ll/ee5HdGRhrHhR5GGN0r4LGZBaj8oFDJitBTJzIZgFcmU0Y8ytWMZMzJOaXUSrUs5RxKnrxmbb5YXO9VGUhtpXldhEUogFr3IzIsvlpmdosVcGVGXFWp2oU9kLFL3dEkSz6NHEY1sjSRdIuDFWEhd8KxFqsRi1uM/nz9/zpxnwlESONdg6dKlbsaMGS4EHFHtjFIDHwKOo46l4TxSuxgDzi+rE2jg+BaFruOX4HXa0Nnf1lwAPufZeF8/r6zD97WK2qFnGjBxTw5qNGPxT+5T/r7/7RawFC3j4vTp09koCxkeHjqbHJqArmH5UrFKKksnxrK7FuRIs8STfBZv+luugXZ2pR/pP9Ois4z+TiMzUUkUjD0iEi1fzX8GmXyuxUBRcaUfykV0YZnlJGKQpOiGB76x5GeWkWWJc3mOrK6S7xdND+W5N6XyaRgtWJFe13GkaZnKOsYqGdOVVVbGupsyA/l7emTLHi7vwTdirNEt0qxnzAvBFcnQF16xh/TMpUuXHDowhlA9vQVraQhkudRdzOnK+04ZSP3DUhVSP61YsaLtd/ks7ZgtPcXqPqEafHkdqa84X6aCeL7YWlv6edGFHb+ZFICPlljHhg0bKuk0CSvVznWsotRu433alNdFrqG45ejoaPCaUkWERpLXjzFL2Rpllp7PJU2a/v7Ab8N05/9t27Z16KUqoFGsxnI9EosS2niSYg9SpU6B4JgTrvVW1flt1sT+0ADIJU2maXzcUTraGCRaL1Wp9rUMk16PMom8QhruxzvZIegJjFU7LLCePfS8uaQdPny4jTTL0dbee5mYokQsXTIWNY46kuMbnt8Kmec+LGWtOVIl9cT1rCB0V8WqkjAsRwta93TbwNYoGKsUSChN44lgBNCoHLHzquYKrU6qZ8lolCIN0Rh6cP0Q3U6I6IXILYOQI513hJaSKAorFpuHXJNfVlpRtmYBk1Su1obZr5dnKAO+L10Hrj3WZW+E3qh6IszE37F6EB+68mGpvKm4eb9bFrlzrok7fvr0Kfv727dvWRmdVTJHw0qiiCUSZ6wCK+7XL/AcsgNyL74DQQ730sv78Su7+t/A36MdY0sW5o40ahslXr58aZ5HtZB8GH64m9EmMZ7FpYw4T6QnrZfgenrhFxaSiSGXtPnz57e9TkNZLvTjeqhr734CNtrK41L40sUQckmj1lGKQ0rC37x544r8eNXRpnVE3ZZY7zXo8NomiO0ZUCj2uHz58rbXoZ6gc0uA+F6ZeKS/jhRDUq8MKrTho9fEkihMmhxtBI1DxKFY9XLpVcSkfoi8JGnToZO5sU5aiDQIW716ddt7ZLYtMQlhECdBGXZZMWldY5BHm5xgAroWj4C0hbYkSc/jBmggIrXJWlZM6pSETsEPGqZOndr2uuuR5rF169a2HoHPdurUKZM4CO1WTPqaDaAd+GFGKdIQkxAn9RuEWcTRyN2KSUgiSgF5aWzPTeA/lN5rZubMmR2bE4SIC4nJoltgAV/dVefZm72AtctUCJU2CMJ327hxY9t7EHbkyJFseq+EJSY16RPo3Dkq1kkr7+q0bNmyDuLQcZBEPYmHVdOBiJyIlrRDq41YPWfXOxUysi5fvtyaj+2BpcnsUV/oSoEMOk2CQGlr4ckhBwaetBhjCwH0ZHtJROPJkyc7UjcYLDjmrH7ADTEBXFfOYmB0k9oYBOjJ8b4aOYSe7QkKcYhFlq3QYLQhSidNmtS2RATwy8YOM3EQJsUjKiaWZ+vZToUQgzhkHXudb/PW5YMHD9yZM2faPsMwoc7RciYJXbGuBqJ1UIGKKLv915jsvgtJxCZDubdXr165mzdvtr1Hz5LONA8jrUwKPqsmVesKa49S3Q4WxmRPUEYdTjgiUcfUwLx589ySJUva3oMkP6IYddq6HMS4o55xBJBUeRjzfa4Zdeg56QZ43LhxoyPo7Lf1kNt7oO8wWAbNwaYjIv5lhyS7kRf96dvm5Jah8vfvX3flyhX35cuX6HfzFHOToS1H4BenCaHvO8pr8iDuwoUL7tevX+b5ZdbBair0xkFIlFDlW4ZknEClsp/TzXyAKVOmmHWFVSbDNw1l1+4f90U6IY/q4V27dpnE9bJ+v87QEydjqx/UamVVPRG+mwkNTYN+9tjkwzEx+atCm/X9WvWtDtAb68Wy9LXa1UmvCDDIpPkyOQ5ZwSzJ4jMrvFcr0rSjOUh+GcT4LSg5ugkW1Io0/SCDQBojh0hPlaJdah+tkVYrnTZowP8iq1F1TgMBBauufyB33x1v+NWFYmT5KmppgHC+NkAgbmRkpD3yn9QIseXymoTQFGQmIOKTxiZIWpvAatenVqRVXf2nTrAWMsPnKrMZHz6bJq5jvce6QK8J1cQNgKxlJapMPdZSR64/UivS9NztpkVEdKcrs5alhhWP9NeqlfWopzhZScI6QxseegZRGeg5a8C3Re1Mfl1ScP36ddcUaMuv24iOJtz7sbUjTS4qBvKmstYJoUauiuD3k5qhyr7QdUHMeCgLa1Ear9NquemdXgmum4fvJ6w1lqsuDhNrg1qSpleJK7K3TF0Q2jSd94uSZ60kK1e3qyVpQK6PVWXp2/FC3mp6jBhKKOiY2h3gtUV64TWM6wDETRPLDfSakXmH3w8g9Jlug8ZtTt4kVF0kLUYYmCCtD/DrQ5YhMGbA9L3ucdjh0y8kOHW5gU/VEEmJTcL4Pz/f7mgoAbYkAAAAAElFTkSuQmCC"]
    }
  ]
}'
Risposta ​
json
{
  "model": "llava",
  "created_at": "2023-12-13T22:42:50. 203334Z",
  "message": {
    "role": "assistant",
    "content": " The image features a cute, little pig with an angry facial expression. It's wearing a heart on its shirt and is waving in the air. This scene appears to be part of a drawing or sketching project. ",
    "images": null
  },
  "done": true,
  "total_duration": 1668506709,
  "load_duration": 1986209,
  "prompt_eval_count": 26,
  "prompt_eval_duration": 359682000,
  "eval_count": 83,
  "eval_duration": 1303285000
}

Richiesta di chat (output riproducibili) ​

Richiesta ​
shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [
    {
      "role": "user",
      "content": "Hello! "
    }
  ],
  "options": {
    "seed": 101,
    "temperature": 0
  }
}'
Risposta ​
json
{
  "model": "llama3. 2",
  "created_at": "2023-12-12T14:13:43. 416799Z",
  "message": {
    "role": "assistant",
    "content": "Hello! How are you today? "
  },
  "done": true,
  "total_duration": 5191566416,
  "load_duration": 2154458,
  "prompt_eval_count": 26,
  "prompt_eval_duration": 383809000,
  "eval_count": 298,
  "eval_duration": 4799921000
}

Richiesta di chat (con strumenti) ​

Richiesta ​
shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [
    {
      "role": "user",
      "content": "What is the weather today in Paris? "
    }
  ],
  "stream": false,
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_current_weather",
        "description": "Get the current weather for a location",
        "parameters": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "The location to get the weather for, e. g. San Francisco, CA"
            },
            "format": {
              "type": "string",
              "description": "The format to return the weather in, e. g. 'celsius' or 'fahrenheit'",
              "enum": ["celsius", "fahrenheit"]
            }
          },
          "required": ["location", "format"]
        }
      }
    }
  ]
}'
Risposta ​
json
{
  "model": "llama3. 2",
  "created_at": "2024-07-22T20:33:28. 123648Z",
  "message": {
    "role": "assistant",
    "content": "",
    "tool_calls": [
      {
        "function": {
          "name": "get_current_weather",
          "arguments": {
            "format": "celsius",
            "location": "Paris, FR"
          }
        }
      }
    ]
  },
  "done_reason": "stop",
  "done": true,
  "total_duration": 885095291,
  "load_duration": 3753500,
  "prompt_eval_count": 122,
  "prompt_eval_duration": 328493000,
  "eval_count": 33,
  "eval_duration": 552222000
}

Carica un modello ​

Se l'array dei messaggi è vuoto, il modello verrà caricato in memoria. ##### Richiesta

shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": []
}'
Risposta ​
json
{
  "model": "llama3. 2",
  "created_at": "2024-09-12T21:17:29. 110811Z",
  "message": {
    "role": "assistant",
    "content": ""
  },
  "done_reason": "load",
  "done": true
}

Scarica un modello ​

Se l'array dei messaggi è vuoto e il parametro keep_alive è impostato su 0, un modello verrà scaricato dalla memoria. ##### Richiesta

shell
curl http://localhost:11434/api/chat -d '{
  "model": "llama3. 2",
  "messages": [],
  "keep_alive": 0
}'
Risposta ​

Viene restituito un singolo oggetto JSON:

json
{
  "model": "llama3. 2",
  "created_at": "2024-09-12T21:33:17. 547535Z",
  "message": {
    "role": "assistant",
    "content": ""
  },
  "done_reason": "unload",
  "done": true
}

Creare un modello ​

POST /api/create

Creare un modello da:

  • un altro modello;
  • una directory safetensors; o
  • un file GGUF.

Se si sta creando un modello da una directory safetensors o da un file GGUF, è necessario creare un blob per ciascuno dei file e poi utilizzare il nome del file e l'impronta SHA256 associata a ogni blob nel campo files.

Parametri ​

  • model: nome del modello da creare
  • from: (facoltativo) nome di un modello esistente da cui creare il nuovo modello
  • files: (facoltativo) un dizionario di nomi di file e impronte SHA256 di blob da cui creare il modello
  • adapters: (facoltativo) un dizionario di nomi di file e impronte SHA256 di blob per adattatori LORA
  • template: (facoltativo) il template di prompt per il modello
  • renderer: (facoltativo) il nome del renderer per il modello
  • parser: (facoltativo) il nome del parser per il modello
  • license: (facoltativo) una stringa o un elenco di stringhe contenente la licenza o le licenze per il modello
  • system: (facoltativo) una stringa contenente il prompt di sistema per il modello
  • parameters: (facoltativo) un dizionario di parametri per il modello (vedi Modelfile per un elenco di parametri)
  • messages: (facoltativo) un elenco di oggetti messaggio utilizzati per creare una conversazione
  • stream: (facoltativo) se false la risposta verrà restituita come un singolo oggetto di risposta, invece che come flusso di oggetti
  • quantize (facoltativo): quantizza un modello non quantizzato (ad es. float16)

Tipi di quantizzazione ​

TipoConsigliato
q4_K_M*
q4_K_S
q8_0*

Esempi ​

Creare un nuovo modello ​

Creare un nuovo modello a partire da un modello esistente.

Richiesta ​
shell
curl http://localhost:11434/api/create -d '{
  "model": "mario",
  "from": "llama3.2",
  "system": "You are Mario from Super Mario Bros."
}'
Risposta ​

Viene restituito un flusso di oggetti JSON:

json
{"status":"reading model metadata"}
{"status":"creating system layer"}
{"status":"using already created layer sha256:22f7f8ef5f4c791c1b03d7eb414399294764d7cc82c7e94aa81a1feb80a983a2"}
{"status":"using already created layer sha256:8c17c2ebb0ea011be9981cc3922db8ca8fa61e828c5d3f44cb6ae342bf80460b"}
{"status":"using already created layer sha256:7c23fb36d80141c4ab8cdbb61ee4790102ebd2bf7aeff414453177d4f2110e5d"}
{"status":"using already created layer sha256:2e0493f67d0c8c9c68a8aeacdf6a38a2151cb3c4c1d42accf296e19810527988"}
{"status":"using already created layer sha256:2759286baa875dc22de5394b4a925701b1896a7e3f8e53275c36f75a877a82c9"}
{"status":"writing layer sha256:df30045fe90f0d750db82a058109cecd6d4de9c90a3d75b19c09e5f64580bb42"}
{"status":"writing layer sha256:f18a68eb09bf925bb1b669490407c1b1251c5db98dc4d3d81f3088498ea55690"}
{"status":"writing manifest"}
{"status":"success"}

Quantizzare un modello ​

Quantizza un modello non quantizzato.

Richiesta ​
shell
curl http://localhost:11434/api/create -d '{
  "model": "llama3.2:quantized",
  "from": "llama3.2:3b-instruct-fp16",
  "quantize": "q4_K_M"
}'
Risposta ​

Viene restituito un flusso di oggetti JSON:

json
{"status":"quantizing F16 model to Q4_K_M","digest":"0","total":6433687776,"completed":12302}
{"status":"quantizing F16 model to Q4_K_M","digest":"0","total":6433687776,"completed":6433687552}
{"status":"verifying conversion"}
{"status":"creating new layer sha256:fb7f4f211b89c6c4928ff4ddb73db9f9c0cfca3e000c3e40d6cf27ddc6ca72eb"}
{"status":"using existing layer sha256:966de95ca8a62200913e3f8bfbf84c8494536f1b94b49166851e76644e966396"}
{"status":"using existing layer sha256:fcc5a6bec9daf9b561a68827b67ab6088e1dba9d1fa2a50d7bbcc8384e0a265d"}
{"status":"using existing layer sha256:a70ff7e570d97baaf4e62ac6e6ad9975e04caa6d900d3742d37698494479e0cd"}
{"status":"using existing layer sha256:56bb8bd477a519ffa694fc449c2413c6f0e1d3b1c88fa7e3c9d88d3ae49d4dcb"}
{"status":"writing manifest"}
{"status":"success"}

Creare un modello da GGUF ​

Creare un modello da un file GGUF. Il parametro files deve essere compilato con il nome del file e l'impronta SHA256 del file GGUF che si desidera utilizzare. Utilizzare /api/blobs/:digest per caricare il file GGUF sul server prima di chiamare questa API.

Richiesta ​
shell
curl http://localhost:11434/api/create -d '{
  "model": "my-gguf-model",
  "files": {
    "test.gguf": "sha256:432f310a77f4650a88d0fd59ecdd7cebed8d684bafea53cbff0473542964f0c3"
  }
}'
Risposta ​

Viene restituito un flusso di oggetti JSON:

json
{"status":"parsing GGUF"}
{"status":"using existing layer sha256:432f310a77f4650a88d0fd59ecdd7cebed8d684bafea53cbff0473542964f0c3"}
{"status":"writing manifest"}
{"status":"success"}

Creare un modello da una directory Safetensors ​

Il parametro files deve includere un dizionario di file per il modello safetensors che contiene i nomi dei file e l'impronta SHA256 di ciascun file. Utilizzare /api/blobs/:digest per caricare prima ciascuno dei file sul server prima di chiamare questa API. I file rimarranno nella cache fino al riavvio del server Ollama.

Richiesta ​
shell
curl http://localhost:11434/api/create -d '{
  "model": "fred",
  "files": {
    "config.json": "sha256:dd3443e529fb2290423a0c65c2d633e67b419d273f170259e27297219828e389",
    "generation_config.json": "sha256:88effbb63300dbbc7390143fbbdd9d9fa50587b37e8bfd16c8c90d4970a74a36",
    "special_tokens_map.json": "sha256:b7455f0e8f00539108837bfa586c4fbf424e31f8717819a6798be74bef813d05",
    "tokenizer.json": "sha256:bbc1904d35169c542dffbe1f7589a5994ec7426d9e5b609d07bab876f32e97ab",
    "tokenizer_config.json": "sha256:24e8a6dc2547164b7002e3125f10b415105644fcf02bf9ad8b674c87b1eaaed6",
    "model.safetensors": "sha256:1ff795ff6a07e6a68085d206fb84417da2f083f68391c2843cd2b8ac6df8538f"
  }
}'
Risposta ​

Viene restituito un flusso di oggetti JSON:

shell
{"status":"converting model"}
{"status":"creating new layer sha256:05ca5b813af4a53d2c2922933936e398958855c44ee534858fcfd830940618b6"}
{"status":"using autodetected template llama3-instruct"}
{"status":"using existing layer sha256:56bb8bd477a519ffa694fc449c2413c6f0e1d3b1c88fa7e3c9d88d3ae49d4dcb"}
{"status":"writing manifest"}
{"status":"success"}

Verifica se un Blob esiste ​

shell
HEAD /api/blobs/:digest

Verifica che il blob di file (Binary Large Object) utilizzato per creare un modello esista sul server. Questa operazione controlla il tuo server Ollama e non ollama.com.

Parametri di query ​

  • digest: il digest SHA256 del blob

Esempi ​

Richiesta ​

shell
curl -I http://localhost:11434/api/blobs/sha256:29fdb92e57cf0827ded04ae6461b5931d01fa595843f55d36f5b275a52087dd2

Risposta ​

Restituisce 200 OK se il blob esiste, 404 Not Found se non esiste.

Carica un Blob ​

POST /api/blobs/:digest

Carica un file sul server Ollama per creare un "blob" (Binary Large Object).

Parametri di query ​

  • digest: il digest SHA256 previsto del file

Esempi ​

Richiesta ​

shell
curl -T model.gguf -X POST http://localhost:11434/api/blobs/sha256:29fdb92e57cf0827ded04ae6461b5931d01fa595843f55d36f5b275a52087dd2

Risposta ​

Restituisce 201 Created se il blob è stato creato correttamente, 400 Bad Request se il digest utilizzato non è quello previsto.

Elenca modelli locali ​

GET /api/tags

Elenca i modelli disponibili localmente.

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/tags

Risposta ​

Verrà restituito un singolo oggetto JSON.

json
{
  "models": [
    {
      "name": "deepseek-r1:latest",
      "model": "deepseek-r1:latest",
      "modified_at": "2025-05-10T08:06:48.639712648-07:00",
      "size": 4683075271,
      "digest": "0a8c266910232fd3291e71e5ba1e058cc5af9d411192cf88b6d30e92b6e73163",
      "details": {
        "parent_model": "",
        "format": "gguf",
        "family": "qwen2",
        "families": ["qwen2"],
        "parameter_size": "7.6B",
        "quantization_level": "Q4_K_M"
      }
    },
    {
      "name": "llama3.2:latest",
      "model": "llama3.2:latest",
      "modified_at": "2025-05-04T17:37:44.706015396-07:00",
      "size": 2019393189,
      "digest": "a80c4f17acd55265feec403c7aef86be0c25983ab279d83f3bcd3abbcb5b8b72",
      "details": {
        "parent_model": "",
        "format": "gguf",
        "family": "llama",
        "families": ["llama"],
        "parameter_size": "3.2B",
        "quantization_level": "Q4_K_M"
      }
    }
  ]
}

Mostra informazioni sul modello ​

POST /api/show

Mostra informazioni su un modello, inclusi dettagli, modelfile, template, parametri, licenza, prompt di sistema.

Parametri ​

  • model: nome del modello da visualizzare
  • verbose: (opzionale) se impostato su true, restituisce i dati completi per i campi di risposta dettagliati

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/show -d '{
  "model": "llava"
}'

Risposta ​

json5
{
  modelfile: '# Modelfile generated by "ollama show"\n# To build a new Modelfile based on this one, replace the FROM line with:\n# FROM llava:latest\n\nFROM /Users/matt/.ollama/models/blobs/sha256:200765e1283640ffbd013184bf496e261032fa75b99498a9613be4e94d63ad52\nTEMPLATE """{{ .System }}\nUSER: {{ .Prompt }}\nASSISTANT: """\nPARAMETER num_ctx 4096\nPARAMETER stop "\u003c/s\u003e"\nPARAMETER stop "USER:"\nPARAMETER stop "ASSISTANT:"',
  parameters: 'num_keep                       24\nstop                           "<|start_header_id|>"\nstop                           "<|end_header_id|>"\nstop                           "<|eot_id|>"',
  template: "{{ if .System }}<|start_header_id|>system<|end_header_id|>\n\n{{ .System }}<|eot_id|>{{ end }}{{ if .Prompt }}<|start_header_id|>user<|end_header_id|>\n\n{{ .Prompt }}<|eot_id|>{{ end }}<|start_header_id|>assistant<|end_header_id|>\n\n{{ .Response }}<|eot_id|>",
  details: {
    parent_model: "",
    format: "gguf",
    family: "llama",
    families: ["llama"],
    parameter_size: "8.0B",
    quantization_level: "Q4_0",
  },
  model_info: {
    "general.architecture": "llama",
    "general.file_type": 2,
    "general.parameter_count": 8030261248,
    "general.quantization_version": 2,
    "llama.attention.head_count": 32,
    "llama.attention.head_count_kv": 8,
    "llama.attention.layer_norm_rms_epsilon": 0.00001,
    "llama.block_count": 32,
    "llama.context_length": 8192,
    "llama.embedding_length": 4096,
    "llama.feed_forward_length": 14336,
    "llama.rope.dimension_count": 128,
    "llama.rope.freq_base": 500000,
    "llama.vocab_size": 128256,
    "tokenizer.ggml.bos_token_id": 128000,
    "tokenizer.ggml.eos_token_id": 128009,
    "tokenizer.ggml.merges": [], // populates if `verbose=true`
    "tokenizer.ggml.model": "gpt2",
    "tokenizer.ggml.pre": "llama-bpe",
    "tokenizer.ggml.token_type": [], // populates if `verbose=true`
    "tokenizer.ggml.tokens": [], // populates if `verbose=true`
  },
  capabilities: ["completion", "vision"],
}

Copia un modello ​

POST /api/copy

Copia un modello. Crea un modello con un nome diverso a partire da un modello esistente.

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/copy -d '{
  "source": "llama3.2",
  "destination": "llama3-backup"
}'

Risposta ​

Restituisce 200 OK se l'operazione ha esito positivo, o 404 Not Found se il modello di origine non esiste.

Elimina un modello ​

DELETE /api/delete

Elimina un modello e i suoi dati.

Parametri ​

  • model: nome del modello da eliminare

Esempi ​

Richiesta ​

shell
curl -X DELETE http://localhost:11434/api/delete -d '{
  "model": "llama3:13b"
}'

Risposta ​

Restituisce 200 OK se l'operazione ha esito positivo, 404 Not Found se il modello da eliminare non esiste.

Scarica un modello ​

POST /api/pull

Scarica un modello dalla libreria ollama. I download annullati vengono ripresi dal punto in cui erano stati interrotti e le chiamate multiple condivideranno lo stesso progresso di download.

Parametri ​

  • model: nome del modello da scaricare
  • insecure: (opzionale) consente connessioni non sicure alla libreria. Utilizzare solo se si sta scaricando dalla propria libreria durante lo sviluppo.
  • stream: (opzionale) se impostato su false, la risposta verrà restituita come un singolo oggetto di risposta, invece di un flusso di oggetti

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/pull -d '{
  "model": "llama3.2"
}'

Risposta ​

Se stream non è specificato o impostato su true, viene restituito un flusso di oggetti JSON:

Il primo oggetto è il manifesto:

json
{
  "status": "pulling manifest"
}

Successivamente viene restituita una serie di risposte di download. Finché il download non è completato, la chiave completed potrebbe non essere inclusa. Il numero di file da scaricare dipende dal numero di layer specificati nel manifesto.

json
{
  "status": "pulling digestname",
  "digest": "digestname",
  "total": 2142590208,
  "completed": 241970
}

Dopo che tutti i file sono stati scaricati, le risposte finali sono:

json
{
    "status": "verifying sha256 digest"
}
{
    "status": "writing manifest"
}
{
    "status": "removing any unused layers"
}
{
    "status": "success"
}

se stream è impostato su false, la risposta è un singolo oggetto JSON:

json
{
  "status": "success"
}

Carica un modello ​

POST /api/push

Carica un modello in una libreria di modelli. È necessario registrarsi su ollama.ai e aggiungere prima una chiave pubblica.

Parametri ​

  • model: nome del modello da caricare nella forma <namespace>/<model>:<tag>
  • insecure: (opzionale) consente connessioni non sicure alla libreria. Utilizzare solo se si sta caricando sulla propria libreria durante lo sviluppo.
  • stream: (opzionale) se impostato su false, la risposta verrà restituita come un singolo oggetto di risposta, invece di un flusso di oggetti

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/push -d '{
  "model": "mattw/pygmalion:latest"
}'

Risposta ​

Se stream non è specificato o impostato su true, viene restituito un flusso di oggetti JSON:

json
{ "status": "retrieving manifest" }

e poi:

json
{
  "status": "starting upload",
  "digest": "sha256:bc07c81de745696fdf5afca05e065818a8149fb0c77266fb584d9b2cba3711ab",
  "total": 1928429856
}

Successivamente viene restituita una serie di risposte di caricamento:

json
{
  "status": "starting upload",
  "digest": "sha256:bc07c81de745696fdf5afca05e065818a8149fb0c77266fb584d9b2cba3711ab",
  "total": 1928429856
}

Infine, quando il caricamento è completato:

json
{"status":"pushing manifest"}
{"status":"success"}

Se stream è impostato su false, la risposta è un singolo oggetto JSON:

json
{ "status": "success" }

Genera Embedding ​

POST /api/embed

Genera embedding da un modello

Parametri ​

  • model: nome del modello da cui generare gli embedding
  • input: testo o lista di testi per cui generare gli embedding

Parametri avanzati:

  • truncate: tronca la fine di ogni input per adattarlo alla lunghezza del contesto. Restituisce un errore se impostato su false e la lunghezza del contesto viene superata. Il valore predefinito è true
  • options: parametri aggiuntivi del modello elencati nella documentazione del Modelfile come temperature
  • keep_alive: controlla per quanto tempo il modello rimarrà caricato in memoria dopo la richiesta (predefinito: 5m)
  • dimensions: numero di dimensioni per l'embedding

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/embed -d '{
  "model": "all-minilm",
  "input": "Why is the sky blue?"
}'

Risposta ​

json
{
  "model": "all-minilm",
  "embeddings": [
    [
      0.010071029, -0.0017594862, 0.05007221, 0.04692972, 0.054916814,
      0.008599704, 0.105441414, -0.025878139, 0.12958129, 0.031952348
    ]
  ],
  "total_duration": 14143917,
  "load_duration": 1019500,
  "prompt_eval_count": 8
}

Richiesta (input multiplo) ​

shell
curl http://localhost:11434/api/embed -d '{
  "model": "all-minilm",
  "input": ["Why is the sky blue?", "Why is the grass green?"]
}'

Risposta ​

json
{
  "model": "all-minilm",
  "embeddings": [
    [
      0.010071029, -0.0017594862, 0.05007221, 0.04692972, 0.054916814,
      0.008599704, 0.105441414, -0.025878139, 0.12958129, 0.031952348
    ],
    [
      -0.0098027075, 0.06042469, 0.025257962, -0.006364387, 0.07272725,
      0.017194884, 0.09032035, -0.051705178, 0.09951512, 0.09072481
    ]
  ]
}

Elenca modelli in esecuzione ​

GET /api/ps

Elenca i modelli attualmente caricati in memoria.

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/ps

Risposta ​

Verrà restituito un singolo oggetto JSON.

json
{
  "models": [
    {
      "name": "mistral:latest",
      "model": "mistral:latest",
      "size": 5137025024,
      "digest": "2ae6f6dd7a3dd734790bbbf58b8909a606e0e7e97e94b7604e0aa7ae4490e6d8",
      "details": {
        "parent_model": "",
        "format": "gguf",
        "family": "llama",
        "families": ["llama"],
        "parameter_size": "7.2B",
        "quantization_level": "Q4_0"
      },
      "expires_at": "2024-06-04T14:38:31.83753-07:00",
      "size_vram": 5137025024
    }
  ]
}

Genera Embedding ​

Nota: questo endpoint è stato sostituito da /api/embed

POST /api/embeddings

Genera embedding da un modello

Parametri ​

  • model: nome del modello da cui generare gli embedding
  • prompt: testo per cui generare gli embedding

Parametri avanzati:

  • options: parametri aggiuntivi del modello elencati nella documentazione del Modelfile come temperature
  • keep_alive: controlla per quanto tempo il modello rimarrà caricato in memoria dopo la richiesta (predefinito: 5m)

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/embeddings -d '{
  "model": "all-minilm",
  "prompt": "Here is an article about llamas..."
}'

Risposta ​

json
{
  "embedding": [
    0.5670403838157654, 0.009260174818336964, 0.23178744316101074,
    -0.2916173040866852, -0.8924556970596313, 0.8785552978515625,
    -0.34576427936553955, 0.5742510557174683, -0.04222835972905159,
    -0.137906014919281
  ]
}

Versione ​

GET /api/version

Recupera la versione di Ollama

Esempi ​

Richiesta ​

shell
curl http://localhost:11434/api/version

Risposta ​

json
{
  "version": "0.5.1"
}

Funzionalità sperimentali ​

Generazione di immagini (sperimentale) ​

WARNING

La generazione di immagini è sperimentale e potrebbe cambiare nelle versioni future.

La generazione di immagini è ora supportata tramite l'endpoint standard /api/generate quando si utilizzano modelli di generazione di immagini. L'API rileva automaticamente quando viene utilizzato un modello di generazione di immagini.

Vedi la sezione Generate a completion per la documentazione completa dell'API. I parametri sperimentali per la generazione di immagini (width, height, steps) sono documentati lì.

Esempio ​

Richiesta ​
shell
curl http://localhost:11434/api/generate -d '{
  "model": "x/z-image-turbo",
  "prompt": "a sunset over mountains",
  "width": 1024,
  "height": 768
}'
Risposta (in streaming) ​

Aggiornamenti di progresso durante la generazione:

json
{
  "model": "x/z-image-turbo",
  "created_at": "2024-01-15T10:30:00.000000Z",
  "completed": 5,
  "total": 20,
  "done": false
}
Risposta finale ​
json
{
  "model": "x/z-image-turbo",
  "created_at": "2024-01-15T10:30:15.000000Z",
  "image": "iVBORw0KGgoAAAANSUhEUg...",
  "done": true,
  "done_reason": "stop",
  "total_duration": 15000000000,
  "load_duration": 2000000000
}