Wiki-Quellcode von Ollama
Zuletzt geändert von René Schmidt am 2025/03/16 11:12
Verstecke letzte Bearbeiter
| author | version | line-number | content |
|---|---|---|---|
![]() |
1.1 | 1 | [[~[~[image:https://private-user-images.githubusercontent.com/3325447/254932576-0d0b44e2-8f4a-4e99-9b52-a5c1c741c8f7.png?jwt=eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3NDIwNjQ2NzMsIm5iZiI6MTc0MjA2NDM3MywicGF0aCI6Ii8zMzI1NDQ3LzI1NDkzMjU3Ni0wZDBiNDRlMi04ZjRhLTRlOTktOWI1Mi1hNWMxYzc0MWM4ZjcucG5nP1gtQW16LUFsZ29yaXRobT1BV1M0LUhNQUMtU0hBMjU2JlgtQW16LUNyZWRlbnRpYWw9QUtJQVZDT0RZTFNBNTNQUUs0WkElMkYyMDI1MDMxNSUyRnVzLWVhc3QtMSUyRnMzJTJGYXdzNF9yZXF1ZXN0JlgtQW16LURhdGU9MjAyNTAzMTVUMTg0NjEzWiZYLUFtei1FeHBpcmVzPTMwMCZYLUFtei1TaWduYXR1cmU9NzM1ODM3ZjQ2MmZkOTE2MTBmYTJiOTAyMjVhMzVmOTcyYTRhNTI4MWNhYTY1NmJjMWU4ZGU2MGU3OTM4ZWUwZiZYLUFtei1TaWduZWRIZWFkZXJzPWhvc3QifQ.-CoNQVw6zTaJlUbBkKR_ofPwIhsKfHx0r9hKD6HOSVY~|~|alt="ollama" height="200px"~]~]>>url:https://ollama.com]] |
| 2 | |||
| 3 | = Ollama = | ||
| 4 | |||
| 5 | |||
| 6 | Get up and running with large language models. | ||
| 7 | |||
| 8 | === macOS === | ||
| 9 | |||
| 10 | |||
| 11 | [[Download>>url:https://ollama.com/download/Ollama-darwin.zip]] | ||
| 12 | |||
| 13 | === Windows === | ||
| 14 | |||
| 15 | |||
| 16 | [[Download>>url:https://ollama.com/download/OllamaSetup.exe]] | ||
| 17 | |||
| 18 | === Linux === | ||
| 19 | |||
| 20 | |||
| 21 | {{{curl -fsSL https://ollama.com/install.sh | sh}}} | ||
| 22 | |||
| 23 | [[Manual install instructions>>url:https://github.com/ollama/ollama/blob/main/docs/linux.md]] | ||
| 24 | |||
| 25 | === Docker === | ||
| 26 | |||
| 27 | |||
| 28 | The official [[Ollama Docker image>>url:https://hub.docker.com/r/ollama/ollama]] ollama/ollama is available on Docker Hub. | ||
| 29 | |||
| 30 | === Libraries === | ||
| 31 | |||
| 32 | |||
| 33 | * [[ollama-python>>url:https://github.com/ollama/ollama-python]] | ||
| 34 | * [[ollama-js>>url:https://github.com/ollama/ollama-js]] | ||
| 35 | |||
| 36 | === Community === | ||
| 37 | |||
| 38 | |||
| 39 | * [[Discord>>url:https://discord.gg/ollama]] | ||
| 40 | * [[Reddit>>url:https://reddit.com/r/ollama]] | ||
| 41 | |||
| 42 | == Quickstart == | ||
| 43 | |||
| 44 | |||
| 45 | To run and chat with [[Llama 3.2>>url:https://ollama.com/library/llama3.2]]: | ||
| 46 | |||
| 47 | {{{ollama run llama3.2}}} | ||
| 48 | |||
| 49 | == Model library == | ||
| 50 | |||
| 51 | |||
| 52 | Ollama supports a list of models available on [[ollama.com/library>>url:https://ollama.com/library]] | ||
| 53 | |||
| 54 | Here are some example models that can be downloaded: | ||
| 55 | |||
| 56 | |=Model|=Parameters|=Size|=Download | ||
| 57 | |Gemma 3|1B|815MB|ollama run gemma3:1b | ||
| 58 | |Gemma 3|4B|3.3GB|ollama run gemma3 | ||
| 59 | |Gemma 3|12B|8.1GB|ollama run gemma3:12b | ||
| 60 | |Gemma 3|27B|17GB|ollama run gemma3:27b | ||
| 61 | |QwQ|32B|20GB|ollama run qwq | ||
| 62 | |DeepSeek-R1|7B|4.7GB|ollama run deepseek-r1 | ||
| 63 | |DeepSeek-R1|671B|404GB|ollama run deepseek-r1:671b | ||
| 64 | |Llama 3.3|70B|43GB|ollama run llama3.3 | ||
| 65 | |Llama 3.2|3B|2.0GB|ollama run llama3.2 | ||
| 66 | |Llama 3.2|1B|1.3GB|ollama run llama3.2:1b | ||
| 67 | |Llama 3.2 Vision|11B|7.9GB|ollama run llama3.2-vision | ||
| 68 | |Llama 3.2 Vision|90B|55GB|ollama run llama3.2-vision:90b | ||
| 69 | |Llama 3.1|8B|4.7GB|ollama run llama3.1 | ||
| 70 | |Llama 3.1|405B|231GB|ollama run llama3.1:405b | ||
| 71 | |Phi 4|14B|9.1GB|ollama run phi4 | ||
| 72 | |Phi 4 Mini|3.8B|2.5GB|ollama run phi4-mini | ||
| 73 | |Mistral|7B|4.1GB|ollama run mistral | ||
| 74 | |Moondream 2|1.4B|829MB|ollama run moondream | ||
| 75 | |Neural Chat|7B|4.1GB|ollama run neural-chat | ||
| 76 | |Starling|7B|4.1GB|ollama run starling-lm | ||
| 77 | |Code Llama|7B|3.8GB|ollama run codellama | ||
| 78 | |Llama 2 Uncensored|7B|3.8GB|ollama run llama2-uncensored | ||
| 79 | |LLaVA|7B|4.5GB|ollama run llava | ||
| 80 | |Granite-3.2|8B|4.9GB|ollama run granite3.2 | ||
| 81 | |||
| 82 | |||
| 83 | Note | ||
| 84 | |||
| 85 | You should have at least 8 GB of RAM available to run the 7B models, 16 GB to run the 13B models, and 32 GB to run the 33B models. | ||
| 86 | |||
| 87 | == Customize a model == | ||
| 88 | |||
| 89 | |||
| 90 | === Import from GGUF === | ||
| 91 | |||
| 92 | |||
| 93 | Ollama supports importing GGUF models in the Modelfile: | ||
| 94 | |||
| 95 | 1. ((( | ||
| 96 | Create a file named Modelfile, with a FROM instruction with the local filepath to the model you want to import. | ||
| 97 | |||
| 98 | {{{FROM ./vicuna-33b.Q4_0.gguf | ||
| 99 | }}} | ||
| 100 | ))) | ||
| 101 | |||
| 102 | Create the model in Ollama | ||
| 103 | |||
| 104 | {{{ollama create example -f Modelfile}}} | ||
| 105 | |||
| 106 | Run the model | ||
| 107 | |||
| 108 | {{{ollama run example}}} | ||
| 109 | |||
| 110 | === Import from Safetensors === | ||
| 111 | |||
| 112 | |||
| 113 | See the [[guide>>url:https://github.com/ollama/ollama/blob/main/docs/import.md]] on importing models for more information. | ||
| 114 | |||
| 115 | === Customize a prompt === | ||
| 116 | |||
| 117 | |||
| 118 | Models from the Ollama library can be customized with a prompt. For example, to customize the llama3.2 model: | ||
| 119 | |||
| 120 | {{{ollama pull llama3.2}}} | ||
| 121 | |||
| 122 | Create a Modelfile: | ||
| 123 | |||
| 124 | {{{FROM llama3.2 | ||
| 125 | |||
| 126 | # set the temperature to 1 [higher is more creative, lower is more coherent] | ||
| 127 | PARAMETER temperature 1 | ||
| 128 | |||
| 129 | # set the system message | ||
| 130 | SYSTEM """ | ||
| 131 | You are Mario from Super Mario Bros. Answer as Mario, the assistant, only. | ||
| 132 | """ | ||
| 133 | }}} | ||
| 134 | |||
| 135 | Next, create and run the model: | ||
| 136 | |||
| 137 | {{{ollama create mario -f ./Modelfile | ||
| 138 | ollama run mario | ||
| 139 | >>> hi | ||
| 140 | Hello! It's your friend Mario. | ||
| 141 | }}} | ||
| 142 | |||
| 143 | For more information on working with a Modelfile, see the [[Modelfile>>url:https://github.com/ollama/ollama/blob/main/docs/modelfile.md]] documentation. | ||
| 144 | |||
| 145 | == CLI Reference == | ||
| 146 | |||
| 147 | |||
| 148 | === Create a model === | ||
| 149 | |||
| 150 | |||
| 151 | ollama create is used to create a model from a Modelfile. | ||
| 152 | |||
| 153 | {{{ollama create mymodel -f ./Modelfile}}} | ||
| 154 | |||
| 155 | === Pull a model === | ||
| 156 | |||
| 157 | |||
| 158 | {{{ollama pull llama3.2}}} | ||
| 159 | |||
| 160 | >This command can also be used to update a local model. Only the diff will be pulled. | ||
| 161 | |||
| 162 | === Remove a model === | ||
| 163 | |||
| 164 | |||
| 165 | {{{ollama rm llama3.2}}} | ||
| 166 | |||
| 167 | === Copy a model === | ||
| 168 | |||
| 169 | |||
| 170 | {{{ollama cp llama3.2 my-model}}} | ||
| 171 | |||
| 172 | === Multiline input === | ||
| 173 | |||
| 174 | |||
| 175 | For multiline input, you can wrap text with """: | ||
| 176 | |||
| 177 | {{{>>> """Hello, | ||
| 178 | ... world! | ||
| 179 | ... """ | ||
| 180 | I'm a basic program that prints the famous "Hello, world!" message to the console. | ||
| 181 | }}} | ||
| 182 | |||
| 183 | === Multimodal models === | ||
| 184 | |||
| 185 | |||
| 186 | {{{ollama run llava "What's in this image? /Users/jmorgan/Desktop/smile.png" | ||
| 187 | }}} | ||
| 188 | |||
| 189 | >**Output**: The image features a yellow smiley face, which is likely the central focus of the picture. | ||
| 190 | |||
| 191 | === Pass the prompt as an argument === | ||
| 192 | |||
| 193 | |||
| 194 | {{{ollama run llama3.2 "Summarize this file: $(cat README.md)"}}} | ||
| 195 | |||
| 196 | >**Output**: Ollama is a lightweight, extensible framework for building and running language models on the local machine. It provides a simple API for creating, running, and managing models, as well as a library of pre-built models that can be easily used in a variety of applications. | ||
| 197 | |||
| 198 | === Show model information === | ||
| 199 | |||
| 200 | |||
| 201 | {{{ollama show llama3.2}}} | ||
| 202 | |||
| 203 | === List models on your computer === | ||
| 204 | |||
| 205 | |||
| 206 | {{{ollama list}}} | ||
| 207 | |||
| 208 | === List which models are currently loaded === | ||
| 209 | |||
| 210 | |||
| 211 | {{{ollama ps}}} | ||
| 212 | |||
| 213 | === Stop a model which is currently running === | ||
| 214 | |||
| 215 | |||
| 216 | {{{ollama stop llama3.2}}} | ||
| 217 | |||
| 218 | === Start Ollama === | ||
| 219 | |||
| 220 | |||
| 221 | ollama serve is used when you want to start ollama without running the desktop application. | ||
| 222 | |||
| 223 | == Building == | ||
| 224 | |||
| 225 | |||
| 226 | See the [[developer guide>>url:https://github.com/ollama/ollama/blob/main/docs/development.md]] | ||
| 227 | |||
| 228 | === Running local builds === | ||
| 229 | |||
| 230 | |||
| 231 | Next, start the server: | ||
| 232 | |||
| 233 | {{{./ollama serve}}} | ||
| 234 | |||
| 235 | Finally, in a separate shell, run a model: | ||
| 236 | |||
| 237 | {{{./ollama run llama3.2}}} | ||
| 238 | |||
| 239 | == REST API == | ||
| 240 | |||
| 241 | |||
| 242 | Ollama has a REST API for running and managing models. | ||
| 243 | |||
| 244 | === Generate a response === | ||
| 245 | |||
| 246 | |||
| 247 | {{{curl http://localhost:11434/api/generate -d '{ | ||
| 248 | "model": "llama3.2", | ||
| 249 | "prompt":"Why is the sky blue?" | ||
| 250 | }'}}} | ||
| 251 | |||
| 252 | === Chat with a model === | ||
| 253 | |||
| 254 | |||
| 255 | {{{curl http://localhost:11434/api/chat -d '{ | ||
| 256 | "model": "llama3.2", | ||
| 257 | "messages": [ | ||
| 258 | { "role": "user", "content": "why is the sky blue?" } | ||
| 259 | ] | ||
| 260 | }'}}} | ||
| 261 | |||
| 262 | See the [[API documentation>>url:https://github.com/ollama/ollama/blob/main/docs/api.md]] for all endpoints. |
