ollama

Commit Graph

Author	SHA1	Message	Date
Devon Rifkin	c87b910232	WIP: stable ordering for tool args Right now we deserialize tool call definitions' arguments into golang maps. These purposefully don't have a predictable iteration order, whereas we want to maintain the order the user originally provided. Unstable rendering of arguments means that we break the kv cache, which this change fixes. There's no way to build this in a fully backwards compatible way when executing existing templates exactly as they are. We get around this by rewriting templates dynamically just before they're rendered. This is fragile, but perhaps the least bad option?	2025-10-07 15:38:58 -07:00
Devon Rifkin	bc71278670	Merge pull request #12509 from ollama/drifkin/oai-compat-refactor openai: refactor to split compat layer and middleware	2025-10-06 16:22:08 -07:00
Daniel Hiltgen	918231931c	win: fix build script (#12513 )	2025-10-06 14:46:45 -07:00
Daniel Hiltgen	04c1849878	discovery: prevent dup OLLAMA_LIBRARY_PATH (#12514 ) This variable isn't currently documented or intended as something the user can override, but if the user happens to set OLLAMA_LIBRARY_PATH we were doubling this in the subprocess environment which will cause problems with the new bootstrap discovery logic.	2025-10-06 14:36:44 -07:00
Devon Rifkin	2c2f4deaa9	openai: refactor to split compat layer and middleware This makes the core openai compat layer independent of the middleware that adapts it to our particular gin routes	2025-10-05 14:18:56 -07:00
Daniel Hiltgen	292767afb4	CI: fix win arm build (#12502 ) Resolve subtle erroraction stickiness difference between x86 and arm builder setup	2025-10-04 11:46:45 -07:00
Daniel Hiltgen	ae5e0f0889	CI: replace clang compiler for windows (#12495 )	2025-10-04 09:18:42 -07:00
Jesse Gross	19e6796eac	llm: Support KV cache quantization with gpt-oss With the new version of GGML in #12245, KV cache quantization no longer causes a fallback to CPU.	2025-10-03 16:31:58 -07:00
Grace	33801c1597	Fixed Deepseek2 adding nil tensor error	2025-10-03 14:20:06 -07:00
Daniel Hiltgen	e4340667e3	Workaround broken NVIDIA iGPU free VRAM data (#12490 ) The CUDA APIs for reporting free VRAM are useless on NVIDIA iGPU systems as they only return the kernels actual free memory and ignore buff/cache allocations which on a typical system will quickly fill up most of the free system memory. As a result, we incorrectly think there's very little available for GPU allocations which is wrong.	2025-10-03 12:17:21 -07:00
Patrick Devine	2fa1e92a99	test: add template error test (#12489 )	2025-10-03 12:05:34 -07:00
Daniel Hiltgen	07e36761c3	ci: place rocm windows in correct runner dir (#12487 )	2025-10-03 07:28:40 -07:00
Daniel Hiltgen	c29fb007c0	CI: temporarily disable clang install (#12486 ) This will likely yield builds that have problems with unicode characters but at least we can start testing the release while we try to find an alternate clang compiler for windows, or mingw ships a fixed version.	2025-10-02 20:31:18 -07:00
Daniel Hiltgen	730ed6e9e1	ci: fix windows build (#12485 )	2025-10-02 19:16:01 -07:00
Daniel Hiltgen	dc06601677	ci: fix windows build (#12484 )	2025-10-02 18:59:26 -07:00
Patrick Devine	1ed2881ef0	templates: fix crash in improperly defined templates (#12483 )	2025-10-02 17:25:55 -07:00
Jesse Gross	0bda72892c	llm: Enable flash attention by default for qwen3 and qwen3moe	2025-10-02 17:04:10 -07:00
Daniel Hiltgen	55ca827267	AMD: block running on unsupported gfx900/gfx906 (#12481 )	2025-10-02 16:53:05 -07:00
Daniel Hiltgen	c68f367ef6	Update GGML to b6646 (#12245 ) Notable EOLs with this change: - MacOS v12 and v13 are no longer supported (v14+ required) - AMD gfx900 and gfx906 are no longer supported	2025-10-02 14:47:10 -07:00
Jesse Gross	fdb109469f	llm: Allow overriding flash attention setting As we automatically enable flash attention for more models, there are likely some cases where we get it wrong. This allows setting OLLAMA_FLASH_ATTENTION=0 to disable it, even for models that usually have flash attention.	2025-10-02 12:07:20 -07:00
Daniel Hiltgen	05a43e078a	fix panic on bootstrapDevices (#12475 ) Wrong index variable was used.	2025-10-01 17:39:29 -07:00
Daniel Hiltgen	bc8909fb38	Use runners for GPU discovery (#12090 ) This revamps how we discover GPUs in the system by leveraging the Ollama runner. This should eliminate inconsistency between our GPU discovery and the runners capabilities at runtime, particularly for cases where we try to filter out unsupported GPUs. Now the runner does that implicitly based on the actual device list. In some cases free VRAM reporting can be unreliable which can leaad to scheduling mistakes, so this also includes a patch to leverage more reliable VRAM reporting libraries if available. Automatic workarounds have been removed as only one GPU leveraged this, which is now documented. This GPU will soon fall off the support matrix with the next ROCm bump. Additional cleanup of the scheduler and discovery packages can be done in the future once we have switched on the new memory management code, and removed support for the llama runner.	2025-10-01 15:12:32 -07:00
Devon Rifkin	6b50f2b9cd	Merge pull request #12461 from ollama/drifkin/qwen3-coder-tweaks qwen3-coder: fix tool definition type rendering	2025-09-30 19:47:44 -07:00
Michael Yang	35ac4eb12c	fix keep alive this reference to keep alive was missed in #12041 so chat has a diffferent behaviour than generate	2025-09-30 17:22:28 -07:00
Jesse Gross	3d0b1734c0	ggml: Preallocate CUDA pool memory The GGML CUDA backend allocates additional memory for intermediate results during calculation. This memory isn't currently allocated during worst case graph reservation and therefore not included in scheduling. This means that as these buffers potentially grow with context length, we could crash. This extends the memory allocation system down layer from the GGML graph to the CUDA layer, preallocating the worst case memory there as well. Fixes #11753	2025-09-30 15:04:43 -07:00
Jesse Gross	efaee8c2d6	ggml: Backport scale kernel fixes The GGML scale kernel uses signed 32-bit ints to represent the number of elements in the tensor. For large images, mistral-small3.2 overflows this, triggering CUDA errors due to negative arguments. Currently, this can happen when the user passes a large image to mistral-small3.2. However, with upcoming changes to reserve CUDA memory, it happens every time mistral-small is loaded as we reserve using a worst case batch. This patch is part of an upstream GGML commit and should be removed after GGML is updated past 0a1b398 "ggml: add ops for WAN video model (cuda && cpu) (#15669)". Fixes #10388	2025-09-30 15:04:43 -07:00
Jesse Gross	734b57da0e	ggml: Remove allocation status reporting For each memory allocation we report the size of the (attempted) allocation and whether it succeeded or failed. The latter status reporting proved to be not that useful in practice as systems such as Windows can automatically overflow from VRAM into RAM, resultings in successful allocations even when there isn't enough memory where we wanted. As a result, this information is only used for debug logging, which isn't worthwhile enough for the amount of code. It also isn't fully accurate, as multiple allocations may result in partial failures.	2025-09-30 15:04:43 -07:00
Devon Rifkin	83021fcf0f	qwen3-coder: fix tool definition type rendering	2025-09-30 15:03:15 -07:00
Michael Yang	0469861d9d	build: call find_package to instantiate library paths	2025-09-30 13:12:46 -07:00
羊撅撅	c47154c08d	fix: correct condition for AMDGPU_TARGETS filtering logic (#12412 )	2025-09-26 11:38:47 -07:00
Patrick Devine	b04e46da3e	bugfix: restore the current runOptions if loading fails in the CLI (#12402 ) There are two bugs when using `/load <model>` for a model that doesn't exist, namely: 1. it will not restore the current model settings if the current model is a thinking model; and 2. it will crash is the current model is a non-thinking model This bug fix saves the current runOptions and then restores them if the model load doesn't happen. It also fixes the crash happening for non-thinking models.	2025-09-25 18:30:45 -07:00
Devon Rifkin	34efbbd3f0	Merge pull request #12417 from ollama/drifkin/qwen3-coder-unicode parsers: fix unicode handling for qwen3-coder	2025-09-25 15:56:34 -07:00
Devon Rifkin	05ba4ca1f4	parsers: fix unicode handling for qwen3-coder When trimming whitespace at the end of every chunk, we were iterating backwards over the string byte-by-byte instead of rune-by-rune. As an example of how this can cause corruption, suppose we have the multi-byte character ✅ (`"\u2705"`), which is represented in utf-8 as the three bytes `0xE2 0x9C 0x85`. It happens that `0x85` is NEL, which passes `unicode.IsSpace()`. Because we were iterating byte-by-byte, this caused us to mistakenly slice in the middle of the rune, removing `0x85` and leaving `0xE2 0x9C`, which beyond being the incorrect place to slice, is not even a valid utf-8 character. `trailingWhitespaceLen()` was modified to count from the end in a rune-aware way. Tests with various multibyte unicode characters were also added. Fixes: #12414	2025-09-25 15:47:46 -07:00
Patrick Devine	5a56ff3cf0	cli: add device signin flow when doing ollama push (#12405 )	2025-09-25 15:04:43 -07:00
Gabe Goodhart	2fba04b5fb	tools: handle the case where a tool call sends "arguments" or "parameters" as a serialized json string (#12413 )	2025-09-25 14:37:39 -07:00
Grace	fbd82ba5bb	Grace/deepseek v3 migration (#12385 ) * init deepseek model file * temp removal of flash attention implementation * shapes and proper, can make a pass * query, key, value have good cosine similarity, but the max diff is a bit high * Attention block is working! ** with eager for now, have not added the mask line * Attention block is working! ** with eager for now, have not added the mask line * working MoE at around 0.95 cosine sim * added cosine similarity function * Starting end to end structure * Trying (and failing) to get rope to work, going to test full thing on tater * running on tater36... just not the right outputs * we have the right values for rope... but its still not working? * chnage Extrapolation Factor to 1 * removed adding residuals twice, removed normalization from shared expert, refactored Norms (Attention, MLP) to be outside the (Attention, MLP) blocks and in the Transformer block instead, add cache setLayer * Temporary modelfiles for cpu * change kpass intermediate step to kv, two layer outputs [0,1] look fine * this calls for 16 chicken nuggets * whoops * cleaning up code * delete stuff we dont need * getting rid of debug statements for llama cpp * working with long contexts * fix long context view error * reverting some changes I made for files that are not apart of pr * Added proper tokenizer for deeepseek3 * clean up model and go test * remove Modelfile * not passing the tests * whoops * how to pass the ci tests * resolving some of the comments * rename * linted and renamed deepseek3 -> deepseek2 * remove name go * addressed changes - main change was adopting qwen3 naming scheme * I cannot with linters * clean up logs * clean up logs --------- Co-authored-by: Grace Guo <graceguo@Graces-MBP.localdomain> Co-authored-by: Grace Guo <graceguo@Graces-MacBook-Pro.local> Co-authored-by: graceguo <graceguo@tater36.localdomain>	2025-09-24 15:19:47 -07:00
Michael Yang	2e742544bf	prefer ollama engine for qwen3moe (#12374 )	2025-09-24 11:21:32 -07:00
Devon Rifkin	bbb195a6ff	Merge pull request #12393 from ollama/drifkin/fix-built-ins harmony: don't sanitize built-ins	2025-09-23 23:45:31 -07:00
Devon Rifkin	fd88cd7cb0	harmony: don't sanitize built-ins In #11910 we started sanitizing function names, but we accidentally were modifying built-ins like `browser.open` to `browser_open`. This was removing the special prompt rendering for built-ins, but this wasn't immediately apparent since the models seem to be reasonably good at remembering the built-ins even when presented with these slightly renamed version. This fix prevents built-ins from ever being renamed.	2025-09-23 23:34:55 -07:00
Michael Yang	e1979c571a	fix: leaf alt name (#12390 ) a leaf node with an alternative name gets all its alternatives names added into the same branch rather than creating branches themselves	2025-09-23 17:50:53 -07:00
Michael Yang	bf78ed6ee9	add pre:, suf: to tags (#12274 )	2025-09-23 16:08:57 -07:00
Michael Yang	a40d427bce	multi-regexp pretokenizer (#12325 )	2025-09-23 13:21:47 -07:00
Patrick Devine	64883e3c4c	auth: fix problems with the ollama keypairs (#12373 ) * auth: fix problems with the ollama keypairs This change adds several fixes including: - reading in the pubkey files correctly - fixing the push unit test to create a keypair file in a temp directory - not return 500 errors for normal status error	2025-09-22 23:20:20 -07:00
Devon Rifkin	41efdd4048	Merge pull request #12339 from ollama/drifkin/harmony-refactor-to-builtin harmony: remove special casing in routes.go	2025-09-22 13:13:40 -07:00
Daniel Hiltgen	c23e6f4cae	tests: add single threaded history test (#12295 ) * tests: add single threaded history test Also tidies up some existing tests to handle more model output variation * test: add support for testing specific architectures	2025-09-22 11:23:14 -07:00
jmorganca	af060eb250	docs: update cloud.md for cloud models	2025-09-22 13:09:17 -03:00
jmorganca	ae5c33008e	docs: move turbo.md to cloud.md	2025-09-22 13:09:17 -03:00
Devon Rifkin	3677842ff1	Merge pull request #12358 from ollama/drifkin/qwen3-coder-ampersands parsers: fix `&`s in qwen3coder parameter values	2025-09-20 12:40:33 -07:00
Devon Rifkin	242df70a75	parsers: fix `&`s in qwen3coder parameter values In <https://github.com/ollama/ollama/issues/12357> we that the model will output tool calls such as ``` <function=shell> <parameter=command> pwd && ls -la </parameter> </function> ``` We parse this using the approach of transforming into valid xml and then using an xml parser. While we do transform the function and parameter names, we weren't escaping the parameter values (which in this example are invalid since `pwd && ls -la` contains unescaped ampersands). This has been fixed by first transforming the tags in the same way, and then walking the transformed string and escaping the text in between the tags. This also fixes a case where `<` in the middle of a parameter value would cause an xml parse failure. Fixes: #12357	2025-09-20 12:11:38 -07:00
Patrick Devine	dba39b2eee	gemma: fix rope scaling for qat models (#12348 ) * gemma: fix rope scaling for qat models * gofumpt yourself	2025-09-19 15:04:40 -07:00

1 2 3 4 5 ...

4613 Commits All Branches Search

4613 Commits

All Branches