Vulnerabilities
Ollama Security Crisis: Update to v0.17.1 Immediately to Block Unauthenticated Memory Leaks
Cyber RTMay 11, 20263 min read

Cybersecurity researchers revealed a critical vulnerability in Ollama, an open-source framework for running large language models locally, affecting over 300,000 servers globally. The flaw, CVE-2026-7482, allows remote attackers to leak process memory. Additionally, two unpatched vulnerabilities in Ollama's Windows update mechanism enable persistent code execution. Users should apply fixes, limit network access, and disable automatic updates to mitigate risks.
Cybersecurity researchers have identified a critical vulnerability in Ollama, a popular open-source framework for running large language models (LLMs) locally. This vulnerability, known as CVE-2026-7482 and codenamed Bleeding Llama, could allow a remote, unauthenticated attacker to access the entire process memory of affected servers. With a CVSS score of 9.1, this flaw is considered highly severe and potentially impacts over 300,000 servers worldwide.
Ollama's vulnerability is linked to an out-of-bounds read flaw in its GGUF model loader. The issue arises when the /api/create endpoint accepts a GGUF file with a tensor offset and size that exceed the file's actual length. This leads to the server reading beyond the allocated heap buffer during quantization, posing a significant security risk. GGUF, or GPT-Generated Unified Format, is a file format used to store LLMs for local execution, similar to formats like PyTorch and ONNX.
The root of the problem lies in Ollama's use of the unsafe package in the "WriteTo()" function when creating a model from a GGUF file. This bypasses the memory safety guarantees of the programming language, allowing for potential exploitation. An attacker could craft a GGUF file with an inflated tensor shape and send it to an exposed Ollama server, triggering the out-of-bounds heap read during model creation via the /api/create endpoint.
Successful exploitation could result in the leakage of sensitive data from the Ollama process memory, including environment variables, API keys, and user conversation data. This data could then be exfiltrated by uploading the resulting model artifact through the /api/push endpoint to an attacker-controlled registry. The attack chain involves three steps: uploading a crafted GGUF file, activating model creation to exploit the vulnerability, and exfiltrating data to an external server.
To mitigate the risk, users are advised to apply the latest security patches, restrict network access, and audit running instances for internet exposure. Additionally, deploying an authentication proxy or API gateway in front of Ollama instances is recommended, as the REST API lacks built-in authentication. These measures can help protect against potential data breaches and unauthorized access.
In a related development, researchers at Striga have uncovered two unpatched vulnerabilities in Ollama's Windows update mechanism, which could lead to persistent code execution. These vulnerabilities, CVE-2026-42248 and CVE-2026-42249, involve a missing signature verification and a path traversal issue, respectively. They allow an attacker to execute arbitrary code at every login by influencing update responses.
To exploit these flaws, an attacker would need control over an update server accessible by the victim's Ollama client. This could result in an arbitrary executable being installed in the Windows Startup folder without proper signature verification. Users are advised to disable automatic updates and remove any Ollama shortcuts from the Startup folder to prevent silent, persistent code execution.
Overall, these vulnerabilities highlight the importance of robust security practices and timely updates in software development. Users and organizations relying on Ollama should take immediate action to secure their systems and protect sensitive data from potential exploitation.


