What attackers actually do to exposed Ollama servers
A honeypot run for 84 days logged 290,000+ interactions on exposed Ollama-style endpoints, revealing real attacker behavior beyond simple scanning.
What is this?
On September 24, 2026, researchers Karina Elzer, Niklas Netterstrøm Johansen, and Emmanouil Vasilomanolakis published “OllamaDrama” on arXiv (2609.29757), describing Ollure, a honeypot built to emulate the Ollama API and measure, empirically, what attackers do once they find exposed LLM infrastructure on the open internet. Unlike the exposure scans covered here in May 2026 (“One million exposed AI services”), which counted how many misconfigured deployments exist, Ollure watches what happens after discovery: it ran for 84 days across four cloud and university network locations, logging 290,887 interactions from 2,793 unique source IPs. It corroborates smaller, independent observations, including a March–April 2026 honeypot run documented by researcher Marco Pedrinazzi that logged 6,461 events from 324 IPs over 32 days.
How it works
Ollure is a low-to-medium interaction honeypot: it speaks the Ollama HTTP API convincingly enough to be fingerprinted as a real deployment, but it has no backend LLM actually generating text, so its exposure is safe by construction. Traffic split into two broad phases. Most volume was reconnaissance: automated discovery scans, service fingerprinting, and model-enumeration requests probing which models a target claims to serve (/api/tags, /api/show). A smaller but more interesting share was active exploitation. Reported categories include abuse of model-management endpoints (/api/pull, /api/create, /api/delete) — in some cases crafting malicious “modelfiles” to attempt local file disclosure — path traversal and SSRF probes against fields that accept URLs, injected RCE-style payloads, cryptomining setup attempts, resource-exhaustion requests, prompt-injection strings aimed at any downstream LLM behavior, attempts to extract configuration and secrets, and traffic consistent with agent frameworks probing tool-use surfaces rather than a human typing prompts. Specific payload strings are not reproduced here, consistent with the paper’s own responsible framing of the dataset.
Why it matters
The gap this closes is significant for risk modeling: security teams have had scan-based evidence that thousands of Ollama, Open WebUI, Flowise, and similar self-hosted AI stacks sit exposed with no authentication, but not evidence of what happens next. Ollure shows the answer is “quite a lot, quite fast” — reconnaissance begins within hours of exposure, and a meaningful fraction of visitors go on to attempt concrete exploitation rather than passive scanning. The presence of SSRF probes and modelfile-based file-disclosure attempts is a reminder that Ollama’s management API was designed for a trusted single-operator context, not for exposure to the open internet, and that treating an LLM runtime as “just an API that answers prompts” understates its actual attack surface.
Defenses
Do not expose Ollama, or any self-hosted inference server’s management API, directly to the internet: bind to localhost or a private network and put authentication in front of any external access. Disable or firewall model-management endpoints (/api/pull, /api/create, /api/push, /api/delete) separately from inference endpoints where the software allows it, since they carry file- and network-side-effect risk that a simple prompt endpoint does not. Validate and restrict any user-suppliable URL fields (used in pull/push-style operations) against SSRF, including blocking cloud metadata IP ranges. Run the inference process in a container or sandbox without Docker socket or Kubernetes service-account access, since credential-harvesting attempts specifically target those. Monitor for the reconnaissance signature described here — bursts of /api/tags and “hello”-style liveness probes — as an early indicator that a deployment has been discovered, not a benign false positive to ignore.
Status
| Item | Detail |
|---|---|
| Honeypot | Ollure (low/medium interaction, no backend LLM) |
| Study duration | 84 days, 4 network locations |
| Interactions logged | 290,887 from 2,793 unique IPs |
| Paper | arXiv:2609.29757, submitted 2026-09-24 |
| Corroborating study | InTheCyber Ollama honeypot, 6,461 events / 324 IPs, Mar–Apr 2026 |